[["This paper constructs a theoretical model which captures the recent slowing-down of Chinese economy. In contrast with the previous literature which largely confines its focus on the resource misallocation between inefficient state-owned enterprises (SOEs) and more efficient private firms under a closed economy setting, this paper re-examines the dynamics of the growth of Chinese economy from the perspective of an open economy. In particular, this paper incorporates heterogeneous outputs and relative prices into the model, where private firms are assumed to be the major exporters and the remaining large SOEs create increasing import demand from the home country. By adding downward sloping world demand curve, our paper predicts a turning point during the transition process, as the falling relative price for exports starts to constrain and eventually slow down the growth; SOEs begin to co-exist with private firms in the economy before it is fully transformed. Our paper provides a theoretical foundation in terms of understanding the current dynamics and institutional change of Chinese economy. Additionally, this paper also provides quantitative evidence on the effects of financial development during the China's economic transition process. --------------------------------------------------------------------------------","Our paper is related to several studies in the previous literature. The closest work to our paper is done by Song et al. (2011) who also constructs a growth model to capture the dynamics of the transition of Chinese economy. The model developed in this paper differs from the one by Song et al. (2011) in the following aspects: First, instead of assuming the homogeneous output in the baseline model, our model incorporates heterogeneous outputs and relative prices into the analysis, with private firms assumed to be the major exporters and the remaining large SOEs create increasing import demand from the home country. Second, we add more realistic intratemporal decisions so the domestic consumption/investment now consists of two different goods. (SOE and Private firm good) These two extensions allow us to illustrate the relevance of the potential occurrence of a turning point within Chinese economy, after which the economy would slow down caused by the co-existence between SOE sectors and private firms. This is drastically contrasted from Song et al. (2011)'s model in which they predict the remaining SOE sectors would be crowded out by high-productivity private firms and the whole economic transition process would be completed smoothly. Lin et al. (1994) argues that the China's economic miracle is due to the state's appropriate adoption of comparative advantage following development strategies such that the labour-abundance factor endowment structure of Chinese economy is fully utilized. Our paper partially agrees with his views; however, it also goes one step further to ask whether such comparative-following advantage development strategies are sustainable in the long run. This is because according to the argument in our paper, comparative advantage following development strategies would inevitably make the exporting sectors become the priority within the economy and the over-focus on the biased growth pattern toward exporting would lower the relative price of the home country due to the more intense competition; this could finally lower the profitability of most of private firms which are largely concentrated within the exporting sectors. The decline in profitability of private firms could make the economic growth of Chinese economy in the long run become unsustainable. Papers by Lardy (2007), Kuijs (2005) and Aziz (2006) have rationalized various of reasons of why Chinese economy has failed to translated itself into the consumption-driven economic growth pattern and been heavily cling to the investment and exporting biased growth mode. For instance, both Kuijs (2005) and Aziz (2006) illustrate the relevance of the role of income disparity in determining current growth pattern of Chinese economy. Other scholars include Riedel et al. (2007) and Boyreau-Debray and Wei (2005) who further argue the underdeveloped financial sector in China might be one of the fundamental causes of the failure of transformation for Chinese economy to step toward the consumption-driven economic growth pattern. What sets our paper apart from their work is that we consider the resource allocation between high- productivity private firms and low productivity SOEs as the main factor in explaining the long-run economic performance of Chinese economy and we also incorporate the role of financial sectors in affecting the growth of Chinese economy into the proposed framework. Hsieh and Peter (2009) demonstrate slow growth of total factor productivity (TFP) of Chinese economy since the 2008 is largely caused by the resource misallocation across private and state sectors, which in turn resulted into the lower aggregate total factor productivity of Chinese economy. We partially agree with their views in the sense the damping effect of state-sectors on private counterparts would certainly lead to the resources flowing into low productivity state sectors, which leads to the resource misallocation within Chinese economy. Nevertheless, what makes our paper distinctive from their work is we demonstrate in a growth model to show that resource misallocation would not only lower the aggregate TFP of Chinese economy, but also directly triggered the inferior performance of exports as well as the lower productivity of private firms, which both factors contribute to the slowing-down of Chinese economy. Paper by Hsieh and Song (2015) further confirm the aforementioned points that, arguing restructuring of large SOEs and shrinkage of state sectors since the late 1990s to 2008 has been responsible for the dramatic growth of Chinese economy during the period including 20 percent of aggregate TFP growth. Paper by Lin et al. (2016) presents a rather optimistic outlook regarding the future growth prospects of Chinese economy. They argue the current slowing-down of Chinese economy is more caused by the structural factors external to the Chinese economy, such as the decline of growth of the rest of the world economy. Although it is true the sluggish growth of other countries with particular reference to Western economies has been persistently responsible for the lower growth of Chinese economy, they seem to ignore the structural factors stemming from the internal issues of Chinese economy such as the expansion of state-sectors and its ensuing crowding out effect on private sectors. Our paper argues that the internal and external factors which contribute to the slowing-down of Chinese economy are inherently intertwined. This is because according to the model developed in our paper, once the private sectors are crowded out by state sectors, the external export would be also negatively affected as most of private sectors concentrate within the exporting sectors. The rest of the paper will be organized as follows: The second section offers the empirical motivation for our theoretical model. Sections 3 and 4 describe the theoretical model and model simulation results in greater detail. Reverse trend between GDP growth and large SOEs ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Regarding the empirical motivation of this study, the slowing-down of Chinese economic growth in the recent years is one of the first observations from the data which triggers the interests of this study. The following Fig. 1 shows the decline in China's annual GDP growth rate (in constant prices) over the past decade. From Fig. 1, it is apparent China's annual GDP Growth Rate (in constant prices) has declined from 11.4 percent in 2005 to 6.6 percent in 2018. One of the argument proposed in this paper is that the resource misallocation caused by the co-existence between private sectors and state sectors is crucial for us to understand the current slowing-down of Chinese economy. Therefore, it is important to look at the evolution of the state-sector in Chinese economy. With respect to the proportions of state sector within the Chinese economy, the data also exhibits some interesting patterns: Fig. 2 has a adverse trend to Fig. 1. From 2005 to 2007, the number of SOEs is decreasing while the GDP growth rate increases in this period. From 2011 to 2015, number of SOEs increase and GDP growth rate decreases in this period. This implies that state sector in China is represented by large SOEs which have co-existed with private sectors in the past decade, but which might potentially contribute to the resource misallocation within Chinese economy. One of the most important features of such resource misallocation is embodied by the huge liabilities against large SOEs. Fig. 3 shows the increasing tendency of the amount of both liabilities and assets of large SOEs in China since 2005. Note: Prior to 2007, this number includes all the industrial SOEs. From 2007 to 2011, large SOEs are defined as SOEs whose operating incomes are above 5 million RMB. After 2011, large SOEs are those whose operating incomes are above 20 million RMB. Note: All the number has been transferred to the price level of 1999. The dramatic increase in the amount of liabilities of large SOEs in China since the past decade signifies the fact the financial sector including banking system is biased towards large SOEs. This gives large SOEs the priority in of bank loan lending as banks are also stated owned. Impact of co-existence on GDP growth ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In this section, the paper explores the effect of co-existence between SOEs and privately owned enterprises (POEs) on the GDP growth using empirical analysis. This paper assumes the co-existence of SOEs impedes POEs’ performance, reduces the export, and decreases GDP growth. Data and model construction This paper uses the economic data of 31 provinces in mainland China1 between 2005 and 2017. All the variables in this paper are collected from the regional economic database within the Chinese Stock Market and Accounting Research Database (CSMAR). The industrial data about private firms begin in 2005 and ends at 2017 in the database so that we choose the data from 2005 to 2017. Besides, most of variables are converted to the price level of 1999 to reduce the impact of inflation. To alleviate the impact of extreme value, all the variables are winsorized at the 1% level. To measure the co-existence between SOEs and POEs, this paper employs the ratio of operating sales in state-owned industrial firms to that in big-and- medium-sized industrial firms. A higher value of co-existence represents SOEs occupy a large proportion in the economy. All the variables definitions are summarized in Table 1 and the summary statistics are shown in Table 2. According to the summary statistics of Table 2, private industrial firms have average 9.6% return on total assets and 54.7% leverage in the sample. Crowding out effect of co-existence on POEs’ performance This paper employs ROA of private firms as the first proxy of economic performance of POEs and shows the result in Table 3. Model 1 only controls year fixed effect and model 2 include the remaining control variables. Coefficients of co-existence are negative and significant at the 1% level, indicating a negative effect of co- existence on earnings capacity of POEs. In other words, POEs in the provinces with high co-existence have lower earnings capacity. The regression result for the second proxy for economic performance of POEs is the operating sales of private industrial firm is presented in Table 4. Co-existence have negative and significant coefficients in model 1 and model 2, proving that co-existence is detrimental for the private firms' sales. Combining the results found in Table 3, this paper concludes that co-existence has crowding out effect for private firms’ performance. Impeding effect of co-existence on export In this section, the paper explores the impact of co-existence on exportation performance, which is a key driver for economic growth. In Table 5, co-existence has negative and significant coefficients with export, illustrating that co- existence has a negative relationship with exportation. In other words, a higher level of co-existence reduces the total exportation. In the next section, this paper examines the impact of co-existence on total exportation and importation. Results are presented in Table 6. Coefficients of co-existence are negative and significant at 1% level, similar to the results of Table 5. This result demonstrates that high co-existence would damage the exportation. Impeding effect of co-existence on GDP In the prior analysis, the empirical results show high levels of co-existence is detrimental to the development of POEs and exportation, implying co-existence might damage the economic development. To examine this expectation, this paper carries out further empirical analysis and presents the result in Table 7 and Table 8. Table 7 harnesses GDP as the dependent variable and shows negative and significant coefficients of co-existence. This result indicates high co-existence impedes the GDP. We employ GDP per capita as the alternative measure for economic development and shows the result in Table 8 still has negative and significant coefficients, in line with prior analysis. These results prove the detrimental effect of co-existence on economic development. Our empirical results find a higher level of co-existence impedes POEs’ performance, export performance, and economic growth.","Developing a theoretical framework which is suitable for transition economies like China would be the key for our analysis. Song et al. (2011) considers a small-open economy model (SSZ model), which provides a good explanation for the Chinese growth story over the past 20 years. The micro-foundations of their model relies on the OLG model, but they extend the simple OLG structure with heterogeneous agents (workers and entrepreneurs) and divide the industrial sector into two types of producers: F firms and E firms. F firms are the less efficient, financially integrated, and state owned enterprises, while the E firms are more productive, credit-constrained, private firms. After economic transition takes off, resources are reallocated from less efficient firms to more efficient ones within the industrial sector. During the transition periods (i.e. resources are reallocated from F firms to E firms), the economy keeps accumulating foreign assets, as the aggregate domestic investment shrinks overtime. The financial development can be reflected by the falling iceberg costs, or alternatively, an increase in the financial access for E firms. Transition dynamics ~~~~~~~~~~~~~~~~~~~ It also guarantees that E firms prefer delegation to direct control and young entrepreneurs are motivated to invest their savings in the production of E firms rather than deposit them in the financial intermediaries. The threshold productivity level is negatively related to the current relative price. Therefore, when the relative price falls, the speed of transition gradually slows down and the transition will eventually be ceased at a certain point. In addition to the capital growth in E firms, employment growth is also positively related to the growth of price levels. Given the complexity of the dynamics, quantitative simulation is required to examine the both capital and employment growth in E firms. This implies that the capital growth rate of F firms relies on the employment shares of E firms in this economy. If the employment share of E firms grows faster than the population growth rate (i.e. during the economic transition), the capital growth in F firms declines, and vice versa. Scenario 1 - adding heterogeneous outputs and relative price ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The calibrated parameters are given in Table 9. Imbs and Mejean (2010) provides cross- country evidence on the export price elasticity, with the estimates for most of the country lie between −3 and −4. We choose the upper bound, −3, in the simulation which provides reasonable dynamics for most of the variables. Furthermore, it seems natural for us to assume the initial relative price to be 1 in both extensions so they can be consistent with the baseline model at the initial position. This assumption is also important for the initialisation of all the other economic variables in all our simulations. In order to avoid sharp price falls, we take the log form of the total exports in our price simulation and set an upper limit for the relative price at 1 which further smooths the price movement during the simulation14. In this simulation, the growing production from E firms will shift the world supply curve to the right and lower the relative price after the transition takeoff. This further decreases the rate of return to capital in E firms and lessens their profitability and entrepreneurial savings. Therefore, the transition process is slowed down. Results ~~~~~~~ In the baseline model, the transition process is primarily reflected by the growth of the employment levels in E firms. During economic transition, we expect the efficient E firms to gradually outgrow the inefficient F firms. This is directly reflected by the baseline simulation in Panel E and Panel F (Fig. 4). On the contrary, our model experiences economic growth slow down, the transition slows down after year 1997 in first extension and, for the second extension, this slowdown comes after year 2000. The slow down in economic growth and trade activities are indeed observed in China in recent years. Hence our model better fits the growth experience post 2008 financial crisis. On the other hand, the employment of F firms are more persistent in our model. E firms face difficulties to carry on the transition after certain point and F firms keep the dominant position on employment levels. According to the assumption there is no unemployment in this economy16 and the amount of working-age workers is fixed, the employment shares of E firms and F firms always add up to the full employment level. This verifies our claims in the propositions. In our model, the full transition equilibrium is not reached, which means that life-time earnings of workers are permanently lower than the baseline simulation at new equilibrium, due to the existence of less efficient F firms. This provides answers to the puzzling differences between the aggregate consumption/saving of workers in Figs. 6 and 7. The wage differential and entrepreneurial savings are the main drivers to trigger the economic transition in the SSZ model; now both factors increase by much less in our model, which inevitably slows down the transition process. Panel E and Panel F in Fig. 5 present staggering differences on the capital account positions among three simulations. The capital flows in the model are determined by the differences between the domestic savings and domestic investment. There are drivers behind this: 1. F firms are financially integrated and its production technology requires higher levels of investment than E firms. More importantly, the F firms remain dominant; 2. The increasing savings from the entrepreneurs (managers) are the major sources for the future investment in E firm, as they are financially constrained, but now the falling relative price discourages the savings from entrepreneurs by lowering the rate of return in E firms and managerial wage. Therefore, our model predicts the possibility of capital inflow rather than outflow19 after certain tipping point. This suggests a possibility to rebalance the position of foreign reserves in a country20. To understand the causes, we need to analyse in more detail the aggregate saving and investment dynamics in the model. As the aggregate investment levels are driven by changes of aggregate capital levels, the trend of aggregate investment in both firms should be consistent with the trend of the aggregate capital levels in both firms. This is confirmed by Panel D and Panel E (Fig. 6). The pattern of aggregate investment is dominated by the pattern of the aggregate investment of F firms in our extensions, while in the baseline, this was dominated by the aggregate investment of E firms. On the other hand, as we have explained earlier, the savings of entrepreneurs grow much slower than the baseline model because of the falling rate of return to capital and managerial wage. This is indeed observed in Panel A (Fig. 6), and the savings of workers increase by much more in our model. Given the large number of workers within the economy, the trend of aggregate savings is inevitably dominated by their savings. The changes in asset levels follow the law of motion. The difference in aggregate asset positions between our model and the baseline simulation is entirely due to the lower life-time earnings of workers, as in our model, the transition process reaches an equilibrium level with the existence of both E firms and inefficient F firms. With less lifetime earnings, the workers need to accumulate more wealth than the baseline to smooth their consumption across time. The aggregate asset positions of workers coincide between two extensions is a direct consequence of the identical long run labour wage. The differences in the saving levels of workers are due to the different capital stock levels in E firms and F firms. We noticed that the savings of workers are higher in the extensions, but they rise less dramatically than the aggregate investment in F firms, which explains the predictions of net capital inflow in our extensions. The divergence between the consumption of workers in our extensions and the baseline (Panel B in Fig. 7) is also caused by the differences of lifetime earnings between the simulation of our extensions and the baseline simulation. With less earnings, the workers consume less in our model. The lower consumption levels for entrepreneurs follow a similar logic, as we have seen in Fig. 5, the managerial wage grows by much less, which results lower consumption. With lower consumption, the asset accumulation for entrepreneurs appears to grow faster in our model at the beginning. However, the baseline soon surpasses both scenarios, as E firms grow much faster and offer increasingly higher managerial wages over time, which enable the wealth accumulation process to speed up after a certain point.","The economic growth in China has been heavily reliant on investment and exports for quite some time. The failure to translate the benefits of economic growth into domestic consumption calls for urgent action to rebalance the Chinese economy21. Our model explores the theoretical possibilities of the Chinese growth experiences with the coexistence of both SOEs and private firms. Our work also enriches the understanding about the interactions and interdependence between economic transition and international trade. We first introduce heterogeneous outputs into the SSZ model by assuming SOEs (F firms) and private firms (E firms) produce two different products, with the latter specialises in exporting only. Another major difference is the inclusion of relative price between two goods and the downward sloping world demand curve. The relative price levels is endogenised into many key variables, such as the managerial wage, the rate of returns to capital in E firms, the saving rate of entrepreneurs, and E firm's capital per effective unit of labour. Therefore, the private firms cannot completely outgrow the SOEs during the transition due to the falling relative price. If the increasing exports from home country lowers the relative price to a threshold level the transition process could even come to a halt. In addition, we develop a more realistic set of micro-level foundations by allowing intratemporal decisions for the consumption demand of workers and entrepreneurs and the investment demand of E firms and F firms. The essential idea is to allow domestic consumers and firms to consume and investment two different goods domestically so that a falling relative price will also raise its domestic demand. Hence the price movements are less dramatic than the first extension. Another important result is related to China's accumulation of foreign reserves and net capital flows. In contrast to the popular argument that the growing foreign reserves in China are due to the manipulation of its exchange rates and export-led growth, the financial repression plays a key role in understanding the foreign reserve accumulation in China. Our model predicts a failure of full transition. Hence it is compatible with the recent empirical observation. China's foreign reserves have been falling continuously from US$3.8 trillion in 2014 to US$3.1 trillion in 2017. Net capital outflow at the beginning of economic transition can be reversed if the transition hits the turning point."],["Two firms engage in price competition to attract buyers located on a network. The value of the good of either firm to any buyer depends on the number of neighbors on the network who adopt the same good. When the size of externalities increases linearly with the number of adoptions, we identify the set of pricing strategies that are consistent with an equilibrium in which one of the firms monopolizes the market. The set includes marginal cost (MC) pricing as well as bipartition pricing, which offers discounts to some buyers and charges markups to others. We show that MC pricing fails to be an equilibrium under non-linear externalities in a general network, but identify conditions for an equilibrium with bipartition pricing to be robust against perturbations in the externalities from linearity. The analysis is applied to platform competition in a two-sided market under local and approximately linear externalities. --------------------------------------------------------------------------------","Goods have network externalities when their value to each user depends on the adoption decisions of others. Consumption goods such as social networking services, online game and auction platforms, and social and entertainment events all exhibit positive network externalities through direct interaction among their users, whereas goods such as computer OS and housing developments also exhibit positive externalities through the provision of complementary goods such as applications packages, shops and public transportation. Externalities are important also for intermediate goods. As argued by Carvalho (2014), for example, modern production is an intricate network of firms. In such a network, a single supplier of inputs serves multiple downstream firms who are themselves linked with each other in the form of mutual production of final products or technology transfers. A downstream firm would find a higher value for the same inputs as used by other downstream firms that are linked to it. Despite their importance in reality, our understanding of network externalities is limited when those goods are supplied competitively. The objective of this paper is to study price competition in the presence of such externalities: We formulate a highly stylized model of price competition in which users located at the nodes of a network experience positive externalities when their neighbors adopt the same good or technology. In our model, two symmetric firms each supply goods or technologies that are incompatible with each other. Users of either good experience larger positive externalities when more of their neighbors in the network adopt the same good. In stage 1, the two firms post prices simultaneously. The prices can be perfectly discriminatory and negative, and are publicly observable. In stage 2, the buyers simultaneously decide which good to adopt or not to adopt either. When no network externalities are present, it is clear that the unique subgame perfect equilibrium (SPE) of this game has both firms offer c, the constant marginal cost (MC) of production, to all buyers. We find that marginal-cost pricing is consistent with an SPE with monopolization by one of the firms in an arbitrary network when the externalities are linear in the number of neighbors adopting the same good. In contrast with the case with no externalities, however, we show that various pricing strategies are consistent with an SPE under linear externalities. In effect, the necessary and sufficient condition for a price vector to be consistent with a monopolization SPE is that the sum of markups and markdowns for any subset of buyers is less than or equal to the externality benefits they collectively enjoy with the rest of the network. With no markup or markdown to any buyer, MC pricing satisfies the conditions for an SPE under linear externalities, and so does a class of bipartition pricing, which entails price discrimination based on a binary partition of the buyer set: Markups are charged to the buyers in one subset and the extracted surplus is used to subsidize the buyers in the other subset. The size of a markup or markdown to each buyer is proportional to the number of his neighbors in the other subset. When the externalities are non-linear, on the other hand, we show that MC pricing is consistent with an SPE only when the buyer network is either a cycle or complete.1 This observation leads us to the study of SPE pricing strategies that are robust against slight perturbations in the externalities from linearity. Specifically, we define SPE pricing strategies under linear externalities to be robust if there exists a non-degenerate set of approximately linear externalities under which there exists an SPE pricing strategy that is “close” to the original pricing strategy. We show that in a large class of networks, bipartition pricing given some binary partition of buyers is in fact robust. One important class of networks that admits a natural interpretation of robust bipartition pricing identified above is the class of networks that corresponds to two- sided markets: In these markets, the two firms are platforms that compete in offering marketplace to agents on two sides, and agents on one side experience externalities only from the adoption decisions of agents on the other side. We show that there exists an SPE with cross-subsidization in which all agents on one side are charged markups while all agents on the other side are offered discounts. Furthermore, the size of the markup or discount to each agent is approximately proportional to the number of users on the other side of the market who are directly connected to them.2 As mentioned above, the key assumption of our model is the ability of the firms to perfectly price discriminate the buyers. In the application of the current model to a consumption good market, hence, we interpret each buyer as a collection of consumers in a particular segment of the market differentiated by geographical locations, gender, past activity records, etc., which is referred to as a side in Jullien (2011).3 According to this interpretation, pricing is uniform within each side, but discriminatory across different sides. On the other hand, it is possible to think of each buyer as a single agent in intermediate goods markets where the number of participants is typically not so large. One good example is provided by the international competition in the sales of infrastructures that has recently become a major form of international trade. In a market for high-speed rails, for example, there are typically a small number of firms capable of providing a system, a buyer is either a country or a region and hence is also limited in number, and the good has externalities since it is not simply a physical product but includes the operation and management of the system: Countries contemplating the adoption of a high-speed rail would be concerned with the rail system adopted by their neighbors if future connection between their systems is anticipated.4 It is well recognized that games with adoption externalities possess multiple equilibria. In our model, this corresponds to the potential multiplicity of Nash equilibria (NE) in the subgame played by the buyers after the posting of the prices by the firms. Our analysis supposes that each firm expects the most pessimistic scenario when contemplating deviations. Specifically, we consider two extreme NE as follows: The A-maximal NE is one in which the set of buyers who choose A is maximal in the sense that buyer i chooses A in it as long as he chooses A in some NE. The B-maximal equilibrium is one in which the set of buyers who choose B is maximal in the same sense. We suppose that the buyers coordinate on the B-maximal equilibrium when firm A deviates, and the A-maximal equilibrium when firm B deviates. Although we avoid the discussion of how such coordination is achieved, this assumption is shown to support the broadest spectrum of SPE by minimizing the profitability of deviations, and hence is useful for the identification of the set of SPE prices. The paper is organized as follows: After discussing the related literature in Section 2, we formulate a model of price competition in Section 3. Section 4 considers the subgame played by the buyers in stage 2, and Section 5 presents a preliminary analysis of the two-stage game involving the firms. An example illustrating the discussion is presented in Section 6. Section 7 provides a characterization of SPE pricing strategies under linear externalities, whereas Section 8 analyzes the feasibility of MC pricing under non-linear externalities. The robustness of an SPE under linear externalities is studied in Section 9. Section 10 concludes with a discussion. All the proofs are collected in the Appendix.","This paper contributes to two strands of literature. First, it contributes to the literature on network competition and two-sided markets through the introduction of local network externalities. Beginning with Katz and Shapiro (1985), most work on the topic supposes that the externalities are global in the sense that the adoption decision of any single buyer affects all other buyers equally.5 In the context of two-sided markets, this implies that the participation of any agent on one side of the market equally affects the utility of all participating agents on the other side of the market.6 In contrast, we suppose that the adoption decision of any buyer affects only his neighbors on the network. In two-sided markets, our formulation implies that the participation of any agent may have different effects on different agents on the other side of the market.7 Second, it presents a general analysis of price competition between suppliers of goods with local network externalities. Models of price competition under local network externalities include Banerji and Dutta (2009), Bloch and Quérou (2013), Blume et al. (2009), Chen et al. (2018), Fainmesser and Galeotti (2016b), and Jullien (2011). Blume et al. (2009) and Bloch and Quérou (2013) study price competition under local network externalities when market segmentation among the firms is exogenously given. Banerji and Dutta (2009) study price competition using a graph representation of local externalities when there is no price discrimination. Chen et al. (2018) formulate a model of differentiated duopoly when the degree of network externalities is small enough to guarantee a unique NE in the buyers' game where they choose continuous consumption levels. Fainmesser and Galeotti (2016b) analyze a model in which two firms with differentiated products compete in price when they observe the degree of influence some consumers have on other consumers in the form of network externalities. Most closely related to the present model is Jullien (2011), which considers a model of Stackelberg price competition between two platforms under a very general specification of local externalities.8 Although we focus on a more restricted class of local externalities than in Jullien (2011) in order to derive a more explicit conclusion, there is a significant overlap in the analysis.9 The multiplicity of equilibria is often a central concern in games with network externalities.10 Since the pioneering work of Dybvig and Spatt (1983), this concern has led the literature to focus on such issues as implementing efficient or revenue maximizing equilibria under complete and incomplete information, intertemporal patterns of adoption decisions, as well as the validity of introductory pricing.11 As mentioned in the Introduction, we abstract from the issue by supposing that whenever there is a deviating firm, the buyers coordinate on its least favorable NE. The literature makes different assumptions in this regard. For example, Ambrus and Argenziano (2009) assume that the agents' actions satisfy correlated rationalizability, which implies that they coordinate on the pareto-efficient alternative, and Jullien (2011, Assumption 2) assumes that a change in price offer by one firm to buyers outside its market segment does not affect the decisions of those inside it. One key idea used in the present paper is that of divide-and-conquer, which has been studied by Segal (2003), and Bernstein and Winter (2012) among others in contracting problems in which a single principal offers a contract to the set of agents whose participation decisions create externalities to other agents.12","The following corollary presents easy-to-verify sufficient conditions for the requirement in Proposition 9.2, and shows that a robust bipartition SPE exists in a large class of networks.","Our construction of an SPE of the price competition game assumes that a non-deviating firm becomes focal following any deviation by either firm. While this assumption supports the broadest spectrum of SPE and serves our purpose, it is not consistent with, for example, the assumption that the buyers choose the Pareto efficient alternative. However, even if the buyers coordinate on a less extreme NE, or make adoption decisions sequentially, the set of SPE prices will be a subset of the set of SPE prices identified in the present analysis. The essential feature of the market for goods with network externalities is the multiplicity of equilibria. In the present context, this corresponds to the multiplicity of NE in the buyers' subgame. Fundamental multiplicity of equilibria also exists in the pricing game between the firms when the externalities are linear. In this case, any pricing strategy is consistent with an SPE as long as the sum of markups and markdowns for any subset of buyers does not exceed the externality benefits they collectively enjoy with the rest of the network. We show that MC pricing is not consistent with a monopolization SPE unless the network is a cycle or complete under generic externalities. Unfortunately, however, positive identification of SPE pricing strategies is difficult under generic externalities: Equilibrium pricing subtly depends on the network configurations and externality specifications, and it appears difficult to obtain a general principle. We show that bipartition pricing is consistent with an SPE under linear network externalities, and identify conditions under which it is robust to slight perturbations in externalities from linearity. There are a number of interesting extensions of the present model. A firm contemplating competing in the sale of network goods should naturally be concerned about uncertainty associated with the large multiplicity of equilibria as discussed above, and may take action to alleviate the problem before entering such a market. For example, a firm may try to design its product so that it will have stronger externalities than the competing products, or a positive degree of compatibility with them.49 It would also be interesting to consider alternatives to the assumptions of perfect price discrimination, public observability of prices, and perfect knowledge of the firms about the network.50 For example, while standard, the assumption that all the prices are perfectly observed by all buyers is strong, and it is interesting to examine what happens when the firms can choose to post its price to each buyer privately."],["Some social surveys now collect physical measurements and markers derived from biological samples, in addition to self-reported health assessments. This information is expensive to collect; its value in medical epidemiology has been clearly established, but its potential contribution to social science research is less certain. We focused on disability, which results from biological processes but is defined in terms of its implications for social functioning and wellbeing. Using data from waves 2 and 3 of the UK Understanding Society panel survey as our baseline, we estimated predictive models for disability 2–4 years ahead, using a wide range of biomarkers in addition to self-assessed health (SAH) and other socio-economic covariates. We found a quantitatively and statistically significant predictive role for a large set of nurse-collected and blood-based biomarkers, over and above the strong predictive power of self-assessed health. We also applied a latent variable model accounting for the longitudinal nature of observed disability outcomes and measurement error in in SAH and biomarkers. Although SAH performed well as a summary measure, it has shortcomings as a leading indicator of disability, since we found it to be biased in the sense of over- or under-sensitivity to certain biological pathways. --------------------------------------------------------------------------------","An important recent development in research based on large-scale social surveys is the integration of physical health measurements and markers derived from biological samples, in addition to traditional self-reported health assessments. Biomarkers are objectively measured and evaluated as indicators of normal biological or pathogenic processes (Colburn et al., 2001), and they have potential advantages over self-assessments as early indicators of conditions that are below clinical diagnostic thresholds, or are pre- symptomatic and below individuals’ threshold of perception (Colburn et al., 2001). Cardiovascular, metabolic, inflammatory, neuroendocrine and other biomarkers have been shown to be predictors of mortality and morbidity when used alone or alongside self- reported health assessments (Idler and Benyamini, 1997; Gruenewald et al., 2006; Ridker, 2007; Jylhä, 2009; Doiron et al., 2015). They have also been used to reveal the socioeconomic gradient in health risks (Seeman et al., 2004; Lee et al., 2015; Carrieri and Jones, 2017). Despite their advantages, biomarkers impose significant additional costs of collection in the survey context and their potential contribution to economic and social research is not entirely clear. The wider social impacts of ill-health – on quality of life, personal and social functioning, and social costs of disease – depend critically on the duration and severity of disability prior to death, and there has been little research on the role of biomarkers in relation to disability. Disability is associated with loss of employment, early retirement and serious consequences for the families affected (Pudney et al., 2011; Jones, 2016; Christensen and Gupta, 2017) and typically implies long-lasting impairments that may prevent independent living and generate large social costs. This is particularly so in the UK where disability prevalence is well above the European Union average (Jones, 2016) and has been rising rapidly (from 11.9 to 13.3 million over 2013/14–2015/16 (DWP, 2017). There is evidence of an increasing birth-cohort trend in functional difficulties for older individuals of low socio-economic status (Morciano et al., 2015) and developed countries like the UK may face severe problems in supporting the projected future growth in the disabled population and providing public support to people with care needs (Commission on Funding of Care and Support, 2011). A crucial question for researchers and policymakers is whether the demand for care services will be curbed by gains in disability-free life expectancy alongside the projected continuing gains in longevity. An answer to this question requires a better understanding of the processes leading to disability, allowing the development of strategies and screening programmes to address disability more efficiently. The availability of biomarker information in population-representative data may contribute to that better understanding. Despite the importance of disability trends for social policy planning, relatively little is known about the association between biomarkers and future disability. The World Health Organization (WHO) proposed a framework that portrays progression from diseases to functional disabilities (WHO, 1980), and Fried et al. (1991) hypothesized the existence of pre-clinical disability as an intermediate stage in which health impairments have an impact on general functioning. Few studies have explored the predictive role of biomarkers in relation to this disability process, and most are limited by being based on small samples or unrepresentative data (Brex et al., 2002; Reuben et al., 1999; Baylis et al., 2013; Seeman et al., 1994; Kallaur et al., 2017); or focused exclusively on older individuals (Reuben et al., 1999; Seeman et al., 1994; Baylis et al., 2013); or concerned with disability outcomes from a specific disease or condition (Brex et al., 2002; Kallaur et al., 2017). Another study by Pagan et al. (2016) tests the hypothesis that disability is a potential mediator in the link between obesity and job satisfaction. We examined the predictive power of a wide range of biomarkers for future disability and specifically asked whether biomarkers offer incremental value in predicting disability outcomes beyond the contribution of the conventional self-assessed health (SAH) measure. SAH may be associated with disability outcomes in parallel with biomarkers by reflecting the impact on the individual of diagnosis of health conditions defined in terms of elevated biomarkers (Idler et al., 2004; Jylhä, 2009), or through bodily sensations that are sensitive to the biochemical processes measured by biomarkers (a, 2009). Besides confirming the value of biomarkers as leading indicators of disability, we also investigated the success of SAH as an overall summary of biomedical states relevant to disability and examined whether SAH is biased in the sense of over- or under-sensitivity to specific biological pathways. The paper makes several new contributions to the research literature. To the best of our knowledge, this is the first study that provides a comprehensive analysis of this kind. Using baseline data from early waves of the UK Household Longitudinal Study (UKHLS, also known as Understanding Society), we estimate predictive models for disability two to four years ahead, exploiting a large set of nurse- collected and blood-based biomarkers. These biomarkers measure adiposity, grip strength, blood pressure, lung, kidney and liver functions, inflammation, steroid hormone levels, blood sugar and anaemia, giving an unusually broad picture of individuals’ health states. The use of several alternative disability measures demonstrates the robustness of our results. In addition to simple prediction models, we also develop a latent variable (LV) approach which is new to the literature. This LV model has a number of advantages. First, unlike simpler prediction models, it takes into account the longitudinal nature of our data on disability, allowing for correlation between disability variables two to four years after baseline. Second, it addresses measurement error bias by allowing for measurement noise in both SAH (Crossley and Kennedy, 2002) and the biomarkers (Zang et al., 2015). Measurement error in biomarkers normally causes attenuation of the estimated impact of the biological pathways on disability, and the LV model is expected to give more accurate estimates of these impacts. This may be of particular importance for developing policy strategies and interventions to reduce the personal and social burden of disability. Third, our LV approach is set up to identify any distinct dimension of health which influences future disability and is captured by the biomarkers but missed by SAH. The pattern of factor loadings tells us what underlying aspects of health tend to be under- or over-represented by the SAH measure and can therefore guide the interpretation of research findings related to SAH.","The UKHLS is a large, nationally representative panel survey, running continuously from the initial wave in 2009–10, with each panel member interviewed annually. Its predecessor, the British Household Panel Survey (BHPS) was incorporated into the UKHLS from wave 2. A set of physical health measures and non-fasted blood samples were collected by nurses, five months on average after the wave 2 interview for UKHLS respondents and similarly at wave 3 for the BHPS sample (McFall et al., 2014). Respondents were eligible for nurse visits if, at the relevant wave, they were aged 16 or over, lived in England, Wales or Scotland and were not pregnant. Blood sample collections were further restricted to those who had no clotting or bleeding disorders and no history of fits. Participants gave informed written consent for their blood to be taken and stored for future scientific analysis. The UKHLS has been approved by the University of Essex Ethics Committee and the nurse data collection by the National Research Ethics Service (10/H0604/2). We define wave 2 as the baseline for the main UKHLS panel and wave 3 as the baseline for the BHPS sub- panel and we refer to the timing of this baseline observation as t = 0; these baseline observations were spread through calendar years 2010–2013. We used baseline data on personal and household characteristics and bio-medical measures as predictors of disability observed subsequently at waves 4–6 for the UKHLS sample (t = 2, … 4) or waves 5–6 for the BHPS sample (t = 2, 3), where t denotes years since the baseline main interview. We did not use data from t = 1 since the time gap between collection of biomarkers and interview at t = 1 was less than 6 months for 75% of the t = 1 sample. There were 15,632 and 5053 UKHLS and BHPS respondents who participated in the wave 2 or wave 3 nurse visits. For the UKHLS group, 13,404, 12,719 and 11,434 were followed up at waves 4, 5 and 6 respectively and had non-missing information on disability; 4513 and 4113 of the BHPS subsample were followed up at waves 5 and 6. We further conditioned the analysis on the absence of reported disability at baseline by excluding those reporting the relevant type of disability at baseline. Exclusion of these cases reduced the potential samples by 25%, 13% and 8% depending on the disability concept used. A detailed summary of available sample sizes is given in the Supplement Table S1. Biomarkers ~~~~~~~~~~ Measures of adiposity, grip strength, heart rate, blood pressure and lung function were collected during visits by trained nurses. We used the waist-to-height ratio (WHR) to measure adiposity. Grip strength was measured (in kg) using a hand dynamometer (McFall et al., 2014) and we took the highest reading from three repeated measurements for the dominant hand. Higher levels of grip strength are indicative of better physical functioning. Three repeated measurements of resting heart rate (HR), systolic and diastolic blood pressure (SBP, DBP) were taken at intervals of one minute (McFall et al., 2014). We skipped the first reading, believed to impose upward biases, and computed HR, SBP and DBP as the average of the second and third readings. HR, which is sometimes regarded as a measure of fitness rather than health, was used as a continuous variable, and we also used a binary hypertension indicator recording SBP > 140, DBP > 90 and/or current use of anti-hypertensive medications (Johnston et al., 2009). Lung function, assessed using spirometry equipment, was measured by the total amount of air forcibly blown out after a full inspiration (forced vital capacity; FVC), higher values indicating better lung function (Gray et al., 2013). Forced expiratory volume in one second (FEV1) is often used as an alternative to FVC. However, different equipment and measurement protocols were used in Scotland than in England and Wales, and comparison of matched samples showed FEV1 to be seriously affected by this, while FVC measures appeared comparable. Consequently we retained the Scottish sample and used FVC as our lung function measure. We used blood-based biomarkers specific to inflammation, steroid hormone, cholesterol, blood sugar, kidney function, liver function and anaemia. C-reactive protein (CRP) rises as part of the immune response to infection and is associated with general chronic or systemic inflammation. We excluded those with a CRP over 10 mg/L, because those values may reflect current transient infections rather than chronic processes (Pearson et al., 2003). Dihydroepiandrosterone suphate (DHEAS) is the most common steroid hormone in the body, considered as one of the primary mechanisms through which psychosocial stressors may affect individual health. Low levels of DHEAS are associated with cardiovascular (CVD) risks and all-cause mortality (Ohlsson et al., 2010). High-density lipoprotein cholesterol (HDL) is known as “good” cholesterol, low levels being associated with increased CVD risks (Wannamethee et al., 2000). Glycated haemoglobin (HbA1c) measures blood sugar, and is regarded as a validated diagnostic test for diabetes (WHO, 2011). The estimated glomerular filtration rate (EGFR), calculated from the serum creatinine concentration, measures kidney function; higher EGFR levels indicate better kidney function (Levey et al., 2009). We used albumin, the main protein made by the liver, as a liver function test; low albumin levels suggest impaired liver function (Howard and Sparks, 2016). Haemoglobin (Hgb), is an iron-containing protein responsible for carrying oxygen throughout the body, and was used to proxy anaemia status; lower levels of Hgb are suggestive of anaemia (Balarajan et al., 2012). In addition to specific markers, we also used two composite summary measures. One was an index of multi-system risk that measures the wear and tear on the body, approximating the allostatic load (Seeman et al., 2008). Our index combined the selected biomarkers for inflammation, blood pressure, HR, HbA1c, HDL cholesterol, albumin, DHEAS and the WHR (Seeman et al., 2008; Howard and Sparks, 2016). HDL, Albumin and DHEAS were converted to negative values to reflect ill health, and then each biomarker was transformed into z-scores and summed to calculate the overall index. The second index was a cumulative risk score for CVD, created by adding the relevant z-scores for WHR, blood pressure, HbA1c and CRP (Walsemann et al., 2016). A summary of all biomarkers by reported future disability state is given in the Supplement Tables S3 and S4. Self-assessed health ~~~~~~~~~~~~~~~~~~~~ SAH is considered a summary measure capturing the way that numerous aspects of health, both subjective and objective, are combined within the perceptual framework of the individual respondent (a, 2009). The SAH question asked respondents to rate their health on a five-point scale from “excellent” to “poor”. It was collected in the self-completion instrument at baseline, approximately five months prior to biomarker collection. We group the lowest two SAH categories because of their small sample size, giving a four-point scale ranging from 1 = “excellent” to 4 = “fair” or “poor”. Disability measures ~~~~~~~~~~~~~~~~~~~ Our measures of disability were collected at UKHLS waves 4–6, so disability outcomes were observed for prediction horizons of t = 2 …4 years for the main UKHLS sample or t = 2, 3 years for BHPS respondents. Respondents were asked about any long-standing physical or mental impairment that they might have and then the consequent functional difficulties, from a list of twelve provided (see Supplement Table S2 for the full list). We used the report of any functional difficulty as a dichotomous variable and the number of functional difficulties (coded as an ordinal variable: 0, 1, 2 or 3+) as an indicator of severity. Specific difficulty with mobility is also examined as a separate dichotomous indicator because of its relatively high prevalence and significance for functioning and independence in later life (Guralnik et al., 1993). Our fourth disability measure came from the income module of the UKHLS questionnaire, constructed as a binary indicator of whether the respondent received income from private disability insurance or any of the UK disability benefit programmes. For all programmes, receipt requires a decision to apply for the benefit, the ability to craft a good-quality application and a positive assessment of need by the programme administrators. Consequently, while the benefit receipt indicator involves a rigorous external assessment of severity, that assessment is confounded to some degree by the incentive and capacity to apply for the benefit, which has a strong socioeconomic gradient (Hancock et al., 2016). Covariates ~~~~~~~~~~ The explanatory covariates that we used in our models have been found to be associated with disability (Hernández-Quevedo et al., 2008; Morciano et al., 2015), and also directly with biomarkers (Carrieri and Jones, 2017). The covariates were collected at baseline and are described and summarised in the Supplement Table S5. Gender and polynomials in age were used to capture demographic influences. Three indicators of socio-economic status were included: educational attainment, home ownership and household income. We excluded disability benefits from income to avoid spurious correlation arising from the fact that disability creates eligibility for those benefits (see Morciano et al. (2015)). We also controlled for marital status, household composition, national and urban dummies. To assess the impact on future disability of lung function over and above smoking status, we estimated models with and without the inclusion of smoking variables.","Our analysis uses a sample of individuals who had no observed history of disability at baseline defined as the time of biomarker collection (t = 0). Conditioning the analysis in this way focuses attention on the transition from full physical functioning to disability, and it avoids the complications raised by the fact that, for people already disabled at baseline, we do not observe the evolution of their health and disability prior to joining the Understanding Society panel. In addition to the three binary indicators, we also extended model (1) to analyse the reported number of functional difficulties as a 4-level ordered probit model. Average marginal impacts of SAH and biomarkers on the probabilities of reporting 2+ or 3+ difficulties were then constructed from the estimated ordered probit. We did not use all 12 specific biomarkers simultaneously as predictors in model (1), for three reasons. First, there was a significant loss of usable data when all markers are required to be observed. Second, the full set of health measures (SAH and biomarkers) displayed a substantial degree of collinearity, so there would have been a further loss of statistical precision. A third policy-relevant reason for considering each biomarker separately is that, in practice, it is unlikely that any screening programme would simultaneously check blood pressure, adiposity, blood sugar, cholesterol, haemoglobin, hormone levels, liver, kidney and lung function; so there is a practical interest in the predictive power of each specific marker on its own. However, we also separately used the two composite indexes for allostatic load and CVD risk to consider the potential performance of more comprehensive tests. Although models like (1) are fairly standard in the literature, they are vulnerable to measurement error bias. Biomarkers and SAH are best seen as noisy indicators of the relevant health concepts rather than direct observations of those concepts. SAH may be subject to transient random variations in mood and perception, while biomarkers are affected by random variations in blood samples and measurement processes. Measurement noise may bias estimates of predictive models like (1), usually causing attenuation of the estimated impact of biomarkers and SAH. Our alternative latent variable approach offers a way of dealing with measurement error, by exploiting the multiplicity of biomarkers that reflect to varying degrees the underlying health state. Our aim here is to develop a form of the LV model that gives a clear indication of the predictive value of the biomarker information beyond that contained in the SAH measure. We used a two-factor structural LV model in which a latent variable hi0 reflects a dimension of general health at baseline, measured to varying degrees by SAH and the set of twelve biomarkers (not the indexes for allostatic load and CVD risk). To capture the incremental contribution of the biomarkers, we specified a second latent health variable, bi0 representing any dimension of baseline health that the biomarkers succeed in measuring, but which is not captured by SAH. The outcome variables, Di2, Di3, Di4, represent disability 2, 3 and 4 years after baseline; the correlation between them is captured by an unobserved random effect, ui, which may have different impacts in different periods. The resulting model has the structure shown in Fig. 1 and set out algebraically in the Supplement. Simple predictive models ~~~~~~~~~~~~~~~~~~~~~~~~ The results for model (1) at horizon t = 4 are presented in Table 1, which shows the percentage impact on the predicted number of people classified as disabled, of a 1-standard deviation increase in the relevant biomarker. Columns three (one or more functional difficulties), six (mobility) and seven (benefits) of the table were derived from binary probit models; each cell in those columns was based on a separate model for the relevant combination of biomarker and disability measure. The cells in each row of columns four and five came from the same ordered probit model for the number of reported disabilities, with the impact of each biomarker evaluated respectively at the 2+ and 3+ thresholds. First note that, when SAH was excluded from the prediction models, almost all biomarkers had substantial and statistically significant (at least at the 5% level) predictive power. The exceptions were for hypertension, DHEAS, EGFR and Hgb, where estimates were insignificant at the 5% level for some disability measures. When dummy variables for SAH status were introduced into the model, the magnitudes of the biomarker coefficients fell, mostly by 20–40%, but generally remained statistically significant. Consequently, the SAH measure succeeded in capturing some of the predictive power of the more objective measures, but far from all. There was some variation across biomarkers but, overall, the nurse-collected and blood-based biomarkers were comparable in terms of their predictive power. As expected, biomarkers for which higher values represent worse health states had a positive percentage impact on future disability and vice versa. Note that lung function, which emerged as a strong influence on disability, was not acting as a proxy for smoking, since inclusion of smoking variables left the estimated effect practically unchanged. This is of particular interest given recent evidence of the strong predictive value of smoking on future disability (Bengtsson and Nilsson, 2018). To save space, the results for shorter prediction horizons (t = 2, 3) are reported in the Supplement (Tables S6–S8); they show systematic predictive power for most of the biomarkers, with the effects mostly rising for longer prediction horizons. To test whether the predictive role of each biomarker and/or SAH vary by demographic and socioeconomic status (household income, education and house tenure), we tested the relevant interaction terms across our different model specifications. The tests found no systematic differences in the predictive role of biomarkers and SAH on future disability by age, gender or socioeconomic status. For example, P-values for the interactions of allostatic load with gender range between 0.471 and 0.716 across the 4-year prediction models of the different disability outcomes; for interactions with age, P-values ranged between 0.211 and 0.898. We also found no systematic interactions of our socioeconomic status variables with allostatic load (P-values between 0.170 and 0.905). Both composite biomarker measures gave strong effects, with allostatic load having the strongest impact. The results for the benefit receipt measure of disability were an interesting exception to this: when SAH was included in the model, the magnitude of the biomarker effect halved and retained significance only at the 10% level. There may be two behavioural factors involved in that result. One is justification bias – receipt of benefit may lead some respondents to report a worse state of health in SAH to justify their receipt of disability benefit. Alternatively, some people may be reluctant to accept or admit that their health is poor, leading them both to under-represent their current health difficulties in SAH and avoid claiming their potential entitlement to disability benefit. Both of these behaviours would be likely to strengthen the empirical SAH effect relative to the estimated effect of allostatic load or CVD risk indexes (which are more highly correlated with SAH than individual biomarkers, and therefore more affected by bias in SAH). Figs. 2 and 3 compare the magnitude of the SAH and biomarker impacts calculated from models where both SAH and the relevant biomarker were use as predictors (together with the covariates X). The SAH impact was calculated as the mean predicted impact of switching the individual from the best (“excellent”) to worst (“poor/very poor”) category of SAH; with the exception of the binary hypertension marker, the biomarker effect was calculated as the mean impact of switching from a value approximately 1.5 standard deviations better to 1.5 standard deviations worse than the mean (between the 5th and 95th percentiles). Impacts are calculated using the covariates X for each sampled individual and then averaged. Figs. 2 and 3 (see also Table S9 of the Supplement) show that the biomarkers with statistically significant impacts shown in Table 1 made predictive contributions of about 20–25% of the SAH effect in most cases. For disability measures based on the number of disabilities reported and mobility difficulties, allostatic load made the largest contribution to prediction, both absolutely and as a proportion of the SAH impact: the impact of a three standard deviation change in allostatic load was approximately one third that of the hypothetical SAH shift. For the mobility-based disability criterion, allostatic load and WHR had the largest impacts in absolute terms (roughly 40% of the SAH effect). In accordance with the results in Table 1, allostatic load contributed less to the prediction of benefit receipt, while the markers for grip strength, adiposity and lung function all gave substantially greater impacts (around one third of the SAH effect). Despite some differences between disability definitions in the pattern of results, the overall conclusion seems robust – biomarkers made a contribution to prediction of disability four years ahead that is significant both statistically and in terms of absolute magnitude. But that contribution is moderate in comparison with the information contained in SAH. There is a risk that the act of observation could change the behaviour being observed – that formal or informal feedback about the respondents’ biomarker levels may prompt additional GP consultations and treatments for previously undiagnosed health conditions, or behavioural changes (Zhao et al., 2013). In the UKHLS, survey nurses were instructed not to discuss or interpret respondents’ results in general or in relation to other people in the survey, and blood tests results were not available to survey participants. However, participants received a Measurement Record Card with their blood pressures, height, weight, waist circumference, percent body fat, and grip strength. The survey protocol (National Centre for Social Research, 2010; McFall et al., 2014) also specified tailored blood pressure feedback: respondents were informed and advised to visit a GP within 2 months, 2 weeks or 5 days if their blood pressure was mildly raised (140≤ systolic blood pressure <160 or 90≤ diastolic blood pressure <100), moderately raised (160–180 or 100–115) or considerably raised (over 180 or over 115), respectively. Those with normal blood pressure measurements received reassuring feedback. To explore the robustness of our results to the possibility of feedback effects, we carried out a number of sensitivity analyses. First, we added the categorial blood pressure feedback variable to all our predictive models, and found that the results remained practically identical to those presented in Table 1 and Figs. 2 and 3. We also interacted the categorical blood pressure feedback variable with our biomarkers, finding the interaction terms statistically insignificant (at the 10% level) across all models. For example, the P-values of interaction terms between our composite biomarker measure (allostatic load) and the categorical blood pressure feedback variable range from 0.678 and 0.904 across the different disability outcome models. Thus the predictive role of biomarkers does not systemically differ between those who were informed of an elevated blood pressure and those with normal blood pressure levels. We also re-estimated the prediction models after excluding respondents who received feedback of mildly, moderately or considerably raised blood pressure, showing only minor differences to our base case results (compare Tables S9 and S10 of the Supplement). Overall, this evidence alleviates concerns about potential distortions from survey feedback effects – although it is disappointing from the policy perspective that there is so little evidence of feedback generating health improvements. This bleak result is consistent with recent evidence suggesting that screening programmes providing health information are often relatively ineffective as a means of disease control (Chang et al., 2018; Kim et al., 2019). Do SAH and biomarkers measure the same thing? Latent variable results ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our finding that SAH is a strong predictor of future disability parallels similar evidence on mortality risks (Glei et al., 2014; Idler and Benyamini, 1997; Jylhä, 2009), but it raises the issue of how SAH relates to more objective biological aspects of health. SAH is a subjective response that may reflect bodily sensations produced by biological disease processes, but also potentially many other things such as mood, self-image, past contacts with the healthcare system, and health shocks to friends and relatives (Idler and Benyamini, 1997; Jylhä, 2009). If used for policy purposes, SAH might also be reported subconsciously or strategically to achieve or justify a particular outcome. Even in its relation to biological processes, SAH may be biased in particular directions since not all disease processes are equally apparent to the sufferer. No single biomarker or composite biomarker index can plausibly act as a direct comparator in a conventional measurement error validation study (Bound et al., 2001). A more productive approach is to use biomarker information alongside SAH to indicate the biological pathways to which SAH is insufficiently or excessively sensitive.1 We use the LV approach outlined in Section 3 to integrate SAH and biomarkers within an appropriately broad measurement setting. In practice biomarkers are also potentially error-prone measures, although not vulnerable to the behavioural reporting biases that may affect SAH. Random measurement errors in biomarkers used as predictors cause bias for the purposes of estimating causal links between health and future disability. But that bias is not a problem if we are concerned with prediction in the context of policy applications such as screening or monitoring programmes, since the biomarkers available to programme administrators are subject to the same measurement error process – the (causally biased) prediction model still gives the best prediction of future disability conditional on biomarker information as observed by the administrator. However, to understand the relationship between SAH and biological disease pathways we need estimates more closely related to the true causal processes. An advantage of the LV model is that allows for measurement noise in both SAH and biomarkers and so avoids measurement error bias stemming from classical random measurement error. Tables 2 and 3 summarise results from the LV model. Table 2 gives the estimated factor loadings relating observable biomarkers to the two latent dimensions of health. Except for the binary hypertension measure, the loadings were normalised to give the impacts of latent health on observed indicators in standard deviation units. The loadings of the primary latent health factor hi0 are mostly as expected, with disability risk raised significantly by WHR, hypertension, CRP and HbA1c; and lowered by grip strength, lung capacity, HDL cholesterol, DHEAS, EGFR and Albumin. The loadings on the second latent factor b must be interpreted carefully in relation to the SAH loadings. For any given marker, if the loading on h has the correct sign and the loading on b has the same sign, then this implies that the health concept implicitly captured by SAH understates the role of the biological pathway which the marker measures. Under this interpretation, our finding is that SAH strongly understates the importance of grip strength, lung function, DHEAS and liver function and weakly understates the effect of CRP. If the loading on h has the correct sign and the loading on b has the opposite sign, then SAH can be interpreted as over-emphasising the pathway measured by the marker. This was the case for WHR, hypertension and HDL cholesterol and, more weakly, for HbA1c and EGFR. In two cases, HR and Hgb, the loading on h was small, with an a priori wrong sign, which was strongly reversed in the loading for b. This finding can be interpreted to mean that, in these dimensions, SAH is a potentially misleading indicator of future disability. Our finding from the simple predictive model (1) was that biomarkers provide significant and substantial predictive power which is nevertheless moderate in relation to SAH. The same result is also evident in results from the LV model. Table 3 gives summary statistics for the predictive probabilities of future disability conditional only on SAH and covariates (π1(X, S)), and conditional on SAH, covariates and twelve biomarkers (π2(X, S, B1 … B12)). These predictive probabilities were calculated for all disability definitions and prediction horizons. For three of the predictors, the mean predicted probability of disability rose as we varied the prediction horizon from t = 2 to t = 4, by approximately 20% (1 or more functional difficulties), 28% (mobility problem) or 62% (benefit receipt). This rise in disability risk over time is a natural reflection of the cumulative character of disability prevalence. However, there was no such rise in prevalence measured as the proportion of individuals reporting two or more functional difficulties. For all four disability measures, the predictive probabilities also became substantially more variable (their standard deviations increased) as the prediction horizon lengthened. For (almost) all of the predictions, there was a rise in the variability of the predictor when we expanded the baseline predictor set by adding the biomarkers. This reflects the fact that adding biomarkers gives a more detailed and diverse picture of each individual's disability risk. The correlations between π1 and π2 were generally high, reflecting the good – but not perfect – performance of SAH as a general health proxy.","We investigated the predictive power of objective nurse-collected biomeasures and blood- based biomarkers for functional disability, following individuals who reported no disability at baseline for up to four years after collection of the health measures. We used a wide range of biomarkers, and alternative measures of disability, covering the existence and number of functional disabilities, mobility difficulties and receipt of disability benefits. For almost all of the biomarkers, we found 4-year-ahead predictive effects that were substantial in magnitude and statistically significant. When SAH was introduced as a predictor alongside the biomarkers, the magnitude of the biomarker effects fell, in most cases by 20–40%, but remained important in magnitude and highly statistically significant for most of the biomarkers examined. Although there were some differences across biomarkers and disability measures, we found that measures of adiposity, grip strength, heart rate, lung functioning, cholesterol levels, inflammation, blood sugar and anaemia had strong predictive power for future disability risk, over and above SAH. How should we interpret these predictive results? Causality and predictability are not the same thing, and there is a possibility that the association between health measures and future disability outcomes is partly the result of unobserved factors. In our view, causality can never be established unambiguously in observational data but, compared to cross-section analysis, the 4-year gap between our initial health measure and disability outcomes makes it much more likely that the predictive effects are causal rather than proxy in nature. This separation in time removes the possibility of reverse causation and weakens any ability of baseline biomarker levels to proxy omitted heterogeneity. Moreover, the most likely non-causal channels of association appear not to be important here – we found no evidence of an effect of health information fed back to respondents, nor of smoking behaviour on disability outcomes. It is also important not to over-emphasise the importance of causality. Predictability is extremely important for some major policy purposes – for example, a successful screening programme needs a good predictor to identify priority population groups, and that does not necessarily require a complete structural understanding of all the biochemical, behavioural and environmental processes involved in determining the target outcomes. In addition to simple predictive models, we also developed a new latent variable approach capable of incorporating large numbers of biomarkers and longitudinal observation of disability outcomes, while allowing for measurement error in SAH and biomarkers. This approach allowed us to identify distortions in SAH as a measure of health, by detecting an additional predictive factor that is revealed by the biomarkers but not by SAH. The corresponding factor loadings indicate dimensions of biological function that are given too much or too little weight by SAH in predicting disability. We found that SAH is excessively sensitive to the biological pathways reflected in adiposity, hypertension and cholesterol, and insufficiently sensitive to strength, lung function, hormonal balance and liver function. Nevertheless, SAH emerged as a good general health proxy in the sense that, when SAH and biomarkers were both used as predictors, the estimated biomarker impacts on future disability, although substantial absolutely, were moderate in comparison with the effects of SAH. For example, using a composite summary measure to proxy allostatic load (our strongest biomarker predictor), we found that moving from the best (excellent) to worst (fair/poor) category of SAH increased the risk of disability 4 years later by 5–18 percentage points on average, depending on the disability concept used, while an increase in allostatic load from the fifth to the lowest 95th percentile (roughly a 3-standard deviation rise) increased disability risk by 2–7 points. Limitations ~~~~~~~~~~~ Key strengths of our analysis come from its use of UKHLS data which allowed us to use a large, nationally representative sample covering all adult ages. The bio-social character of UKHLS provided a wide range of nurse-collected and blood-based biomarkers, in addition to SAH, disability indicators and extensive measures of household characteristics and socio-economic status. This adds breadth and depth to the small body of evidence that already exists on biomarkers as predictors of future disability. Existing studies are more limited in terms of the range of biomarkers used and also the study population, which is mainly restricted to older people, nonrepresentative samples or specific patient groups (Brex et al., 2002; Reuben et al., 1999; Baylis et al., 2013; Seeman et al., 1994). As far as we are aware, ours is the first study that makes an explicit evaluation of subjective SAH against objective biomarker information in relation to disability. Despite these advantages, there are limitations. First, the available data follow individuals for a relatively short time horizon. We have found evidence of a rise in the estimated effect of biomarkers as the length of the prediction horizon increases, suggesting that our results may understate the full long-term predictive role of biomarkers. Second, functional disability is a slippery, hard-to-measure concept and the measures used in our analysis are necessarily limited. Our use of a range of alternative disability measures alleviates these concerns to some degree, but a complete solution requires further research. Finally, although we used an unusually extensive set of biomarkers, the multidimensional nature of the biomedical processes underlying disability means that there may remain significant aspects of physical health that are not covered by our analysis."],["It is an ongoing debate how to increase the adoption of energy-efficient light bulbs and household appliances in the presence of the so-called ‘energy efficiency gap’. One measure to support consumers’ decision-making towards the purchase of more efficient appliances is the display of energy-related information in the form of energy-efficiency labels on electric consumer products. Another measure is to educate consumers in order to increase their level of energy and investment literacy. Thus, two questions arise when it comes to the display of energy-related information on appliances: (1) What kind of information should be displayed to enable consumers to make rational and efficient choices? (2) What abilities and prior knowledge do consumers need to possess to be able to process this information? In this paper, using a series of (recursive) bivariate probit models and three samples of 583, 877 and 1375 households from three major Swiss urban areas, we show how displaying information on the future energy consumption of electrical appliances in monetary terms (CHF), rather than in physical units (kWh), increases the probability that an individual makes a calculation and identifies the appliance with the lowest lifetime cost. In addition, our econometric results suggest that individuals with a higher level of energy and, in particular, investment literacy are more likely to perform an optimization rather than relying on a decision-making heuristic. These individuals are also more likely to identify the most (cost-)efficient appliance. --------------------------------------------------------------------------------","The residential sector consumes nearly 30% of the total final energy consumption in Switzerland and about 58% of the energy end-use consumption of households is based on fossil fuels (BFE, 2015). Improving the energy efficiency in the residential sector is therefore one of the strategies to reduce total fossil energy consumption and related CO2-emissions in Switzerland. While a major effort needs to be made to reduce the consumption of heating fuels, there is also a potential for enhanced energy efficiency in the electricity consumption of Swiss households. One important pillar of reducing residential electricity consumption is to foster the adoption of energy-efficient lighting and household appliances. A low adoption of energy-efficient technologies is often related to the ‘energy-efficiency gap’ (Sanstad and Howarth, 1994a; Howarth and Sanstad, 1995; Allcott and Greenstone, 2012), i.e. the frequent observation that individual decision- makers do not choose the most energy-efficient appliance, even if this appliance is also the most cost-efficient choice from the individual's point of view (minimizing lifetime operating costs).1 The list of potential underlying causes for the ‘energy-efficiency gap’ is long and includes a myriad of market and behavioral failures (Sanstad and Howarth, 1994b; Broberg and Kazukauskas, 2015). A large body of literature studies, for example, (implicit) subjective discount rates and their role for the persistence of the energy efficiency gap (Hausman, 1979; Train, 1985; Coller and Williams, 1999; Harrison et al., 2002; Epper et al., 2011; Bruderer Enzler et al., 2014; Min et al., 2014). In this paper, we abstract from subjective discounting and other market and behavioral failures to focus on those market and behavioral failures that are related to the provision and processing of energy-related information. In order to choose between two similar electrical appliances, a rational, utility-maximizing consumer should solve an optimization problem in order to choose the appliance that minimizes the sum of the purchase price and the present value of future energy costs (Sanstad and Howarth, 1994a,b; Gerarden et al., 2015). This optimization requires knowledge on the purchase prices of the appliances to choose from, the electricity consumption of the respective products, the expected intensity and/or frequency of use, the expected lifetime of the appliance as well as current and future electricity prices. If markets provide too little or inadequate information about these parameters, or if this information is not salient enough to the consumer, this constitutes a barrier to solving the optimization problem (Sanstad and Howarth, 1994a). In fact, in many purchase situations, the information about the energy- efficiency of an appliance and thus about the future energy costs is less salient than the purchase price. However, even if information on the energy consumption of the appliance is provided, the optimization regarding the lifetime cost of an appliance depends on additional information as mentioned above. To carry out the optimization, the consumer needs to gather the required information and then process this information correctly. This creates both ‘information cost’ and ‘optimization cost’ (Conlisk, 1988), given that the consumer needs to deliberate upon the options to choose from. Acknowledging the presence of ‘deliberation cost’ (Pingle, 2015) is equivalent to acknowledging that individuals are ‘boundedly rational’ (Simon, 1959; Sanstad and Howarth, 1994a), which means that they are not always able to acquire and process all the necessary information to trade-off all the alternatives in real decision-making situations. This is because information acquisition is costly and the processing of information is cognitively burdensome. As a consequence of being boundedly rational individuals tend to have problems solving the optimization problem when making an investment decision. Instead, they often rely on simple rules of thumb or decision heuristics (Wilson and Dowlatabadi, 2007; Frederiks et al., 2015), which potentially widens the energy efficiency gap.2 Against the background of the described information-related market and behavioral failures, the research presented in this paper deals with the question how information on future energy consumption should be displayed on products in order to enable consumers to identify the appliance that minimizes lifetime cost. Furthermore, we investigate whether and to what extent cognitive abilities as well as energy and financial literacy support consumers in identifying a cost-efficient appliance. We hereby assume that consumers may follow two different types of decision- making strategies: One is to optimize over the lifetime cost of the appliance. This is in line with the neoclassical concept of a fully rational and informed consumer. The other type of decision-making strategy, which is in line with the concept of bounded rationality, is heuristic decision-making, i.e. choosing an appliance according to a specific and salient characteristic of the appliance, e.g., a low purchase price, a high energy-efficiency rating or a lower physical energy consumption. The choice of the decision-making strategy is determined, on the one hand, by the information that is readily available in the purchase situation, e.g., in the form of information display on the products. On the other hand, it is determined by an individual's ability to process the available information, which is influenced by socioeconomic factors as well as the individual's level of energy and investment literacy. The latter determine the individual- specific deliberation cost. We thus assume, that the deliberation cost are a function of energy and investment literacy, with energy literacy being defined as the individual's prior energy-related knowledge, such as knowledge about energy prices and energy consumption of different appliances, and investment literacy being defined as the individual's cognitive ability to perform an investment analysis. To examine the role of information display, energy and investment literacy on the choice of electrical appliances of boundedly rational consumers, we organized a household survey and conducted two online randomized controlled trials with simple decision tasks. In two experiments, respondents were presented two similar appliances and had to determine which of the two minimizes lifetime cost. With these experiments we thus did not elicit consumers preferences but their ability to calculate the lifetime cost of the appliances.3 Individuals were randomly assigned to a treatment in which yearly energy consumption of the appliances was displayed in either monetary terms or physical units. In the empirical part of the paper, we analyze the individual decisions in the experiments while accounting for the respondents’ energy and investment literacy, their attitudes towards energy conservation, as well as their sociodemographics. Drawing on three samples of 583, 877 and 1375 households from three major Swiss urban areas the data is analyzed in a series of (recursive) bivariate probit models. We find that displaying information on energy consumption in monetary terms rather than in physical units positively influences the probability to perform an investment analysis which in turn increases the probability to choose the most (cost-)efficient appliance. Also a higher level of energy and investment literacy clearly enhances the individuals’ probability to do an investment calculation and to choose the most efficient appliance. This supports the view that both displaying monetary information on future energy consumption as well as consumers’ prior knowledge and cognitive abilities are decisive factors when attempting to reduce the energy-efficiency gap. Investing in consumer education to increase their energy and investment literacy could thus be an important element in a set of policy measures to enhance residential energy efficiency. The remainder of the paper is organized as follows. Section 2 presents a literature review and discusses the theoretical considerations and hypotheses. The dataset and the experimental design is presented in Section 3 and the econometric specifications are presented in Section 4. Section 5 presents the results and Section 6 concludes. Information, rationality and the choice of efficient appliances ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Are individuals able to make fully rational decisions by minimizing total lifetime cost when choosing electric appliances? Or are they ‘boundedly rational’ (Simon, 1959) and therefore lack the cognitive abilities to perform the optimization task? The discussion about rational decision-making in the domain of energy efficiency and the role of information in the choice of efficient appliances is going on for a while (Sanstad and Howarth, 1994a,b). McNeill and Wilkie (1979) and Anderson and Claxton (1982) investigate how the provision of information about energy consumption in either monetary or physical units impacts on the choice of household appliances. While both studies do not find significant effects of information display, more recent studies do so. For example, Heinzle (2012) examines the impact that different ways of disclosing energy-related information have on the choice of TVs in online experiments in Germany. The effects of three different disclosure formats of energy labels on appliance choice are compared: annual energy consumption in terms of physical units (kWh), annual energy operating cost in monetary terms and lifetime energy operating cost in monetary terms. Heinzle (2012) shows that individuals tend to overestimate potential cost savings between two TV sets if provided with information on energy consumption in physical units. When disclosing energy consumption in monetary terms, respondents’ willingness to pay (WTP) for energy efficient TV sets only increased when lifetime energy cost but not when annual operating cost were displayed. The display of annual operating cost in monetary terms even reduced WTP compared to the display in physical units. Similar results are reported in Deutsch (2010): based on a randomized field experiment on a commercially operating price comparison website it is shown that disclosing the life-cycle cost of an appliance instead of the purchase price induces consumers to purchase cooling appliances that are on average 2.5% more efficient than in the absence of life-cycle cost disclosure. Also Newell and Siikamäki (2014) test the effects of different forms of energy efficiency labeling. Among other features, they evaluate the impact of a label including the estimated yearly operating cost versus the impact of a label including physical information on energy use. Based on a choice experiment among 1214 US home owners, they find that providing information on estimated yearly operating cost is more effective in enhancing willingness to pay for more efficient water heaters than providing information about energy use in physical units. These findings are partly in line with the findings in Heinzle (2012), except that in the study of Newell and Siikamäki (2014) the display of annual operating cost (as opposed to lifetime operating cost) had a positive influence on the choice of more efficient appliances. For light bulbs, Min et al. (2014) test the influence of energy labeling on implicit discount rates in an incentivized choice experiment among 168 US residents and also conclude that the provision of information on annual operating costs of the bulbs increases consumers’ WTP for more efficient bulbs. Allcott and Taubinsky (2015) tested the effect of simultaneous information about yearly and lifetime energy cost on consumers’ choices of either compact fluorescent (CFL) or incandescent light bulbs. In both an online and an in-store experiment, they provide information about yearly and lifetime costs of the bulbs to a treatment group. While the treatment increases the market share of CFLs by about 12% in the online experiment, a similar treatment in the in-store experiment seemed less effective. Although the effects of displaying yearly and lifetime energy cost cannot be separated in their experiments, their results further support the view that providing monetary information supports consumers in accounting for operating cost. Other studies consider the role of energy-efficiency rating scales on the choice of appliances. The results of these studies suggest that the information on energy use provided on energy labels tends to be disregarded in the presence of a rating scale. Both the results presented in Waechter et al. (2015) and Hille et al. (2015) indicate that the display of an energy-efficiency rating scale on electric appliances may divert attention from the information on actual energy consumption of the products, suggesting that the EU energy label in its current form is used as a decision-making heuristic by many consumers, rather than supporting a rational and informed decision making.4 Houde (2014) reaches a similar conclusion. He finds evidence that an energy label can act as a substitute for more precise information on energy consumption of an appliance and that consumers that rely on the energy label often overestimate the energy savings associated with the certified product. The role of energy literacy and education for appliance choice ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Another potentially important prerequisite for rational decision-making in the domain of electric appliances is energy literacy, which can be defined as an individual's ability to make informed and deliberate choices in the domain of household energy consumption. In the literature, energy literacy is defined as a comprehensive concept that has a cognitive, an affective and a behavioral dimension (DeWaters and Powers, 2011). According to DeWaters and Powers (2011), energy literacy thus comprises of (1) knowledge about energy production and consumption as well as its impact on the environment and society, (2) attitudes and values towards energy conservation as well as (3) corresponding behavior. In this paper, we use a narrower definition of energy literacy that mainly reflects the individual's knowledge about energy prices and the energy consumption of different household appliances. This is because our main goal is to examine what knowledge and cognitive abilities consumers need to have in order to identify cost-efficient appliances. With respect to investments in energy-efficient appliances, an additional component gets relevant: investment literacy, i.e. the individuals’ ability to perform an investment analysis and hence to correctly evaluate different investment alternatives, for example when choosing appliances or when deciding about energy-saving renovations. Regarding this ability, the literature on financial literacy is informative. Lusardi and Mitchell (2009) show, for example, that more educated people are more likely to correctly answer a question on compound interest and that this indicator of financial literacy has a relevant influence on economic decisions in several domains: inter alia, individuals who know about interest compounding are 15 percentage points more likely to be retirement planners (Lusardi and Mitchell, 2007). In a study on financial literacy in Switzerland, Brown and Graf (2013) find that respondents scoring high on financial literacy are more likely to have an investment related custody account and to make voluntary retirement savings. One of the first studies that investigates the effect of energy and investment literacy on electricity consumption in a large sample is the study of Brounen et al. (2013). They examine the effect of energy and investment literacy on household conservation behavior and energy consumption in an online survey in the Netherlands. Their indicators for energy and investment literacy are three items capturing the households’ awareness of the amount of their monthly gas/electricity bill, the respondents’ choice in a decision between two alternative heating systems with different levels of energy efficiency, and the answer to the question whether the household uses green electricity. They find that older and male respondents are more likely to know about their gas bill and that more educated respondents are more likely to make a rational investment decision in the heating system example. However, Brounen et al. (2013) do not find that energy literacy has an impact on energy conservation behavior among the sampled households in terms of thermostat settings, and also not on the overall electricity and gas consumption of the household. In addition, there seems to be a role for education more generally when it comes to energy-related decision-making. Some studies find a positive correlation between an individual's level of education and energy literacy or energy-related investment literacy. For example, Mills and Schleich (2010) find that education, among other socio-economic characteristics, is positively correlated with knowledge about energy efficiency labels on appliances. In another study, Mills and Schleich (2012) build an energy-related knowledge index and find that the knowledge index increases when the most educated member of the household has an university degree, whereas a high-school degree does not have any effect and vocational training has a negative effect on the knowledge index. Also Nair et al. (2010) report from a Swedish sample that a higher level of education as well as a better knowledge about energy efficiency measures in buildings increases the likelihood that a household invests in building envelope measures. Apart from the study of Brounen et al. (2013), there is only little research about the role of energy literacy and energy-related investment literacy on investment decisions in the domain of residential electricity consumption. In particular, we are not aware of any study that investigates the impact of energy and investment literacy on the choice of household appliances. This paper therefore contributes to the existing literature along three dimensions: first, further empirical evidence is provided on the role of displaying monetary information about the energy consumption of an appliance for the choice of electrical appliances. Second, according to our best knowledge, this is one of the first studies which analyzes the impact of the level of energy literacy on the choice of electrical appliances and the impact of energy literacy, investment literacy and monetary information on the choice of the decision making strategy (optimization vs. heuristic decision making). Lastly, we also analyze the impact of the decision making strategy on the choice of the appliance in a recursive model. Theoretical considerations and hypotheses ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the following, we provide some theoretical reasoning that can explain the role of energy information display, energy literacy and financial literacy on the choice of appliances. According to household production theory (Deaton and Muellbauer, 1980) households purchase inputs such as energy and capital (i.e. household appliances) and combine them to produce outputs which are the desired energy services such as cooked food, washed clothes or hot water. These energy services appear as arguments in the household's utility function (Muth, 1966; Flaig, 1990). The utility function of a household is based on the consumption of both energy services and all other goods and is maximized under the budget constraint. A high level of expenditure for energy services will reduce the options to consume all other goods. Therefore, when facing the choice of light bulbs or household appliances, consumers are confronted with the optimization problem of choosing the appliance that provides the desired energy service at the minimum lifetime cost. Consumers wish to reduce their overall expenditure for energy services and to maximize their opportunities to consume all other goods, otherwise they experience a loss in utility.5 For the minimization of lifetime cost consumers have to consider the purchase price and future operating costs of the appliance, which depend on the energy consumption of the appliance (in Watts), the lifetime of the appliance, the frequency and intensity of use of the appliance as well as on current and future electricity prices.6 Lifetime cost and intensity of use cannot be predicted with certainty in the moment an appliance is purchased, so the individual needs to form expectations regarding lifetime cost and intensity of use for each appliance within the set of appliances to choose from. The process of forming expectations and comparing the expected lifetime cost of several appliances requires time and other resources that can be considered as ‘decision cost’ (Pingle, 2015). To study the role of ‘decision cost’ for the choice of the decision-making strategy and for the choice of an appliance, we provide a simple 2-period-model of expectation formation that explicitly takes decision costs into account. This model is based on the model in Conlisk (1988) and assumes that an individual assesses the expected lifetime cost of an appliance before purchase.7 Similar as in Conlisk (1988), we assume that the individual faces two potential sources of a loss that he or she wants to minimize: (i) experiencing a loss in inter-temporal utility by either underestimating the lifetime cost of an appliance in period 1 and thus not allocating enough of the budget in period 2 on the consumption of the energy service, or, by overestimating the lifetime cost of an appliance in period 1 and thereby restricting consumption of other goods in period 1 itself due to the individual's budget constraint, and (ii) spending too much time and resources on decision making itself. The lifetime cost of the appliance that was estimated after having invested the optimal amount of units T★ for decision making can thus be expressed as L(T★) and can be considered as a realization of the random expectation variable E(L(T)). Next, we assume that we can define R as the rational expectation of L and that individual j's estimate E(L(T)) of R is a weighted combination of two elements: a free estimator f, that is based on a simple rule of thumb and is thus generated without any decision cost, and a costly improvement of f denoted as r(T), that depends on the time spent on deliberation. We further assume that the individual makes his or her decision as if the costly improvement r(T) was as accurate as a sample mean of T independent observations taken from a distribution with mean R and variance σ2 and as if f was as accurate as a sample mean of S independent observations which represents a lower number of draws taken from the same distribution with mean R and variance σ2. The intuition here is that for an individual who performs a detailed analysis, it would imply drawing a large number of T thoughts from the distribution with mean R and variance σ2 thereby increasing the probability that he or she arrives close to R which is the rational expectation of L. However, in a heuristic approach (rule of thumb), the judgment can be considered to be based upon only a small number of S thoughts, which can a priori be assumed to be objectively biased (note that from the point of view of the individual, the guess itself may be considered unbiased; otherwise he or she would make a different guess altogether). It is obvious that the lower bound for the optimal time spent on deliberation T★ is at least 0. This equation gives us some insight into the optimality of different decision- making strategies. Consider the case when an individual has a lower capacity to perform a calculation task (i.e. higher β) and the analysis is costly relative to the size of the problem (large C/σ2) to the extent that the initial guess is reliable (large S). If these conditions hold with enough force, the corner solution T★ = 0 applies and the individual spends no resources on deliberation but only follows the rule of thumb f which in this case is good enough. By definition, α has an upper bound of 1 when the corner solution T★ = 0 applies, i.e. when individual's rule of thumb choice is good enough. Eq. (6) can be seen to cover both the extremes; when T★ goes to infinity, the expected lifetime cost converges to the rational expectation R. On the other hand, when T★ = 0, i.e. when α = 1, the expected lifetime cost is the free estimator f. Depending upon the different parameters, an individual could be lying anywhere between the range of unboundedly rationality and a rule of thumb approximation (Conlisk, 1988). The estimation of the lifetime cost of the appliance will thus be more closer to the rational expectation R as α gets smaller. From Eq. (7), α is the lower the lower β, i.e. the lower the individual's effort needed to perform the estimation of lifetime cost of the appliance the lower C relative to σ2, i.e. the lower the decision making cost related to the complexity of the problem the lower S, i.e. the lower the amount of (costless) best guesses spent on estimating R, i.e. the less reliable the rule of thumb Any individual having to decide between several appliances on offer will first assess the lifetime cost of all the appliances separately and then, in a second step, compare them to identify the one with the minimum lifetime cost. From the above-described model, we derive two hypotheses with respect to whether the individual deliberates or follows a rule of thumb when comparing the appliances (choice of decision-making strategy). As a natural consequence, but not directly related to the above theoretical model, we can identify two more hypotheses in relation to what determines whether the individual successfully identifies the most cost- efficient appliance (choice of cost-efficient appliance). First, we expect that individuals with a higher level of energy and investment literacy are more likely to choose an investment analysis as the decision-making strategy as the process of deliberation is less costly to them, i.e. a lower β in the theoretical model (H1a). We further assume that disclosing yearly energy consumption in monetary terms (CHF) rather than in physical units (kWh) decreases the value of C and hence increases the probability that the individual carries out an investment analysis rather than following a heuristic decision-making strategy as this lowers the per unit cost of deliberation (H1b). In a second step when an individual compares the lifetime costs of two appliances, we expect that an individual that carries out an investment analysis is more likely to identify the more cost-efficient appliance (H2a). In addition, we expect that disclosing yearly energy consumption in monetary terms increases the probability that the individual identifies the more cost-efficient appliance, again, as the display of monetary information lowers the per unit cost of deliberation (H2b). Our hypotheses can be summarized as follows: H1a: The level of energy and investment literacy has a positive impact on the individuals’ ability to follow an optimization strategy rather than a heuristic strategy. H1b: Displaying information on the yearly energy consumption of an appliance in monetary terms rather than in physical units has a positive impact on the individuals’ ability to follow an optimization strategy rather than a heuristic strategy. H2a: Opting for optimization as the decision-making strategy has a positive impact on the probability to identify the most cost-efficient appliance. H2b: Displaying information on the yearly energy consumption of an appliance in monetary terms rather than in physical units has a positive impact on the probability to identify the most cost-efficient appliance.","We use an explanatory research approach in order to examine the role of information, energy literacy, and investment literacy on the choice of the decision-making strategy for choosing an electrical appliance as well as on the ability to identify the appliance with the lowest lifetime cost. For this purpose, we have organized a web-based survey in which two online randomized controlled trials were embedded. The survey was organized in cooperation with three Swiss electricity providers operating in three major urban areas in Switzerland (Lucerne, Bellinzona, Biel/Bienne). The two online experiments were part of the online survey, which was conducted among electricity and gas customers during the year 2015. For this survey, customers of the electricity providers were invited with a letter accompanying one of their electricity (or gas) bills to access an online questionnaire.8 The invitation letter was sent to a total of 50, 000 (Lucerne), 30, 000 (Bellinzona) and 38, 000 (Biel/Bienne) customers of which 1999 (Lucerne), 958 (Bellinzona) and 1308 (Biel) accessed the survey page (corresponding to response rates of 4% (Lucerne), 3.2% (Bellinzona) and 3.4% (Biel/Bienne)). After accounting for the correct target group,9 incomplete responses, and duplicate entries, we have valid and complete data for 1375 (Lucerne), 583 (Bellinzona) and 877 (Biel) survey respondents.10 The three observed samples should relate to the Swiss population living in urban (and sub-urban) areas. Among people who accessed the survey page, dropouts primarily happened either because they did not provide their customer number with the electrical utility (which they had to fill in at the beginning of the survey), or if they were disqualified as not being part of the target group of the survey. Of all respondents who started the online survey, filled in their customer number and were filtered-in as the target group, almost 85% completed the survey and we do not find any significant selection among the sample of usable surveys relative to the target group. Next we compare the available basic demographic characteristics of the three urban areas with the sample population relating to households reached via the survey (Table 1). In terms of gender and age-groups, the population represented by the surveyed households seems comparable to the overall population in the three cities, except in Bellinzona where the share of young population is higher and that of the elderly is lower. The average number of people living in a household seems to be higher in our sample, especially in Bellinzona. Other deviations include the Biel/Bienne sample having a larger living space per person. Bellinzona seems to have more people per room on average whereas Biel/Bienne has less people per room on average than represented by the respective population statistics. Finally, we find that two-person households as compared to one-person households are slightly over-represented across all three samples (data not shown in the table). It is to be noted that the statistics at the city level may not completely reflect the statistics of the surveyed areas, i.e. the service areas of the utilities, which usually also include neighboring municipalities. As the availability of information about socioeconomic variables at the city level is limited, we also compare other crucial attributes with the information available at the national level. The share of respondents who donated money to an environmental organization within the 12 months preceding the survey is largely in line with the share reported for the Swiss population (42%).11 The average Swiss household income in 2013 was CHF 10′052. The household income across our samples is recorded as income categories and about 50% of the households fall in the category representing an household income between CHF 6′000 and CHF 12′000. Considering this discussion and the low response rates, we acknowledge that our sample may not be completely representative of the population living in these regions on all socio- economic dimensions. However, note that a selection issue in this context could be considered less of a concern since we are working in a multivariate framework and are focused primarily on the treatment effects. Nevertheless, generalizations of the results to the entire Swiss population should be considered carefully. One of the two experiments was embedded towards the end of the online surveys. Customers of the utilities in Bellinzona and Lucerne saw Experiment 1, customers of the utility in Biel/Bienne saw Experiment 2. Survey respondents who saw Experiment 1 were asked to imagine a situation in which they had to replace a light bulb in their living room. As a replacement, they were shown two bulbs that differed in their purchase price, power, lifetime, electricity cost as well as energy efficiency rating (A versus C rating). The information display corresponded to the current version of the EU energy label for light bulbs (see Fig. 1). It is important to note that the two described bulbs only differed in price and energy- related characteristics but not in their color temperature and brightness.12 Respondents were randomly assigned to two different treatments. In Treatment 1, the information on energy consumption of the two bulbs was displayed in terms of physical consumption (kWh) per 1000 h, as it is displayed in the current version of the EU energy label. In Treatment 2, energy consumption was again displayed in monetary terms, i.e. in the form of an estimate of the energy cost per 1000 hours (in CHF, see Fig. 1). Additionally, we controlled for order effects by randomly changing the order of presentation of the two light bulbs in both treatments. Respondents were asked which of the two light bulbs would minimize their expenditure for lighting during 8 years of planned usage. Thus, as already discussed in the Introduction, our online experiment is not a hypothetical stated choice experiment but an online randomized controlled trial with the goal of real calculations of lifetime costs of the appliances from an objective point of view under different conditions. In principle, the result of the comparison of lifetime cost will also be driven by the individual's subjective discount rate. Assuming that the average participant of our study is not familiar with the concept of discounting and would need a calculator to incorporate discounting in the analysis, we refrained from providing respondents with a ‘reference discount rate’ that they should use. Instead we assumed that consumers, in case they opted for an investment analysis, would consider the undiscounted future operating cost to evaluate the lifetime cost of the two bulbs. Also Allcott and Taubinsky (2015) present undiscounted operating cost in an online experiment on consumers’ choices of light bulbs, presumably for the same reasons. The respondents who saw Experiment 2 were asked to imagine a situation in which they need to replace their refrigerator. Among two refrigerators, they were asked to identify the refrigerator with the lowest lifetime cost. The two appliances differed only in terms of their purchase price, their energy consumption, and the energy efficiency rating (A+++ versus A+). The remaining characteristics of the two refrigerators were identical. The information was presented in the same way as it is presented in the current version of the EU energy label for refrigerators (see Fig. 2).13 The participants were randomly assigned to one of two different treatments. In Treatment 1, the information on yearly energy consumption of the two appliances was displayed in terms of physical consumption (kWh), as it is displayed in the current version of the EU energy label (see Fig. 2). In Treatment 2, the information on electricity consumption was displayed in monetary terms, i.e. in the form of an estimate of the yearly energy cost. As mentioned, we controlled for order effects by randomly changing the order of the two refrigerators in both treatments. In both experiments, also the order of the answer options was randomized in order to control for order effects in the presentation of the answer options.14 Respondents were asked which of the two refrigerators would minimize their expenditure on the cooling of food and beverages during 10 years of planned usage. Again, the question was not about the respondent's subjective preference for either the one or the other refrigerator, but about which of the two appliances creates less lifetime costs. In Experiment 2, at the Swiss average electricity price of 20 Rp. /kWh, the 100 kWh more energy-efficient refrigerator has an annual energy cost of CHF 20 compared to that of CHF 40 for the 200 kWh refrigerator (Fig. 2). Over 10 years of planned usage, the total energy costs (in the absence of discounting) are CHF 200 and CHF 400 respectively, which from the private perspective does not justify the difference of CHF 500 in the upfront purchase prices. The less energy-efficient refrigerator is hence minimizing the total lifetime costs.15 While the general setup was very similar in both experiments, it has to be highlighted that they also differed in a decisive way: While in Experiment 1 the more energy-efficient bulb was also the one that minimizes lifetime cost, Experiment 2 was designed in a way that the refrigerator with the lower energy consumption, i.e. the more energy-efficient appliance, was not the appliance that minimized lifetime cost. This seems counter-intuitive, as in such a case an ‘energy-efficiency gap’ does not exist. It is perfectly rational for the consumer to choose the less efficient appliance, at least from a private perspective. The reason why this specific setting was chosen was to identify those individuals who performed an investment analysis and to distinguish them from the respondents who followed a heuristic decision-making strategy. If the more expensive (and more energy-efficient) appliance would have had the lowest lifetime cost, individuals applying a decision heuristic (such as choosing based on the energy-efficiency rating provided on the energy label) would have ended up making the same choice as the ones who performed an investment analysis. This would not allow to discriminate between the two groups of decision-makers. In a debriefing question after the choice task, respondents were asked about the decision- making strategy they had adopted when making their choice. Several decision making strategies were offered: one being performing an investment analysis (comparison of lifetime cost) and the other being heuristic decision-making strategies such as comparing the purchasing prices of the two products, comparing the yearly electricity consumption, comparing the energy-efficiency ratings on the labels or making a random choice. The light bulb experiment also had another strategy choice to compare based on the lifetime of the two bulbs. An open answer option “Other reason” was also provided (Fig. 5 in the appendix). Answer options were again randomized to control for order effects. The introduction of this debriefing question gives us the possibility to econometrically analyze the factors that influence the choice to perform an investment analysis, and, therefore, to adopt a rational decision strategy. However, it has to be noticed that in Experiment 2, among all respondents who claimed to have performed an investment analysis as their decision strategy, almost 37% incorrectly selected the more costly fridge as the most (cost-)efficient appliance (in Experiment 1, the number of such cases was about 5.1%). There are two possible explanations for this: either the calculation was not performed correctly or the self-reported decision-making strategy does not necessarily coincide with the actual strategy applied. In Experiment 1, about 28.0% of the respondents claimed to have compared the electricity consumption of the two light bulbs. About 22.9% mentioned that they made an investment calculation before making the choice. Another 24.4% of the respondents reported to have compared the energy-efficiency ratings on the label and 13.0% said they made their choice based on the lifetime of the two bulbs. 2.9% answered that they had compared the purchase prices of the two bulbs and another 7.6% of the respondents either mention other reasons for their choice or report that they made a rather random choice. In Experiment 2, most of the respondents self-reported that they compared the energy-efficiency ratings on the energy labels (45.2% of respondents). 27.7% of the respondents claimed to have compared the electricity consumption of the two refrigerators, and 17.6% claimed that they made an investment calculation to evaluate the two appliances. Only 1.9% of respondents reported to have compared the purchasing prices, 3.0% indicated that they made a rather random choice and 3.5% stated that they had other reasons for their decision.16 Table 2 gives a summary on how the respondents were distributed as the two treatment groups, and also their responses for the two outcomes of focus in this study: the choice of investment analysis as their decision strategy, and the correct identification of the cost-efficient household appliance. In addition to the randomized controlled experiment, the questionnaire included several other questions related to the household's energy consumption as well as questions on sociodemographics of the respondent and other household members. We included information on respondents’ age, gender and level of education, their attitudes towards energy conservation as well as their energy and investment literacy, which is used in the econometric analysis. The gender is represented by a binary variable (FEMALE) which takes value one if the respondent is female. Age of the respondent is captured through three binary variables representing age groups – less than 40 years (AGE40M as reference category), between 40–60 years (AGE40_69) and 60 years or above (AGE60P). The ownership status of the residence (owned or rented) is captured via a binary variable (OWNER) that takes value one if the residence is owned by either the respondent or another member of the household. Monthly gross household income is captured through three binary variables representing income groups –less than CHF 6′000 (HHI6K as reference category), between CHF 6′000 – CHF 12′000 (HHI6_12K) and more than CHF 12′000 (HHI12K).17. The binary variable UNIEDU takes value one if the respondent holds a university degree. The survey language in which the respondent has taken the survey is also accounted for. The dummy variable ITALSP is 1 for survey taken in Italian (applies to the light bulb experiment) and the dummy variable FRENCHSP is 1 for survey taken in French (applies to fridge experiment). All econometric specifications also control for the respondents’ level of energy literacy and investment literacy. Energy literacy is measured by an index (ENLIT_IN) that accounts for several dimensions of energy-related knowledge. This index ranges from 0 to 11 and is based on correct responses to several questions that examine (i) knowledge of the average price of a kilowatt hour of electricity in Switzerland; (ii) knowledge of the usage cost of different household appliances; and (iii) knowledge of the electricity consumption of various household appliances. Investment literacy is measured by a binary variable (INVLIT) that takes the value one if the respondent correctly solved a compound interest rate calculation.18 Compound interest rate calculations are usually used to assess an individual's financial literacy (Lusardi and Mitchell, 2007, 2009; Brown and Graf, 2013). Furthermore, we account for the individual's pro-environmental moral attitude (ATTMORAL) and their concern for free-riding (ATTCONCE) by asking for agreement or disagreement to two statements on a 5-point likert scale. The two statements are “I feel morally obliged to reduce my energy consumption” and “I am not willing to reduce my energy consumption if others don’t do the same”. Each of the two binary variables takes the value one if the respondent chooses ‘agree’ or ‘strongly agree’. Lastly, all specifications control for the treatment effects of displaying the yearly electricity consumption in monetary terms (binary variable TREATCHF), as well as for any effect of the order in which the two appliances are presented (binary variable ORDEFF). An overview of the summary statistics for the variables used in our econometric models for Experiment 1 and Experiment 2 is presented in Table 3. The two samples are found to be similar in terms of ownership, share of higher income households, share of respondents with university education and their attitudes towards environment. The sample in the refrigerator experiment has a larger share of old population than in the light bulb experiment. It also has a higher share of female respondents and low income households. The sample in Experiment 1 appears to have a higher energy-related literacy as well as investment literacy. We also notice that about 23% of respondents in the Experiment 1 employ investment calculation as their decision strategy compared to 18% in the Experiment 2. Unsurprisingly, due to the experimental design, most of the respondents in the light bulb experiment are able to correctly identify the most cost-efficient light bulb. In the refrigerator experiment though, only 1 in 5 respondents correctly identifies the most cost-efficient refrigerator.","The data collected from the survey and from the two experiments allows us to analyze the impact of the display of monetary information about future energy consumption as well as the level of energy and investment literacy on the ability to identify the light bulb and refrigerator with lowest lifetime cost. Further, the debriefing question about the decision making strategy gives us the possibility to analyze the factors that influence the choice to perform an investment analysis. Therefore, from the econometric point of view, we are interested in explaining the impact of information, energy and investment literacy and other socioeconomic factors on two binary outcome variables. In one case, the dependent variable takes the value 1 if someone chose the most cost-efficient appliance and 0 otherwise. In the second case when analyzing the decision strategy, the outcome variable takes the value 1 if someone chose to perform an investment calculation and 0 otherwise. From a methodological point of view, such binary outcome variables are the simplest case of a discrete choice situation and can for example be analyzed using a probit model (Greene, 2003). The separate estimation of two single equation probit models described above is based on the assumption that the two outcome variables can be independently determined. This may or may not hold true in reality. We believe in our context it can be safely assumed that the two outcome variables should be determined jointly rather than individually. In this case, from an econometric point of view, a bivariate probit model could be applied for a simultaneous estimation of the two binary outcome variables. Furthermore, even more precisely, the ability to identify the most cost-efficient appliances could be modeled as a two stage decision process, explaining first the choice to adopt a decision strategy based on an investment calculation, and then the identification of the cost efficient appliance. In this case, an appropriate econometric model for this sequential decision-making process is the recursive bivariate probit as it accounts for the likely endogeneity of the investment calculation variable in the equation related to the identification of the cost efficient appliance.20 Most of the following discussions and results refer to the recursive bivariate probit model, which is our preferred model. Nevertheless, in this paper we still decided to estimate two separate probit models and a bivariate probit model for each of the two experiments for comparison purposes. The exogenous variables in the model comprise socioeconomic characteristics of the respondent (age, sex and university education) and that of the household (household income and if the residence is owned or rented); environment related attitudes; the level of energy and financial literacy of the respondent; the treatment variable (i.e. yearly electricity consumption shown in physical or monetary units) and some other variables like language in which the survey was taken and the order in which the two appliance choices were presented. It is important to note that the model is identified irrespective of whether the exogenous variables x1 and x2 in the two equations are different or not.21 Moreover, such a recursive model of simultaneous equation can be estimated using the full information maximum likelihood (FIML) approach ignoring the simultaneity.22 In a non- linear model, marginal effects are more informative than coefficients, because they inform us how the outcome variable will change when an explanatory variable changes. The marginal effects can be calculated for each observation i or for any specific vector of the regressors. In this study, the marginal effects are calculated for the sample mean. The expression for the marginal effects23 could be derived following Greene (1998) and Kassouf and Hoffmann (2006). For a two equation model, one would obtain a direct effect for variables appearing on the right hand-side of the choice equation (i.e. x2) and an indirect effect for explanatory variables in the decision strategy equation (i.e. x1). The indirect effect on the correct choice occurs via the endogenous decision strategy variable which also appears on the right-hand side of the choice equation. The total effect is then the sum of the direct and the indirect effects. Similarly, the effect of an exogenous continuous variable (say z) is calculated by computing the partial derivative of the unconditional mean function in Eq. (13) with respect to z. The somewhat complicated expression appears in Kassouf and Hoffmann (2006, p. 115) and the analysis appears more generally in Greene (1996, 1998). We are interested in calculating the impact of at least four variables of interest on the choice of decision strategy and in turn on the correct choice of the (cost-)efficient appliance — energy literacy, investment literacy, monetary treatment and the endogenous decision strategy to make an investment analysis. We have used NLOGIT for the model estimation and calculation of the marginal effects presented in this paper.24","Below we present the estimation results for our two online randomized controlled trials. Three probit models were estimated for each of the two experiments for comparison purpose — the single equation probit (Probit), bivariate probit (Biv. Probit) and the recursive bivariate probit (Rec. Biv. Probit).25 These models give us insights into the factors that influence the decision strategy and the identification of the appliance that minimizes total cost over the entire lifetime. Table 4 presents the empirical results from Experiment 1. Following this, Table 5 shows the results from Experiment 2. Finally we compare the results across the two experiments and present the estimated marginal effects for our quantities of interest, i.e. energy literacy, investment literacy, monetary treatment and investment analysis strategy on the identification of the cost-efficient appliances.26 Empirical results ~~~~~~~~~~~~~~~~~ Table 4 shows results of the three model specifications for the light bulb experiment. The slope coefficients in Table 4 are quite similar for the investment calculation equation across the three models. Differences are apparent in the appliance choice equation, particularly when comparing the first two models to the recursive bivariate probit. The correlation coefficient (ρ, shown as RHO(1,2) in the table) between the two error terms is significant only in the recursive bivariate setting. Limiting the discussion to the recursive model which is our preferred framework, we observe a significant negative effect of gender and age (AGE60P) on choosing an investment analysis approach, and in turn on the identification of the cost-efficient appliance. University education, energy and investment literacy, and TREATCHF (i.e. getting the energy consumption information in monetary terms) – all appear to have a significant positive effect on the decision to perform an investment calculation. In the appliance choice equation, AGE60P has a significant and positive slope coefficient likely due to the fact that older respondents could have identified the cost-efficient light bulb via several other decision heuristics. Furthermore, the coefficient on INVCALC is significantly positive which signifies a positive impact of choosing an optimization decision strategy on the probability to identify the cost-efficient appliance. Table 5 shows the results from Experiment 2, the fridge experiment. The coefficients appear similar for the investment calculation equation across the three models and most of the significant effects are also more prominent in magnitude. Higher income levels are seen to be positively associated with the correct identification of the cost-efficient refrigerator in the probit and biprobit settings. The correlation coefficient is significant in both the bivariate and the recursive bivariate setting. We notice that most effects are stronger in Experiment 2, likely due to the fact that making a correct investment calculation was a necessary strategy to identify the cost-efficient refrigerator. Within the recursive bivariate framework, we again find a significant negative effect of gender and age (AGE60P) on choosing an investment analysis strategy. University education has a positive effect. Energy literacy shows a significant positive effect in both equations. Investment literacy and monetary treatment both show a very strong effect on the decision to perform an investment calculation. TREATCHF also has a significant effect in the appliance choice equation now. Moreover, the coefficient on INVCALC is strong and significant highlighting the role of choosing an optimization decision strategy on the identification decision of the cost-efficient refrigerator. Summarizing the most important findings: We observe a positive and significant effect of energy literacy (ENLIT_IN) and investment literacy (INVLIT) on the choice of investment calculation as the decision strategy in both experiments. This supports Hypothesis H1a in that individuals with a higher cognitive ability are more likely to follow an optimization strategy rather than a heuristic strategy. Furthermore, the coefficient on INVCALC is significantly positive which signifies a positive impact of choosing an optimization decision strategy on the appliance choice step, which in turn supports Hypothesis H2a. All these effects are in particular found to be stronger in Experiment 2, wherein making an investment calculation was a necessary strategy to identify the cost-efficient appliance. The effect of being in the treatment with monetary information on yearly electricity consumption is strong and highly significant in almost all model specifications thereby supporting Hypotheses H1b and H2b. As expected, the treatment effects are again stronger in the fridge experiment. No significant order effect is found in either experiment which indicates the absence of any bias due to the order of the presented appliance choices. Marginal effects ~~~~~~~~~~~~~~~~ In Table 6, we present the marginal effects at the sample means for the variables that are particularly important to verify our hypotheses. The marginal effects are computed for the two variables measuring energy and investment literacy, the dummy variable capturing the treatment effect of displaying yearly energy consumption in monetary terms, as well as the endogenous dummy variable of choosing an investment analysis decision strategy. We restrict the discussion to the reported results from our preferred model, which is the recursive bivariate probit. It is shown that the fact that an individual is able to do complex calculations increases the probability to choose the most cost-efficient appliance by 2–3 percentage points in Experiment 1 and by about 19–20 percentage points in Experiment 2. Similarly, an increase in an individual's energy literacy score (measured on a scale of 0 to 11) also increases the probability to choose the most cost-efficient appliance. The higher marginal effect in Experiment 2 (3 percentage points) is likely due to the fact that in this experiment, the most cost-efficient appliance could only be identified when comparing lifetime usage costs of both appliances, which requires some calculation. In Experiment 1, also the comparison of the power of the two light bulbs, their lifetime or the comparison of the energy-efficiency rating on the label lead to the choice of the most cost-efficient appliance. Hence, the ability to make complex calculations was less important here. For the same reason, also the marginal effects of providing the yearly energy consumption in monetary terms rather than physical units are much higher for Experiment 2. While the probability to choose the most cost-efficient light bulb increased by about 4 percentage points when the monetary information was displayed in the light bulb experiment, the probability increased by about 30 percentage points in the case of the refrigerator experiment. This result gives strong support for Hypotheses H1b and H2b, i.e. that the information on yearly energy costs strongly increases the chances that consumers choose the more (cost-)efficient appliance. Lastly, the marginal effect of the endogenous investment calculation decision strategy variable is up to 8 percentage points in Experiment 1 and up to 78 percentage points in Experiment 2. The strong impact supports our Hypothesis H2a that opting for optimization as the decision-making strategy has a positive impact on the probability to identify the most cost-efficient appliance.","In order to examine the role of information display, energy and investment literacy on the choice of electrical appliances of consumers, we have organized a household survey and conducted two online randomized controlled trials among households from three major Swiss urban areas. The first experiment was related to a choice between two light bulbs whereas the second experiment was concerned with making a choice between two refrigerators. The information collected from the survey was analyzed by estimating a series of (recursive) bivariate probit models in order to simultaneously model two binary variables which represent the two stage decision process explaining first the adoption of the decision strategy and then the identification of the cost-efficient appliance. From the survey we observe that more than two-third of the consumers do not perform an investment calculation, which supports the view that a large part of the consumers are boundedly rational and prefer to use shortcuts instead of optimizing their expenditure. Further, from the econometric analysis we observe that displaying yearly energy consumption of appliances in estimated yearly energy cost rather than in physical units increases both the probability that consumers perform an investment analysis and that they identify the most (cost-)efficient appliances. Our results emphasize that informed and rational choices of appliances can be enhanced by the provision of monetary information on yearly energy consumption. Furthermore, we could show that individuals who possess energy-related knowledge and high cognitive abilities, captured by high levels of energy and investment literacy, were more likely to opt for optimization as the decision-making strategy which in turn positively influences the probability to identify the most cost-efficient appliance. It can therefore be concluded, that enhancing an individual's energy-related knowledge and the ability to make complex (investment) calculations seems to be one important prerequisite to empower consumers to make rational and informed energy-related choices. From an energy policy point of view, the results suggest that an improvement in energy efficiency could be reached in three ways. First, with an obligation for the producers of electrical appliances to provide information on the future energy consumption of the product in the form of a monetary estimate. This could follow the example of the EnergyGuide label used in the United States that requires that on certain appliances an estimate of the annual operating energy cost is displayed (US-FTC, 2017). A second strategy would be to educate consumers about the energy consumption of different appliances and how to identify the most efficient appliances by means of brochures and energy literacy courses at schools. As a third strategy, decision support tools such as lifetime-cost-calculators could be provided in stores, mobile applications or through a web-page promoted by the government. All these measures would make use of the insights that energy and investment literacy as well as the display of monetary information seem to improve individual decisions when it comes to the identification of efficient household appliances."],["Using data from 23 sectors in 10 OECD countries over the period 1984–2007 we show that the homogeneity assumption underlying empirical models of capital accumulation may lead to mis-specification. Thus, we adopt a fully disaggregated approach – by asset types and sectors – to estimate the responsiveness of investment to the tax-adjusted user cost of capital. Once all the sources of heterogeneity are accounted for, we find that capital accumulation is significantly affected by changes in the user cost, although the size of the impact is smaller than the unit benchmark. We do not find robust evidence that the long run substitution elasticities are statistically different across asset types. --------------------------------------------------------------------------------","Capital accumulation is crucial for business cycles and economic growth. Understanding its drivers is therefore essential. Among the potential determinants, the literature has extensively investigated the role of the user cost (see Chirinko (1993a, 2008) for comprehensive surveys). Most studies treat capital as a homogeneous good. However, there is motivated concern that the single capital good model inadequately describes the effects of changes in the user cost on capital accumulation, primarily because it neglects compositional shifts in investment. In fact, different capital goods command different prices, display different depreciation patterns, and receive a specific tax treatment. First, market prices vary widely across assets and over time. In some cases, price changes might reflect long-term trends, such as technological progress. For instance, quality improvements in high-tech components have led to a dramatic decline in market prices for computers and similar goods (Greenwood et al., 1997). Secondly, technological features directly affect adjustment costs, which presumably increase with the useful life of the assets. Likewise, the durability of capital goods determines the amount of replacement investment needed to sustain a given level of production, under unchanged technological constraints. Thirdly, the impact of tax policy differs across capital asset types. Tax allowances for depreciation of capital expenditure are typically asset–specific, or defined for relatively narrow asset categories (Clark, 1993; House and Shapiro, 2008). Moreover, even non-targeted tax policy measures translate into different relative changes of the user cost of different assets, depending on its initial level (Cummins et al., 1996). Importantly, both asset and sector specificities matter for the trajectories of capital accumulation. In so far as different sectors are technologically constrained to rely on specific capital assets, investment evolves unevenly across industries. The responsiveness to cost variables changes also if supply is rigid and if the capital assets are not easily redeployable, even within sectors (Goolsbee, 1998). Moreover, increased asset specialization, by reducing the incentives for disinvestment, might alter the sensitivity of investment to its cost. As Desai and Goolsbee (2004) point out, these types of irreversibilities are likely to manifest at the microeconomic level – i.e. at the level of the individual asset and sector – “rather than apply to all assets in all sectors homogeneously” . Abandoning the assumption of homogeneous capital creates challenges for investment modelling. In the context of structural models, the combination of different types of capital goods into a single aggregate imposes unappealing restrictions, on either the level and the shape of adjustment costs (Wildasin, 1984; Chirinko, 1993b), or the degree of substitutability among assets (Hayashi and Inoue, 1991). On the empirical side, aggregation creates issues for the construction and measurement of variables in the first place. Moreover, naturally, econometric models with aggregate variables force homogeneity on the estimated parameters. Likewise, the standard panel data pooled estimators constrain the slopes in the estimating equation to be the same across cross-section units. This might have severe consequences in reduced-form models of capital accumulation resting on the long run cointegrating relationship between the actual and the frictionless level of capital implied by economic theory (Caballero et al., 1995). In this paper, we investigate the consequences of imposing homogeneity when estimating the sensitivity of investment to the tax-adjusted user cost of capital. We use data – including detailed information on business tax incentives – for 23 sectors comprising the market economy in 10 OECD countries over the period 1984–2007. Our setup accommodates heterogeneity across capital asset types and economic sectors. As such, it departs from the bulk of the literature on the substitution elasticity, based on aggregate data (Schaller, 2006; Caballero, 1997; Bond and Xing, 2015). In focusing on asset heterogeneity we take inspiration from Tevlin and Whelan (2003), who reveal the shortcomings of aggregate models due to the rising importance of computers as of the 1990 s in the US. Smith (2008) and Bakhshi et al. (2003) provide similar analyses for the UK. We extend their contributions not only by considering a broader set of assets, countries and sectors, but also by systematically investigating the effects of neglected heterogeneity along these dimensions. Our paper also relates to the recent article by Bond and Xing (2015). They use the same investment data (although a slightly different sample definition) as we do, but still work with aggregate measures of capital. We show that their conclusions do not necessarily survive a finer definition of capital goods. Moreover, we formally examine how heterogeneity affects econometric estimates of the substitution elasticity, something that is inherently different from the focus of their analysis. The remainder of the paper proceeds as follows. Section 2 discusses a way to deal with multiple capital assets in a standard empirical investment model. Section 3 introduces the data, and some stylized facts. Section 4 describes our empirical strategy. The results are in section 5. Finally, section 6 offers some concluding remarks.","The two capital series need not be equal on average, as they can differ up to a stationary error term, e, which captures transitory deviations. In our framework, for a single cointegrating relationship to exist, the capital output ratio and the user cost must be cointegrated. In turn, this imposes constant returns to scale on the production technology, which we assume throughout as in Caballero (1997). Precisely relying on the cointegration between the two capital series, the full specification with short run dynamics can be reparametrized into an error correction model (ECM), as in Bloom et al. (2007). We discuss the empirical implementation of the ECM in Section 4, and now turn to the issue of heterogeneity and aggregation. Aggregation over heterogeneous capital goods creates analogous econometric issues. In the first place, one should ensure consistency between the aggregate measures of quantities and cost variables (Bakhshi et al., 2003). In practice, capital stock series available from the national accounts are obtained additively from the individual series. The corresponding tax-adjusted user cost is built as an aggregate quantity-weighted index of the asset-specific user costs. Fitting an equation for investment requires that, for consistency, the aggregate user cost be built as a price index for investment, though. Thus, the sets of weights will differ unless all capital assets accumulate in proportion to their stock value.3 Even with appropriately defined asset weights, aggregation of the non-price component of the user cost may still be problematic, as it would propagate any measurement errors affecting the tax terms of the single capital assets (Goolsbee, 2000). Overall, the discussion above casts doubt on the fact that ignoring heterogeneity would lead to correctly specified empirical models of capital accumulation. To deal with cross-sectional heterogeneity, we analyze the dynamics of the different capital goods separately and use estimators that allow for heterogeneous parameters. With our data, we can factor in heterogeneity down to the country-sector level.4 Variables and main sources ~~~~~~~~~~~~~~~~~~~~~~~~~~ Our dataset includes 23 sectors (SIC 2-digit classification) adding up to the market economy of 10 OECD economies over the period 1984–2007. Overall, this gives a panel of up to 5060 observations, for 230 country-sector pairs. Details on the sample coverage are reported in Appendix A. Data on production are taken from the EU KLEMS database, which provides harmonized series for capital stock and output at the sector level for European and other advanced economies (O’ Mahony and Timmer, 2009) .5 The stock of fixed capital in the KLEMS data breaks down into several asset types. We focus on the following: computers, communication equipment, transportation equipment, and other machinery and equipment. These assets make up aggregate equipment capital. Adding up structures gives the overall stock of productive physical capital used in the market economy.6 The capital stock series (with 1995 as the base year) are obtained using the Perpetual Inventory method by summing up past real investments, weighted by the relative efficiencies of capital goods at different vintages. Asset-specific depreciation rates are equal across countries and time, but can vary across sectors. They are lowest for structures (the minimum rate is around 2%). When it comes to equipment capital, depreciation rates range from 9% for transportation equipment, to 31.5% for computers. Wear and tear for the different capital aggregates is determined endogenously by the relative importance of the different assets. Average depreciation rates for total capital and equipment are 8% and 14%, respectively, over the sample period. Thus, average values hide significant cross-sectional differences across asset types. Variation along the time-series dimension is equally important. The average economic depreciation rate for total capital increases from 7% to almost 10% over the sample period, while the rate for equipment capital rises from 13% to 16%. This is a result of the dramatic increase in the use of rapidly depreciating equipment capital (see next section). Asset-specific price indices for gross fixed capital formation are also available at the sector-country pair level. The base year for the price indices is 1995. We use value added as a measure of sector output. The corresponding deflator is also taken from the KLEMS database. When it comes to the non-price component of the user cost of capital, the main source for the tax rules is ZEW (2013), which provides disaggregated data by asset and sector according to the KLEMS classification. To fill the gap of missing information in the earlier years of our sample, we have used the International Bureau of Fiscal Documentation (IBFD) and the International Tax Summaries by Coopers and Lybrand. Profit taxes are measured by the headline statutory tax rates on corporate income, augmented by local taxes and surcharges, potentially sector-specific sectors, whenever applicable. Importantly, provisions on tax depreciation allowances and other incentives, such as accelerated depreciation, are also asset-specific. When there are multiple rules under national tax codes, the most efficient scheme is applied. The real discount rate is calculated as the opportunity cost of finance, namely as a weighted average of the cost of equity and the cost of debt, net of CPI inflation. Details on the calculations of the user cost are in Appendix B. Stylized facts ~~~~~~~~~~~~~~ Here we illustrate some key features of the data, which further motivate our preference for a disaggregated approach in modelling investment demand. Fig. 1 plots the capital- output ratio and the user cost of capital for both the aggregate capital stock and equipment capital. The overall capital-output ratio decreases slightly over the sample period. At the same time, the series for equipment capital appears clearly rising as of the mid-1990s, while being relatively flat previously. Taken together, this evidence points to a compositional shift within physical capital. In particular, the increased use of aggregate equipment is accompanied by a decreased importance of structures, which ultimately drives down the aggregate capital-output ratio. The user cost of capital shows a clear downward trend, with the reduction especially marked in the case of aggregate equipment. Fig. 2 depicts the evolution of quantities and prices for disaggregated capital series. There is a clear declining trend in the use of structures (Panel A, left hand side). Likewise, aggregation hides diverging dynamics also within aggregate equipment capital, where a compositional shift towards short-lived high-tech capital assets is apparent. At the same time, other machinery and equipment shows a strong downward trend over the sample period (Panel B, left hand side). Fig. 2 (right hand side column) plots the calculated tax-adjusted user cost of capital for the disaggregated asset series. Again, IT capital assets display similar patterns, with a downward trending user cost, particularly for computers. On the contrary, the user cost of transportation equipment and other machinery and equipment do not show overall clear trends, but rather upward and downward dynamics over shorter sub-periods. The same holds for the user cost of structures (panel A). In logs, the user cost can be expressed as the sum of two components: the relative price of capital and the non-price component, which comprises the cost of finance, the tax term, and the depreciation rate. Fig. 3 depicts the evolution of the relative market price and of the tax term for each of the five assets. As a mirror image of the large rise in volumes, relative prices of both computers and communication equipment (averaged across country-sector pairs) display pronounced negative trends. This is a well-known fact, often taken as evidence of quality improvements stemming from investment-specific technological change (Greenwood et al., 1997). By contrast, the market price for structures is trending upwards, while the relative prices of transportation equipment and other machinery are relatively flat. The tax term of the user cost displays far less heterogeneity across capital assets than the price component (right hand side of Fig. 3). The significant reduction in statutory tax rates on corporate income across OECD countries seemingly lies behind the generalized downward trend observed for the tax term. Short run dynamics are somewhat more volatile, being most likely driven by changes in the asset-specific depreciation allowances. So far we have focused on the dynamic properties of quantities and prices. Taking a look at the cross-sectional variation in the allocation of the capital assets is also useful, as this would give an indication on the degree of heterogeneity in the underlying production technologies across sectors and countries. In Fig. 4 we present box plots for the shares of each capital type into the stock of total capital across sector-country pairs in 2007. Expectedly, structures and other machinery and equipment show the largest median shares. In general, the interquartile range of shares is relatively narrow with respect to the tails of the respective distributions. Moreover, there are quite a few outliers for all assets the short-lived assets.","In addition to coefficient heterogeneity, another important source of concern in estimation is the presence of cross-sectional dependence in the error term due to omitted common factors. Common correlations can arise from macroeconomic shocks affecting all the sectors. Moreover, in our setup, sectoral linkages imply that shocks that are specific to capital-producing sectors propagate throughout the rest of the economy (Foerster et al., 2011). Such interlinkages in the use of capital inputs may effectively transform idiosyncratic shocks into common shocks. While the strength of the amplification mechanism depends on the structure of production linkages between sectors (Horvath, 1998), the scope for transmission clearly increases when an aggregate measure of capital is considered. Common shocks can induce cross-sectional dependence in the residual, and, if correlated with the regressors, result into inconsistent estimates. Likewise, correlation across cross-section units may also lead to significant size distortions in panel unit root tests that assume independence. However, if the extent of cross-sectional dependence of errors is sufficiently weak, or limited to a small number of units, then its consequences in a standard setup are negligible (Chudik and Pesaran, 2014). While the conventional pooled estimators control for the presence of unobserved common factors with the time fixed effects, relaxing the slope homogeneity restriction calls for alternative strategies to deal with such unobservables.10 Specifically, we first implement the Mean Group estimator on cross-sectionally demeaned variables (CDMG), viz., variables measured as deviation from their year-specific average over the whole sample. This procedure eliminates trending components that are common to all sector-country pairs, and thus allows one to deal with common factors affecting capital accumulation, although only imperfectly when slopes are heterogeneous. By augmenting the estimating equations with country-sector specific linear trends we control for group-specific shocks that evolve linearly over time. Time series properties ~~~~~~~~~~~~~~~~~~~~~~ Since the error correction specification rests on the cointegration between the frictionless capital and the actual level of capital, it is important to investigate the time series properties of the variables. To this end, we employ the panel unit root test proposed by Pesaran (2007), which allows for heterogeneity and cross-sectional correlation.11 We run the test for up to three lags, and found that in general the quantity and price variables are non- stationary in levels, both for the raw and the demeaned series. The detailed results are in Appendix C. Investment equations ~~~~~~~~~~~~~~~~~~~~ We estimate Eq. (9) on both aggregate measures of capital and disaggregated asset types. Our composite variables are total capital and aggregate equipment, which differ only because of the inclusion of structures into the former aggregate.12 Subsequently, we split aggregate equipment into its components – computers, communication equipment, transportation equipment and other machinery and equipment. In all cases, we apply the different panel techniques discussed in Section 4. As an extension, we then allow the price and the non-price components of the user cost to have different long run impacts on capital accumulation. Further, we estimate the model for a debt-financed investment. In discussing the estimated coefficients, we focus on the adjustment coefficients and the implied long run substitution elasticity. To put our results in perspective and facilitate comparison with the literature, we test if the substitution elasticity is statistically different from one, the value of the Cobb-Douglas production function. More importantly for our purposes, we also test if the elasticities are equal across asset categories.13 In this way, we can assess whether and to what extent the types of capital respond differently to the price variable. We also perform diagnostic tests for residual stationarity - again using the test proposed by Pesaran (2007) -, and for the presence of cross-sectional correlation (Pesaran, 2004) .14 The root mean squared error (RMSE) statistic is reported as a measure of goodness of fit. Aggregate capital We first estimate the error correction model for total capital and aggregate equipment using the standard two-way fixed effects (2FE) estimator and the MG estimators. The regression results are in Table 1. First, let us look at the models with homogeneous parameters (first and fourth column, respectively). The short run coefficients are all consistent with the underlying theory and significantly different from zero. The point estimates of φ suggest that the speed of adjustment towards the long-run target level is somewhat slower for total capital than for aggregate equipment. This is consistent with previous evidence pointing to a significantly sluggish adjustment of structures (Desai and Goolsbee, 2004; Schaller, 2006). The implied long run substitution elasticity for total capital is statistically different from one at conventional significance levels, whereas the case for rejecting the Cobb-Douglas benchmark is not equally compelling for equipment capital alone. The estimated elasticities are in the upper range of the literature results, in line with studies using firm-level data, such as Cummins et al. (1996), Schaller (2006), and Caballero et al. (1995). The Wald test rejects the hypothesis of equal long run elasticities at 5% level (p-value of 0.012). Residual diagnostics reveal the presence of strong cross- sectional dependence. Moreover, the residuals in the equation for aggregate equipment appear non-stationary, which casts doubt on the validity of the inference drawn for that specification. Allowing for heterogeneous parameters with demeaned variables (second and fifth column in Table 1) results into a faster speed of adjustment and decreased long-run elasticities (in absolute value). In particular, the coefficients are half in size compared to the 2FE estimates. The test for a unit long-run elasticity is rejected for both capital aggregates. Moreover, the Wald test does not reject the hypothesis that the long-run elasticities for the two capital series are equal. While the residuals in both equations appear stationary, demeaning does not alleviate the problem of strong cross-sectional dependence. The estimates from the CCE version of the Mean Group model (third and sixth column) point to long-run elasticities centered on 0.4 rather than one. These results broadly corroborate the findings in Bond and Xing (2015), who estimate elasticities for total capital between −0.5 and −0.3, values within the range of previous findings (Smith, 2008). Again, we cannot reject the hypothesis that the elasticities are statistically equal at conventional significance levels. Importantly, the residuals from the model are well behaved, as they are stationary and reveal only weak cross-sectional dependence. Disaggregated capital We estimate the error correction model in Eq. (9) separately for structures and for the different equipment types – computers, communication equipment, transportation equipment, and other machinery and equipment. We start from the standard 2FE pooled model, and then implement the Mean Group approach. The results are reported in Table 2 and 3, respectively. The speed of adjustment of the capital series to their long-run targets is faster for computers than for other types of equipment, while structures exhibit, expectedly, a sluggish behavior. The long-run substitution elasticity is not statistically different from one in the case of IT capital (computers and communication equipment) and, marginally, transportation equipment. Structures and other machinery and equipment display much lower elasticities (in absolute value), but still statistically significant (see Table 2). The findings are consistent with previous evidence pointing to a relatively high responsiveness of short-lived capital, particularly computers (Tevlin and Whelan, 2003) compared to slowly depreciating assets (Bakhshi et al., 2003). Residual inspection for the different types of equipment does not give fully reassuring results when it comes to stationarity, however. Strong cross- sectional dependence is also an issue for all asset types, except computers. We interpret this result as evidence of the different nature of the unobservable common shocks hitting the different capital goods. Specifically, in the case of computing equipment the shocks seem common to sectors and countries. This is fully consistent with supply side shocks, stemming precisely from the technological improvements reflected in steadfastly declining market prices. As such, these unobservable factors can be adequately controlled for by the time fixed effects in the model with homogeneous parameters. By contrast, unobservable shocks to the other types of capital seemingly have a different nature. Hence, in this case, we expect neglected heterogeneity across countries and sectors to play an important role in contributing to overall cross-sectional dependence. Next, we relax the assumption of homogeneous parameters across country-sector pairs by implementing the Mean Group approach. The results are in Table 3. We first consider variables in deviations from their sample mean in the different years, which allows us to control for unobservables under the maintained assumption that they have common impact on the cross-section units (the corresponding estimates are in the columns with CDMG headings). The estimated parameters of interest are highly significant throughout. The long run elasticities (in absolute value) are half in size compared to the regression with homogeneous parameters. By contrast, the estimates of the error correction term point to a much faster speed of adjustment for all the assets. The residuals are in general stationary. However, in general, the ability to control for strong cross-sectional dependence is not particularly satisfactory. Results from the CCE version of the MG estimator are reported in the CCEMG columns. Strikingly, the speed of adjustment for computers is much faster than the previous estimates would suggest. Looking at the diagnostics shows that the residuals from all the equations are stationary, while cross-sectional dependence also appears significantly reduced for all asset types, except for computers. The combined reading of these regression diagnostics leads us to prefer the CCEMG estimates. The estimated long run substitution elasticities, of the same order of magnitude as the CDMG estimates but with a lower dispersion, are centered on 0.5, a value that does not deviate substantially from the bulk of the results in the literature obtained with different techniques and data samples (Chirinko, 2008). Table 4 reports the p-values of the pairwise Wald statistics testing the equality of the long run elasticities in Tables 2 and 3. The test results for the homogeneous parameter models suggest the clustering of capital assets into two classes. The long run elasticities of the fast depreciating assets are not statistically different from one another, although at varying significance levels. Likewise, structures and other machinery and equipment display statistically similar elasticities. These results are broadly confirmed when the Mean Group estimator with demeaned variables is used. However, our preferred MG estimator with common correlated effects points to a much lower degree of differentiation of the long run elasticities. In particular, the p-values confirm that the hypothesis of equality in general cannot be rejected, although only marginally in the comparison between computers and the long-lived assets, i.e. structures and other machinery and equipment.","Empirical models of investment for aggregate capital may be plagued by inherent biases because of neglected heterogeneity originating from asset and sector specificities. In this paper, we investigate the effects of imposing homogeneity on the long run substitution elasticity using a panel of 23 sectors in 10 OECD countries over the period 1984–2007. We perform the analysis for capital stock aggregates as well as for individual asset types – namely computers, communication equipment, transportation equipment, other machinery and equipment, and structures. We further relax homogeneity by using panel techniques with heterogeneous parameters next to the standard pooled models. We find that the tax-adjusted user cost significantly influences capital accumulation, both for aggregate and disaggregated series. Results from the standard two-way fixed effects model suggest that long-lived assets displays statistically similar long run elasticities, consistent in size with the unit benchmark. We do not find significant differences also among the elasticities short-lived assets, which are, expectedly, also smaller in magnitude. However, conventional panel data models, by imposing parameter homogeneity across countries and sectors, increase the risk of spurious regression and do not correct for cross-sectional correlation in the residuals. In this respect, the homogeneity assumption proves critical for virtually all assets, except computers, for which we can pin down the common supply side nature of technological shocks. Allowing for heterogeneous parameters reduces both the magnitude and the dispersion of the estimated long run elasticities for the different assets types. Once we account for unobserved common factors affecting investment using cross-section averages in the country-industry regressions, we cannot reject the hypothesis that the long run substitution elasticities are statistically similar across asset types. Moreover, we concur with the evidence of a more muted impact than the neoclassical unit benchmark."],["Changes in linkages between growth in the USA, Euro area and China are investigated utilising an iterative procedure for detecting structural breaks in VAR coefficients and disturbance covariance matrix. We find dynamics to be unchanged and, accounting for volatility changes, cross-country correlations are constant until the end of 2007. Although largely isolated from the other large economies until 2007, growth in China is subsequently strongly related to that of the US and the Euro area. The effects are illustrated using generalised impulse responses and forecast error variance decompositions. The increased international synchronisation found may be associated with the effects of the Great Recession on the US and Euro area together with China's extraordinary export growth since joining the World Trade Organisation in 2001. --------------------------------------------------------------------------------","The economic rise of China over the last four decades is well-documented, with its share of world GDP rising from less than 2% in 1979 to almost 15% in 2016, alongside its share of world trade in the export of goods increasing from 0.8% in 1979 to 13% in 2016.1 Indeed, China overtook the US in 2007 to become the world's largest exporter of goods. Although relatively few studies focussed on the role of China in the international economy until its rise was cemented by overtaking Japan as the world's second largest economy in 2009 (by share of world GDP), it is now attracting a great deal of attention. For example, recent studies undertaken within the IMF examine the nature and extent of international spillovers from China, including Arora and Vamvakidis (2011), Blagrave and Vesperoni (2016) and Furceri et al. (2017). Other authors, including Cesa-Bianchi et al. (2012), Dreger and Zhang (2014), Osborn and Vehbi (2015) and Pang and Siklos (2016), also examine how shocks to growth in China affect other economies, while related studies focus on the role played by China for exchange rates and inflation (for example, Granville et al., 2011, Metelli and Natoli, 2017). Although much of this work is motivated by the growing importance of China, empirical analyses nevertheless typically assume constancy over time. The aim of the present paper is to inform discussion about the nature and timing of any change(s) in growth relationships across the world's major economic blocks by applying formal structural break tests to a VAR model for GDP growth in the US, China and the Euro area. Previous studies that consider time-variation in China's relationship with other economies include Fidrmuc et al. (2014), Furceri et al. (2017) and Osborn and Vehbi (2015), but the methods they employ are not designed to pinpoint the nature of change and when this occurred. However, through a structural breaks analysis, we examine evidence for change in the cross-country dynamics of growth, its volatility and the strength of contemporaneous growth linkages. Although methods such as random coefficient models and rolling regressions can be employed to capture change, we prefer to take a structural breaks perspective because it does not require a priori assumptions about the existence or timing of change, and hence may be particularly useful for examining the emergence of China as an economic force. The implications of the breaks we uncover are explored through impulse response functions and forecast error variance decompositions. Following Diebold and Yilmaz (2015), our principal results are not based on any assumed cross-country causal ordering for growth ‘shocks’, but employ the generalised techniques of Koop et al. (1996) and Pesaran and Shin (1998). We employ quarterly data over 1975 to 2015, allowing us to focus on changes in international growth affiliations in the post-Bretton Woods period. Although there would be some advantages in expanding the analysis beyond the US, the Euro area and China, difficulties associated with econometric inference for multiple breaks in a system with a limited amount of data means that parsimony is required in the number of economies included. We study the Euro area as an aggregate, in order to recognise the international importance of this economic region, with aggregate output comparable to the US. Breaks are examined within our three equation system using the iterative testing procedure of Bataa et al. (2013), which not only separates coefficient and covariance breaks, but also further decomposes covariance breaks into variance and correlation breaks. While the broad approach is similar to that employed by Doyle and Faust (2005), who study changes in linkages between G7 countries, ours is more flexible in that we neither specify a priori the number of breaks nor are coefficient and covariance breaks required to be contemporaneous. Further, we separate correlations from volatilities, which is crucial since the former measure the strength of contemporaneous linkages, whereas volatility changes may arise from purely domestic factors. Our results imply that breaks in the contemporaneous correlations of ‘shocks’ are the most important feature of changing international growth affiliations. More specifically, a correlation break around 2007 evidences the growing importance of China, with substantially increased comovement across the three economies after this time. On the other hand, no changes in cross-country dynamic interactions (breaks in the VAR coefficients) are found. Due to the greater integration of China into the international economy, the effect of a one standard deviation ‘shock’ to its growth is associated with strong growth effects for both the US and the Euro area, whereas growth in China was largely isolated from these other economies until 2007. However, the greater integration of China also has the consequence that its growth volatility is also now more closely associated with growth shocks from these other economies. The structure of this paper is as follows. Section 2 discusses data, with Section 3 then outlining our methodology for measuring linkages; an example of the role of volatility breaks and an overview of the methodology employed for econometric inference can be found in the Appendix. Our principal results on growth linkages are presented in Section 4, while Section 5 provides some discussion and conclusions.","Our analysis employs quarterly real GDP growth rates of the US, Euro area and China over the period 1975Q2 to 2015Q2. All data are seasonally adjusted and, except for China before 2011, obtained from the OECD database.2 Data for China starts in 2011Q1 in that database, with growth rates for the earlier period computed using Abeysinghe and Rajaguru's (2004) estimates of real seasonally adjusted quarterly GDP for China. Abeysinghe and Rajaguru (2004) interpolate available annual data through the Chow-Lin technique that exploits information in related quarterly series (namely M1 and total external trade) and observed autocorrelation, and hence the estimated values are anticipated to be more reliable than those based on univariate interpolation. We acknowledge that there is widespread doubt about the quality of historical data relating to the Chinese economy; see, for example, the study of quarterly GDP by Franses and Mees (2013). Nevertheless, there is little that individual researchers can do beyond working with the available data and, despite its limitations, we consider this data to be sufficiently reliable to show the patterns of growth in the real GDP of China. Of course, the Euro area came into existence only in 1999 and its membership has expanded since that date. To maintain a consistent composition, our Euro area data relate to the original ‘Euro 12’ (denoted EU12), namely the twelve countries that comprised the Euro area at the launch of the physical notes and coins in January 2002.3 EU12 is used in preference to an aggregate for the entire Euro area because of the changing country composition of the latter. The growth rate in each case is measured as 100 times the first difference of the log real GDP values. Alongside positive association between US and EU12 growth rates, the rise of China is evident in Fig. 1, with its growth rate typically being substantially above than the others since at least the early 1980s. The Great Recession is clearly visible as a decline in growth for each country around 2008/2009, albeit with that for China remaining positive. The figure also indicates that all three economies may have experienced changes in the volatility of growth over our sample period. Although some changes in patterns may be seen in the figure, it is nevertheless important to undertake formal analysis in order to confirm (or otherwise) their nature, since they could be due to random variation rather than changes in the underlying process. Our analysis employs the quarterly growth rates of Fig. 1. Although some researchers filter GDP growth rate data in order to remove very short run fluctuations and hence concentrate on the so-called business cycle frequencies, such filtering has substantial consequences for the dynamics of the process and hence we prefer to analyse unfiltered growth rate data.","As already explained, our analysis is based on a VAR model for GDP growth in the US, Euro Area and China. In common with many VAR analyses, we employ the tools of impulse response functions and forecast error variance decompositions in order to examine the nature of interactions across variables (in our case, the three economies). However, our analysis is distinctive in two respects. Firstly, employing the methodology of Bataa et al. (2013), we examine whether changes have occurred in the parameters of the VAR; details of the procedure can be found in that paper and is outlined in Appendix 6.2. Sufficient to note here that, although Doyle and Faust (2005) find evidence of breaks in both the VAR coefficients and the covariance matrix for international output growth, such breaks need not occur with the same frequency or at the same dates, as they assume. Previous studies focusing on the univariate properties of output growth imply volatility declines might be anticipated in the early 1980s (see, for example, Sensier and van Dijk, 2004), whereas globalisation may affect dynamic linkages and contemporaneous correlations from the latter part of the century (Kose et al., 2008). Therefore, our analysis first examines whether the coefficients, disturbance volatilities and correlations of our VAR change over time. The second distinctive feature of our analysis is that, when comparing effects over different sub-periods, we allow shocks across economies to be correlated by utilising the generalised methodology associated with Koop et al. (1996), Pesaran and Shin (1998), Diebold and Yilmaz (2015, 2014, 2012), and others. For some VAR analyses, it is plausible to impose restrictions in order to deliver orthogonalised shocks for each equation. However, such restrictions can be difficult to justify for cross-country growth spillovers between the major international economies, and hence we prefer to use generalised measures. Nevertheless, for comparison purposes, we also provide results for the VAR orthogonalised using contemporaneous ordering restrictions, in which the US is ordered first, followed by EU12 and then China. Sections 3.1 and 3.2 describe the linkage measures employed in our analysis. Once the dates of structural breaks are identified, the measures discussed in this section become regime-specific, in that they relate to the estimated model parameter for the specific sub-period of time. When horizons such as one or two years ahead are considered, the measures computed implicitly assume that no structural break occurs within the horizon considered. The calculation of confidence intervals, included in the results of the next section, is discussed in Appendix 6.2. Impulse responses ~~~~~~~~~~~~~~~~~ It is important for the interpretation of both OIRFs and GIRFs to appreciate that these are influenced by both the disturbance correlations and volatilities of the VAR, that is by both P and D of (3). Consequently, the VAR coefficients, disturbance (or shock) standard deviations and correlations all play important roles when measuring the cross- country effects of a shock to growth. In particular, a break in any of the three components will, in general, affect both OIRFs and GIRFs. The simple example in the Appendix provides an illustration of the effects of volatility change in the VAR disturbances on these measures. Growth volatility effects ~~~~~~~~~~~~~~~~~~~~~~~~~ In addition to GIRFs, Pesaran and Shin (1998) define the generalised forecast error variance decomposition (GFEVD). Diebold and Yilmaz (2015, 2014, 2012) build on the GFEVD concept, applying the results to financial markets and, in Diebold and Yilmaz (2015), to the international growth context. Their latter papers (Diebold and Yilmaz, 2015, 2014) refer to GFEVDs as measures of ‘connectedness’, but we prefer to refer to such measures in our context as growth volatility linkages. In any case, some of our definitions differ from the corresponding expressions employed by Diebold and Yilmaz (2015, 2014), as explained below.","We now turn to the principal interest of this paper, namely changes in international growth linkages and China's increasing role in the world economy. Subsection 4.1 provides evidence on the structural breaks in the three-economy VAR model of (1) for the US, EU12 and China, while the implications for international growth linkages are discussed in subsections 4.2 and 4.3. Structural breaks ~~~~~~~~~~~~~~~~~ Structural break testing requires the researcher to set a priori the maximum number of breaks that can occur in the sample period (M) and the minimum percentage (ɛ) of the sample within each regime identified between breaks. We specify these as M = 5 and ɛ = 15%, with the aim of having sufficient observations in each detected regime for reliable inference while also being able to detect important changes during the sample period. It is important to appreciate, however, that these values apply separately when considering coefficients and the covariance matrix, since we employ the methodology of Bataa et al. (2013), outlined in Appendix 6.2. For our sample period, the 15% minimum regime length requires any initial break to occur after the second quarter of 1981 and any final break before the third quarter of 2009, with at least 6 years (24 quarters) between two breaks of the same (coefficient or covariance) form. The maximum of five breaks considered is fairly arbitrary, but appears reasonable in our sample covering four decades. We employ a VAR with p = 1, identified using the Hannan-Quinn criterion and all hypothesis tests are conducted at a 5 percent significance level. Table 1 (panel A) shows an apparent single break in the VAR coefficients in 2009Q2 by the asymptotic WDMax test of Qu and Perron (2007), applied in the iterative coefficient/covariance break testing procedure of Bataa et al. (2013). However, the finite sample bootstrap test of Bataa et al. (2013) finds the single break identified by the asymptotic procedure to be insignificant, with a p-value over 50 percent. Thus we conclude there is no statistically significant change in the VAR coefficients. This initial result is itself notable in the light of the changes in the international economy over the period that we study, and implies that any changes apply within the covariance matrix of the shocks rather than the temporal dynamics. The modelling implication is that all subsequent analysis is based on a VAR with time- invariant coefficients. Panel A of Table 3 presents the estimated VAR coefficients and their significance. These indicate statistically significant growth persistence (positive own lag coefficient) in all three economies. Further, there is evidence of positive Granger causality in growth from the US to EU12, with the reverse (EU12 to the US) coefficient also positive and close to significance at the 5 percent level7. The relative isolation of China from direct dynamic effects originating in the other major economies is seen in the lagged VAR coefficients relating to China (as either the dependent or explanatory variable) being both numerically small and statistically insignificant. In contrast to the constant VAR coefficients, panel B of Table 1 shows three breaks in the VAR covariance matrix, with estimated dates of 1983Q4, 1993Q3 and 2007Q4, and all are highly significant according to both the asymptotic (WDMax and sequential) and the bootstrap tests. Using a sample ending in 2002, Doyle and Faust (2005) identify breaks at similar dates to ours, namely in 1981Q1 and 1992Q2, in their VAR for GDP growth for the G-7 countries. However, the methodology available to Doyle and Faust (2005) considers only coincident breaks across coefficients, variances and correlations, whereas our finding of unchanged coefficients implies that relevant breaks in our VAR are confined to the disturbance covariance matrix. The focus of much of our empirical analysis is, therefore, the nature of changes in the volatilities and cross-country correlations of growth in these major economies and how such changes impact on spillovers between them. When considering cross-market relationships, the finance literature has long recognised the importance of distinguishing between changes in volatility and changes in correlations, since the former may be due to specific market influences whereas the latter measure the strength of interlinkages; for example, see Longin and Solnik (1995). The same considerations apply in our analysis of cross-country growth linkages, and hence it is important to control for volatility changes so that these do not contaminate an examination of correlations. Table 2 therefore decomposes changes in the covariance matrix of our international growth VAR, with volatility considered in panel A and correlations in panel B. A joint bootstrap test that volatility in the three economies is unchanged at each covariance break date is rejected at significance levels of 2 percent or less. Further investigation by considering each country individually indicates that the 1983 volatility break emanates from only the US, which is a manifestation in our data of the so-called Great Moderation (McConnell and Perez-Quiros, 2000; Sensier and van Dijk, 2004; Stock and Watson, 2005), whereas the 1993 break is associated with highly significant volatility reductions for China and the EU12, but the US is unaffected (see the volatility estimates in Panel B of Table 3,8 which are computed with restrictions imposed based on the individual volatility test results of panel A of Table 2 using a 5% significance level). The EU12 volatility reduction around this time has been found in other studies (see, for example, Perez et al., 2006) and can be associated with the move towards greater European integration signalled by the Maastricht Treaty bringing lower growth volatility for the EU12 economy. Interpretation of the volatility decline for China at this date is more difficult, partly because of the lower reliability of China GDP data, particularly in the earlier part of the sample. Finally, the significant change in volatility in 2007Q4 is associated particularly with a substantial increase for EU12, although Table 2 also indicates some evidence (with a p-value of 7.5%) of a change also for the US. This break for the Euro area and (possibly) the US may be associated with the onset of the Great Recession in 2008, with EU12 volatility more than doubling after this break. No break is detected in China at this time. Employing the standard deviation estimates of Table 3, panel B of Table 2 investigates the nature of correlation changes at the covariance break dates. In contrast to the significance of volatility changes at all three dates, correlations alter (according to 5% significance) only at the end of 2007 and hence we recognise only two regimes for the matrix P. Indeed, until the end of 2007, the contemporaneous correlations are small9 and a bootstrap test that the residual correlations for each economy with the other two are jointly zero is not rejected (panel B, Table 2). In other words, prior to the end of 2007, none of the three major economies exhibited statistically significant contemporaneous (within quarter) growth linkages with the other two and the correlations are consequently imposed at zero in panel C of Table 3. Thereafter, the zero correlation null hypothesis is clearly rejected; the resulting correlations of the US with both the EU12 and China are about 0.4, with that between EU12 and China higher at 0.6. The period from the end of 2007 has therefore seen a very substantial change in international growth linkages, from a situation of effectively no contemporaneous association to one of strong and positive cross-economy effects. Our dating of this change is supported by the finding of Fidrmuc et al. (2014) of increased synchronisation of short-term GDP growth between China and individual G7 countries from 2006. Whereas the US and particularly the Euro area have seen relatively weak growth since the onset of the Great Recession, the first decade of the 21st century was notable for China's entry into the World Trade Organisation (December 2001) and the increasing role it has subsequently played in world trade (see the discussion in Section 5). Bearing in mind the small dynamic spillovers to and from China revealed by the VAR coefficients, the period since 2007 is therefore not only one in which China has important interactions with the other two major economies, but the cross-country effects of growth ‘surprises’ are seen quickly, namely within a quarter. For reference, Tables 2 and 3 also provide relevant results for a VAR which ignores covariance breaks. A key consequence of such an analysis is that China would appear to be contemporaneously uncorrelated with the other major economies, with the zero correlation test p-value of nearly 12% in panel B of Table 2 for the no break model. Therefore, in contrast to the strong positive contemporaneous correlations estimated for the period from 2008 in Table 3 (panel C) when covariance breaks are recognised, correlations for China with both the US and EU12 are imposed at zero in the no breaks model. Further, the US-EU12 correlations in the no breaks model are moderate at 0.24. To the extent that our results from 2008 reflect on-going cross-country growth linkages, use of a constant parameter VAR analysis would be seriously misleading. Impulse responses ~~~~~~~~~~~~~~~~~ As noted in Section 3, impulse responses change with any break in the parameters of a VAR model. Based on the identified breaks of subsection 4.1, Figs. 2–4 show growth cumulated GIRFs and OIRFs for a one standard deviation shock applied to each of the three economies. Since the disturbance covariance matrices are diagonal for all sub-periods until 2007Q4 (Table 3), GIRFs and OIRFs coincide until that date. With VAR coefficients constant over time, the estimated impulse responses in each of Figs. 2–4 until 2007Q4 differ only due to the magnitude of the one standard deviation shock applied. Reflecting these size effects, note that the vertical scale sometimes changes. The volatility results in Table 2 imply a single regime from 1984 to 2007 (the period of the Great Moderation) for US shocks, and over 1975Q3 to 1993Q3 for each of EU12 and China shocks. These volatility restrictions are imposed in the results presented, with the sub-periods identified in the headings of the figures reflecting the volatility regimes relevant to the economy of the originating shock. For example, in Fig. 2, a US shock is of magnitude 1.16 percent in the sub-period to 1983Q4, but declines to 0.54 from 1984Q1 onwards. Although no volatility break is detected (using 5% significance) for the US at the Great Recession (2007Q4), the correlation break at this date (see Tables 2 and 3) leads to a new sub-period applying for impulse responses resulting from a US shock. The OIRFs we consider impose a contemporaneous causal ordering of the US, followed by EU12 and then China. Therefore, the GIRFs and OIRFs are identical for the US (the first variable in the ordered VAR) also in the final sub-period and are consequently not shown separately in Fig. 2. For each graph, one and two standard error confidence intervals are included around the estimated responses, with these obtained as discussed in the Appendix subsection 6.2. For reference, each graph also includes corresponding information obtained from a VAR in which constant parameters are assumed (the no breaks model), with that estimated response shown as a blue dotted line and the corresponding confidence intervals by blue shading. Consider, first, GIRFs for US shocks in Fig. 2. As already noted, the magnitudes of the shocks differ over the 1975–1983 and 1984–2007 regimes. Although our model does not find a change in the magnitude of the US shock in 2007Q4 or a change in the VAR coefficients, the width of the confidence intervals for the own US responses substantially increase, due to contemporaneous international linkages now associated with the US. It is also notable that although the point estimates of these own responses from a model with no breaks are reasonably close to those for the 2008Q1-2015Q2 sub-period, the no breaks model implies much tighter confidence intervals. In other words, the no breaks model fails to capture the uncertainty of the most recent period for the US. Responses by the EU12 and China to US shocks, also, of course, shift with the Great Moderation due to the relative sizes of US shocks. Notice also that in both the period before the Great Moderation and since 2007Q4, a one standard deviation US shock leads to a GIRF point estimate for the EU12 growth response of approximately 0.6 percent after about a year. For China, the point estimate response to a US shock is positive only from 2008, from when it is also significant according to the one standard error band. Since the size of the US shock does not change in the latest period, it is now more “potent” due to contemporaneous international linkages. Although the changed role played by China in the world economy is illustrated by it responding to US shocks only from 2008, the GIRFs show US shocks to have positive, and typically highly significant, effects on EU12 growth over the entire sample period. Historical interactions between growth in the US and the Euro area are emphasized by Fig. 3, which shows EU12 shocks to have effects on the US which are significant according to the one standard error bands. Of course, the 1993 volatility reduction for EU12 leads to responses declining in magnitude and the confidence intervals narrowing. The figure includes, for the most recent period, both GIRFs and OIRFs. According to the former, effects of EU12 on the US are largest in the most recent sub-period, due to both increased EU12 volatility and increased US/EU12 correlation (Table 3). However, use of the orthogonalised VAR in the final column diminishes the EU12 role for the US since 2008 relative to the use of GIRFs due to the causality assumed. Except for the period between 1993 and 2007 when growth volatility in EU12 was relatively low, the GIRF point estimates of the own effects of EU12 shocks are relatively constant over time, albeit with wider confidence bands in the post-2007 period. Until the end of 2007, these shocks have inconsequential effects on China, but the substantial correlation of the most recent period leads to positive and significant (according to the one standard error bands) responses after that date, whether measured by GIRFs or OIRFs. Of particular interest for our analysis, Fig. 4 presents impulse responses for China shocks. Throughout the period to 2007, the linkages of China shocks with the other two economies are small in magnitude (indeed, negative for EU12) and not significant according to the one standard error bands; with a diagonal covariance matrix, these effects come only through the VAR coefficients (Panel A of Table 3). With increased contemporaneous correlations from 2008, the GIRFs show strong responses of both other economies to China shocks, with those for EU12 being significant at two standard errors for short lags. With China ordered last in the orthogonalised VAR, the OIRFs in Fig. 4 contrast with the GIRFs for the international spillovers from China shocks for 2008 onwards. In particular, with no contemporaneous effects allowed to flow from China to these other economies, OIRFs show effectively no spillovers from growth in China. We consider such a finding to be implausible and hence concentrate on GIRFs. Therefore, one key result from the GIRFs of Fig. 4 is that China shocks are important for growth in both the US and the Euro area since 2008. This result is driven by the strong positive correlations between the China disturbances and those of the US and EU12 over this period (Panel C of Table 3). Although impulse responses are shown for a model which does not recognise structural breaks, it should be noted that this model has zero correlations for China shocks with those of the US and EU12 and hence the GIRFs for this model find effectively no responses of the other economies to China shocks. This again emphasizes the importance of recognising the possibility of changes in international relationships over our sample period, with breaks in the contemporary correlations of growth shocks being particularly important. Growth volatility ~~~~~~~~~~~~~~~~~ Turning to growth volatility resulting from cross-country shocks, Table 4 provides forecast error variance decompositions in the form of both GFEVDs and (post-2007) OFEVDs. Results for horizons h = 1 and h = 4 are shown, with longer horizons being similar to the latter. The results indicate that, as measured through the GFEVD, growth volatility in all three economies and across all covariance regimes is primarily associated with own shocks. This applies especially for China, where at least 99% of the growth forecast error variance at both horizons considered is associated with own shocks. Although a little lower, the corresponding figures are 90% or more for the US. The lowest percentage applies in EU12, where own shocks are associated with around 85% of volatility at a one year horizon over 1975–1983 and during the European integration phase of 1994–2007. As discussed in subsection 3.2, the use of GFEVDs implies that decompositions do not sum to 100% across shocks unless the covariance matrix is diagonal, which is the case in our model until the end of 2007. However, the positive correlations from 2008 mean that the sums of GFEVDs across shocks for each of the three economies substantially exceed this value in the final sub-period. Although 1983Q4 represents a pure volatility break according to our test results of Table 2, the results in Table 4 indicate that this brings about a marked change in US/EU12 volatility linkages. In particular, whereas US shocks are associated with about 15% of EU12 growth forecast error volatility at h = 4 in the earlier sub-period, the decline in US volatility during the Great Moderation causes this to drop to only 3% over the decade from 1984, before subsequently increasing again. Until 2008, however, volatility in China is effectively isolated from these other major economies. Not surprisingly in the light of the international shock correlations after 2007, GFEVDs show shocks in each of the other countries to be important (and generally statistically significant) for growth volatility in all three economies in the final sub-period. On the other hand, the use of OFEVDs leads to weaker effects to China from both the EU12 and its own shocks in this recent period. The final column block, labelled From Others, shows the total volatility not associated with own innovations, as defined by (7). As discussed in section 3.2, the GFEVD measure we present here excludes all effects associated with own shocks. In 1975–1983 and 1994–2007, the converse of the lower volatility percentage accounted for by own shocks in EU12 is that international growth volatility linkages to this economy from others are larger (and often more statistically significant) than for other economies and other sub-periods. However, the strong post-2007 shock correlations lead to values that are relatively small for all three economies in this period, and these are all less than one standard error in magnitude. Bidirectional comparisons obtained using (6) are also shown in Table 4. The GFEVD results indicate that net volatility linkages from the US to EU12 are positive until the end of 2007, with the exception of the sub-period following the Great Moderation (1984–1993). Over the remaining two sub-periods US shocks are estimated to account for substantially more EU12 forecast error volatility than the EU12 does for the US, underlining the international role played by the US. Although negative, the net US-EU12 GFEVD values are very small post-2007, again reflecting the strong shock correlation during this time. Perhaps surprisingly, the GFEVD estimates indicate that the US has a negative net growth volatility linkage to China (that is, the net value is in the direction of China to the US) over all sub-periods, but the values are relatively small and typically less than one standard error in magnitude. This last comment applies also when net GFEVD volatility spillovers between EU12 and China are considered. Orthogonalisation increases net growth volatility linkages post-2008. Of course, if a constant parameter model is employed, the changes over time discussed above that arise as a result of both volatility and correlation breaks cannot be detected. Nevertheless, the general patterns can be seen of international effects on growth volatility being most marked for EU12, with positive net volatility bilateral linkages from the US to EU12, but net linkages with China being small. An interesting comparison between the two (generalised versus orthogonalised) FEVD approaches is provided by the average of the percentage volatility from others, as defined in (8) and shown in the bottom horizontal block of Table 4. The averages obtained from the GFEVD are relatively small (6% or less) and have not increased in the recent period, suggesting that international growth volatility linkages have remained rather muted throughout our sample period. In contrast, causal ordering due to orthogonalisation suggests a huge increase in growth volatility due to shocks originating in other economies in the period after 2007Q4, being over 20% for EU12 and around 40% for China. Once again, however, we consider these orthogonalised results to be a consequence of the imposition of ordering restrictions that do not reflect current macroeconomic relationships. Indeed, increased contemporaneous correlation but low net growth volatility linkages (as indicated by FEVDs post-2007) are consistent with increased synchronisation of growth across these major economies.","Although it is beyond the scope of the present paper to analyse in detail the reasons why China has become such a force in the world economy, some comments are nevertheless in order. Many recent studies, including Autor et al. (2016), Caporale et al. (2015) and Yao (2014), point to China's remarkable growth since the 1970s being led by exports. After increasing quickly from then until 2008, Yao (2014) also notes that China's share of world exports has subsequently been in line with its share of world GDP, which is compatible with our finding of a new regime of China's integration in the world economy from 2008. Caporale et al. (2015) study the changing composition of China's trade, documenting a shift from labour-intensive to capital- and technology-intensive exports over two decades to 2012; see also Autor et al. (2016). The key to China's increased trade in the current century is its accession to the World Trade Organisation at the end of 2001 (Autor et al., 2016; Yao, 2014). In particular, China's exports of goods rose dramatically after it joined the WTO, and (as noted in the Introduction) it's share overtook the US in 2007. This suggests our finding of greater integration for China with the US and Euro area from 2008 is associated with its role as a major trading nation. Empirical work of the type undertaken here might be extended to other countries, to examine whether the changed affiliations of China with the US and Euro area apply more generally and with similar dates of change. Based on our findings, such analyses might place particular focus on the role of trade and China's accession to the WTO. The results of this study are in line with these analyses and emphasise the key role played by China in the world economy over the last decade or so. Using a VAR model to capture GDP growth interactions across the US, Euro area and China, our principal finding is the substantial increase in the contemporaneous correlations of cross-country disturbances at the end of 2007. Using wavelet analysis, Fidrmuc et al. (2014) also detect increased synchronisation of GDP growth in China with major economies from around this date. Since China's growth was effectively isolated from influences from these other major economies until 2007, this most recent sub-period is very different from earlier ones in terms of the synchronicity of international growth. Consequently, cross-country shock responses to and from China are not only more marked, but these occur more quickly then previously, with the China-Euro area growth relationship being particularly strong. Volatility linkages have also become important since 2008, with China's volatility more strongly associated with growth shocks in the other economies. Interestingly, however, net growth volatility linkages in this period are in the direction of China to the US, underlining its increased role for even the largest economy in the world."],["Environmental conditions in early life are known to have impacts on later health outcomes, but causal mechanisms and potential remedies have been difficult to discern. This paper uses the Nepal Demographic and Health Surveys of 2006 and 2011, combined with earlier NASA satellite observations of variation in the Normalized Difference Vegetation Index (NDVI) at each child's location and time of birth to identify the trimesters of gestation and periods of infancy when climate variation is linked to attained height later in life. We find significant differences by sex: males are most affected by conditions in their second trimester of gestation, and females in the first three months after birth. Each 100-point difference in NDVI at those times is associated with a difference in height-for-age z-score (HAZ) measured at age 12–59 months of 0.088 for boys and 0.054 for girls, an effect size similar to that of moving within the distribution of household wealth by close to one quintile for boys and one decile for girls. The entire seasonal change in NDVI from peak to trough is approximately 200–300 points during the 2000–2011 study period, implying a seasonal effect on HAZ similar to one to three quintiles of household wealth. This effect is observed only in households without toilets; in households with toilets, there is no seasonal fluctuation, implying protection against climatic conditions that facilitate disease transmission. We also use data from the Nepal Living Standards Surveys on district-level agricultural production and marketing, and find a climate effect on child growth only in districts where households’ food consumption derives primarily from their own production. Robustness tests find no evidence of selection effects, and placebo regression results reveal no significant artefactual correlations. The timing and sex-specificity of climatic effects are consistent with previous studies, while the protective effects of household sanitation and food markets are novel indications of mechanisms by which households can gain resilience against adverse climatic conditions. --------------------------------------------------------------------------------","Attained height is among the most important indicators of childhood health or deprivation. Approximately 25 percent of each year’s worldwide cohort of infants grow up to be stunted, and the dietary or disease conditions that limit linear growth in childhood also contribute to poor educational attainment, low earnings, and high mortality rates later in life (IFPRI, 2014; UNICEF, 2015). Stunting rates have been especially high in Nepal, where extreme poverty and political instability led to rates as high as 57 percent in 2001, before declining to 41 percent in 2011 (United Nations, 2013). Despite this improvement, Nepal remains one of the 10 countries in the world with the highest stunting prevalence (UNICEF, 2014), making it a high-priority location for research into increasingly effective ways of protecting children from harmful early-life circumstances. Socioeconomic factors associated with stunting in Nepal are described by Headey and Hoddinott (2015), who show how changes in household and community-level characteristics help explain local variation and the overall improvement in this indicator from 2001 to 2011. Key changes involved both greater household sanitation and access to improved diets, which are pillars of the Nepal government’s multisector nutrition plan (Nepal NPC, 2012). Despite this progress, however, poor sanitation and inadequate food intake remain widespread and are likely to be worsened by rising temperatures and more variable rainfall associated with climate change (IPCC, 2014). This paper uses satellite data on vegetation near each child’s home as an indicator of changing agroclimatic conditions, with randomness in the month of birth providing a natural experiment in the timing of exposure to more or less advantageous circumstances. Our use of variation in birth timing relative to changes in climatic conditions contribute to a rapidly growing body of literature using natural experiments to study the determinants of human health (Angrist and Krueger, 2001; Akresh et al., 2011; Lokshin and Radyakin, 2012; Tiwari et al., 2013; Brown et al., 2014), addressing the timing and mechanisms by which early conditions influence later outcomes (Skoufias and Vinha, 2012; Kumar et al., 2016 ; Schultz-Nielsen et al., 2016). Our identification strategy takes a difference-in-differences approach, testing whether household sanitation and district-level food markets can protect children against the health consequences of unfavorable agroclimatic conditions at sensitive times in their early growth and development. The specific data we use are the Nepal Demographic Health Survey (NDHS) for child health, sanitation, and other household characteristics from 2006 and 2011, combined with Normalized Difference Vegetation Index (NDVI) data from the National Aeronautics and Space Administration (NASA) for 2000–2012 at each child’s location, and the Nepal Living Standard Survey (NLSS) to characterize local agricultural markets for 2003–2004 and 2010–2011. By combining three kinds of data, we are able to identify patterns in attained heights of children observed at 12–59 months of age, and test whether sanitation and food markets limited their association with agroclimatic conditions experienced during gestation and the first year after birth. We find the underlying patterns to be sex-specific, with systematic differences in how later heights relate to NDVI fluctuations that occurred during infancy and pregnancy. These differences are consistent with both gender bias in infant care (Maccini and Yang, 2009) and physiological differences in fetal development before the sex of the child is known (DiPietro and Voegtline, 2015; Rosenfeld, 2015). We find that improved household sanitation and more commercialized food markets limit both kinds of vulnerability, providing significant protection from agroclimatic conditions for both pregnant mothers and infants. Agriculture and climate in Nepal ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Nepal is a landlocked country with a population of approximately 27 million people, of whom about 85 percent live in rural areas (Nepal MoHP, 2012) and are highly reliant on rain-fed agriculture (Nepal MoAD, 2013). The country features three distinct ecological zones: Mountains (52,000 km2), Hills (61,000 km2), and Terai or lowlands (34,000 km2), with varying population densities. The Mountain zone has a dry alpine climate and is situated at the highest altitude (>2500 m), with steep and rugged terrain and short growing seasons. The Hills have a mostly temperate climate (500–2500 m), and the Terai (<500 m) has a mostly subtropical and humid climate (Nepal MoHP 2012). Although the Terai occupies 23 percent of the country’s landmass, it hosts almost half of the population (48 percent) and most of the cultivable land (56 percent) (Nepal MoAD, 2013). The most commonly grown crops are cereals including maize, millet, barley, rice, and wheat (WFP, 2014). Consistent with global trends of increasing temperature and erratic rainfall patterns (NASA, 2015), temperatures in Nepal increased by 1.5 °C over the period from 1978 to 2005 (Krishnamurthy et al., 2013), while rainfall declined in frequency and increased in intensity (Malla, 2008). The impacts of climate trends and fluctuations can be seen through changes in sowing dates, crop duration, crop yields, and management practices (IPCC, 2014). Between 1978 and 2008, the summer months (May–August) became increasingly hot and wet, and winter months (November–February) became colder and drier. During that time, the higher levels of rainfall in summer increased rice yields but decreased yields for other crops, while lower levels of rainfall in winter decreased maize yields (Joshi et al., 2011). For this paper, we use NDVI data to summarize the complex pattern of variation in both rainfall and temperature, providing a simple index of changing agroecological conditions in the area around each child’s home. Seasonality and child nutrition ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Seasonal variation and other climatic changes have a clear link to the nutritional status of children in many contexts, even in industrialized countries (Chodick et al., 2009). In the UK, for example, babies born in winter have significantly lower birth weights, educational attainment, and adult heights, perhaps as a result of low vitamin D levels during early life (Day et al., 2015). In the United States, children conceived in the summer have a higher prevalence of birth defects (McKinnish et al., 2014) and different genetic characteristics (Rietveld and Webbink, 2016). Some seasonal patterns may be the result of selection effects, as Buckles and Hungerman (2013) show in their study of winter births in the United States which occur disproportionately among disadvantaged youths. However, in developing countries, studies have repeatedly found relatively large agroclimatic patterns that cannot be explained by selection effects. For example, in the Democratic Republic of Congo, Darrouzet-Nardi (2015) shows that children born during wet seasons grow up to be shorter, with no evidence for selection effects of adverse birth timing of children with lower levels of household wealth or education. Agroclimatic fluctuations may affect child nutrition through both disease risk and dietary intake. A principal source of variation in both kinds of risk is rainfall: children born during monsoon months in India have lower height and weight than children born during the fall–winter months (Lokshin and Radyakin, 2012). In addition, rainfall fluctuations in Indonesia have been shown to affect child health in both rural and urban areas (Yamauchi 2012; Cornwell and Inder, 2015). These associations often depend on the timing of exposure. For example, Tiwari et al. (2013) show that in Nepal, a child’s weight for age is positively correlated with rainfall in the previous monsoon season, but negatively correlated with rainfall in the current monsoon. Temperature may play an independent role, as suggested by Hu and Li (2016), among others, although Nepal’s complex topography complicates efforts to analyze the effects of spatial variation in temperature. In any case, covariance among climatic variables, agricultural conditions, and dietary intake makes it difficult to distinguish one factor from another. In Malawi, for example, the prevalence of underweight among children under age five rises during the rainy season, which is also the preharvest period, when maize prices are highest (Sassi, 2015). Vulnerability in utero and after birth ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Gestation and the first two years after birth are the most critical periods for child development, and have clear impacts on physical, cognitive, and other outcomes later in life (Almond, 2006; Black et al., 2013; Hoddinott et al., 2013). Adverse conditions in utero and during the first two years of life can cause high perinatal mortality and subsequent stunting (Coffey, 2015) as well as low weight and anemia (Kumar et al., 2016), low economic productivity (Paxson and Schady, 2007; Hoddinott et al., 2013), and less academic success (Schultz-Nielsen et al., 2016). Affected children may also have higher odds of developing cardiovascular disease and other conditions (Popkin et al., 1996; Sawaya et al., 2003), as set forth in the fetal origins hypothesis of Barker (1995). This paper expands on the previous literature by using quarterly variation in NDVI to identify sex-specific differences in the timing of vulnerability before and after birth, and to test for the possible protective effects of improved sanitation and food markets. Our focus on sex differences follows that of Skoufias and Vinha (2012), and our attention to the exact timing of exposure follows Andalon et al. (2014) and Carlson (2015), among others, building on previous work in South Asia using Demographic and Health Surveys (DHS) in both India and Nepal that suggest a stronger link between rainfall variation and child height during early months in infancy than during other periods of a child’s life (Lokshin and Radyakin, 2012; Tiwari et al., 2013). Moreover, focusing on the growing seasons in Nepal, other researchers find that anomalies in vegetation density show higher correlations with stunting during the in utero and infancy phases than in other phases of a child’s development (Shively et al., 2015). The timing and magnitude of vulnerability to agroclimatic conditions could differ by the child’s sex. There is a growing body of evidence suggesting differential effects of prenatal stress on male and female fetuses on perinatal outcomes (Aibar et al., 2012; Mulla et al., 2013; Persson and Fadl, 2014) and adult health (Scholte et al., 2015). In general, female fetuses are more resilient and adaptive to stress than are male fetuses (DiPietro and Voegtline, 2015; Rosenfeld, 2015). Evidence from studies that examined the effects of stressors such as intrauterine lead and pesticide exposure, and maternal alcohol and drug use, suggest that exposed male fetuses are more likely to be born preterm and have poorer scores on developmental assessments than females (Rosenfeld, 2015). Historical data from Danish cemeteries suggest that males’ heights increased during the 19th century while females’ did not (Jørkov, 2015). In malnourished populations today, on average boys are more likely to be stunted than girls, but after birth, when the child’s sex is known, gender discrimination may play an important role in health outcomes. For example, Maccini and Yang (2009) show that Indonesian families commonly protected boys more than girls from early-life shocks, and countries with more gender discrimination in favor of boys have lower rates of stunting in males relative to females (World Bank, 2008) with intergenerational effects (Osmani and Sen, 2003). Protective effects of sanitation and food markets ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Numerous interventions could provide protection against the disease transmission and dietary inadequacy associated with agroclimatic conditions. In this paper, we study two kinds of variables that are of particular interest to policymakers: household sanitation and local food markets. Sanitation has become an increasingly important policy tool in recent years, especially for South Asia (Hammer and Spears, 2016), while the impact of agricultural commercialization on nutrition has been of longstanding concern around the world (Von Braun and Kennedy, 1994). The potential efficacy of household sanitation against stunting operates through preventing fecal–oral transmission of diseases such as diarrhea and enteropathy, which cause both mortality and stunting through loss of nutrients, decreased absorption of nutrients, and a weakened immune system (Checkley et al., 2008; Humphrey 2009). There may also be selection effects, as households with toilets may have other favorable conditions for child development, but the prevention of disease transmission provides a clear causal mechanism to explain how sanitation might improve nutritional status (Coffey and Geruso, 2015) and linear growth (Checkley et al., 2004; Lin et al., 2013). In this study, we focus on the average effects on each child of having a toilet in their own household; future work could address the externalities among households described by Hammer and Spears (2016). The potential efficacy of food markets is likely to operate primarily through dietary intake, as shown for Ethiopia by Abay and Hirvonen (2016). As shown by Puentes et al. (2016), total intake, especially of nutrient- rich foods, can have a major impact on child growth. In this study, we focus on whether households have access to food from elsewhere than their own farm production, by using their district’s share of total food consumption that is either purchased or received in- kind, as opposed to consumed on-farm. We employ district fixed effects to control for other time-invariant district characteristics, thereby isolating the specific effect of being in districts where households are more reliant on local conditions and unable to use markets to improve dietary quality, as in the study by Sibhatu et al. (2015). Identification strategy and data sources ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our study uses a two-stage difference-in-differences design, comparing the association between child height and earlier agroclimatic conditions among children with different birth exposure, in households with and without toilets, and in districts with low and high food market use. The first set of differences exploits a natural experiment in exposure to different levels of NDVI at each stage of early development, while the second splits the sample into two groups to isolate the protective effects of sanitation and food markets. This strategy relies on combining three distinct kinds of information: NDHS data on child height and household characteristics, NDVI data on agroclimatic conditions, and NLSS data on agricultural commercialization. With these datasets, we are able to address a number of potential threats to identification: first, showing that the month of conception is uncorrelated with maternal and household socioeconomic status or other characteristics; second, using fixed effects and statistical controls to narrow the parallel-trends assumption implicit in our method; and third, performing a set of placebo regressions to demonstrate that results are not an artifact of the method. In all regressions, we control for observables known to correlate with maternal health and child size, such as parental education, maternal body mass index (BMI), household wealth, and altitude (Wehby et al., 2010). The NDHS is a comprehensive, nationally representative survey typically conducted every five years to gather data on population and health. The NDHS is carried out under the Ministry of Health and Population and conducted by New ERA as part of the worldwide DHS program. The DHS uses standardized questionnaires and fieldwork to allow comparisons across years in demographics and health and nutrition–related variables. This paper uses two recent rounds of the NDHS conducted in 2006 (February–August) and 2011 (February–June). Both survey rounds use two-stage, stratified sampling, including households in all 75 districts across all ecological zones and development regions. The 2006 survey uses the 2001 population census for a sampling frame, while the 2011 survey uses an updated 2001 population census for a sampling frame that accounts for population growth and internal and external migration. Anthropometrics for children under five years of age were collected from 5237 and 2335 children in 2006 and 2011, respectively (Nepal MoHP, 2007, 2012). The NDVI is a measure of vegetation density resulting from the interaction among rainfall, temperature, and soil fertility over time. Green vegetation, owing to the presence of chlorophyll, absorbs red (visible) light and reflects near infrared light, while sparse vegetation reflects more red light and less near infrared light. NDVI values are measured as NDVI = NIR − RED/NIR + RED. Here, the numerator denotes the normalized difference between red and near infrared light bands, and the denominator is the sum of red and near infrared light bands (Weier and Herring, 2000). Those data are obtained from NASA satellite remote sensors using a Moderate Resolution Imaging Spectroradiometer (MODIS) Climate Modeling Grid (CMG), which provides data at 5 km resolution (Shively et al., 2015). The dataset includes monthly NDVI values for 12 years (February 2000–May 2012) for each child’s location of birth, corresponding with 260 and 289 clusters in 2006 and 2011 NDHS, respectively (Nepal MoHP, 2007, 2012). The full distributions of NDVI levels for each month are shown in Appendix A. This variable provides an attractive measure of green biomass and leaf area (Thenkabail et al., 2009), which in turn depend on available moisture, temperature, and soil fertility as well as human intervention in response to those underlying conditions (Laidler et al., 2008). NDVI is also influenced by factors other than plant growth, such as cloud cover and ground conditions. Although NDVI is not a simple index of climatic conditions or crop production (Shively et al., 2015), it does provide a powerful measure of trends and fluctuations in various agroclimatic conditions that could affect child development. The NLSS is a comprehensive national, multitopic household survey conducted by the Nepal Central Bureau of Statistics (CBS), using survey methods developed and promoted by the World Bank. The NLSS uses multistage and stratified sampling with the Population Census 2001 of Nepal as a sample frame basis. Data collection spans an entire year, unlike the NDHS, in order to capture the effects of seasonality. The purpose of the NLSS is to assess changes in the population’s living standards using a combination of panel data and cross-sectional data for a given survey round. A wide range of topics covered by the NLSS includes poverty and access to finances, health, education, agriculture and rural development, and labor markets (Nepal CBS, 2004, 2011). This paper, however, uses only agriculture and rural development data, which include a range of information on food consumption, food production, and related expenses. In order to align the time frames between NLSS and NDHS data, the NLSS II 2003–2004 is merged with the NDHS 2006, and the NLSS III 2010–2011 is merged with the NDHS 2011 (Shively et al., 2015).","Our main outcome variable is each child’s height-for-age z score (HAZ), defined as the gap between that child’s measured height and the median height of a healthy population at each age and sex, expressed in terms of standard deviations (SDs) of the healthy population (WHO and UNICEF, 2009). Children are classified as stunted when their HAZ is two or more SDs below the median, but here we focus on HAZ scores as a continuous variable to capture variation at every level of attained height. In addition, although the NDHS enumerators measured the length or height of all children under five years of age, here we focus on heights attained between the child’s first and fifth birthdays, as a function of agroclimatic conditions experienced both in utero and during the child’s first year after birth. Control variables include the child’s age in months at the time of measurement, total number of siblings ever born, maternal age, maternal education and BMI, household wealth, altitude and region, urban residence, and survey round. Descriptive statistics ~~~~~~~~~~~~~~~~~~~~~~ Table 1 shows means and SDs for nutritional outcomes and control variables in addition to other variables of interest. All data are from the NDHS 2006 and 2011, except for district-level data on food market participation and distance to market centers, which are obtained from the two waves of the NLSS. Because later regressions will be performed on data split by the sex of the child, we show the pooled dataset of all children aged 12–59 months (N = 6127), and the subsamples of boys (n = 3129) and girls (n = 2998). The summary statistics in Table 1 show that the two subsamples are balanced in all regards, except that the male subsample is larger and the females have more siblings, which is consistent with sex-selective stopping rules by which parents seek additional children until they reach their desired number of boys (Bongaarts, 2013). Turning to agroclimatic conditions, the NDVI levels to which each child is exposed at each stage of development depends on the month and location of birth. Fig. 1 shows the seasonal patterns of NDVI variation in Nepal’s three main regions. Vegetative cover generally peaks during August–October, with greater month-to-month variation occurring in the Terai region. There is greater uncertainty in calculating each month’s mean NDVI for the Mountain region, partly because of smaller sample size: only 87 of the 547 cluster locations for our two NDHS surveys were in the Mountains, while the rest were almost equally distributed between Hills and Terai. Exploratory regressions ~~~~~~~~~~~~~~~~~~~~~~~ Our identification strategy relies on randomness in birth timing relative to variation in agroclimatic conditions around the home. Before proceeding, we test for possible selection effects, asking whether some kinds of mothers are more likely to conceive in certain months of the year. Seasonal patterns of conception could result from seasonal migration of family members, variation in natural fertility, or even deliberate pursuit of conception at more favorable times. Table 2 tests for selection patterns in a very general way, using a multinomial logit model of selection for each month of conception, nine months before the observed month of birth. The results of Table 2 show 12 of 121 coefficients with statistical significance at the 5 percent level or higher. Half of these are the significant coefficients on altitude between February and August. During that period, families at higher altitudes are more likely to conceive a child than families at lower altitudes. Lower altitude places are relatively hot in the summer, so this result is consistent with Levitas et al. (2013) who found sperm counts and motility to be lower at warmer times and places. The remaining six coefficients which are significant at the 95 percent level have no clear pattern, confirming the absence of selection effects on the timing of conception in this context. These results contrast with the strong selection effects in the United States (Buckles and Hungerman, 2013) but are consistent with studies elsewhere, such as Taiwan (Fan et al., 2014), India (Lokshin and Radyakin, 2012), and Nepal itself (Panter-Brick, 1996). Without selection by education or other parental characteristics for specific birth months, differences in attained height by month of birth are most likely the result of environmental exposure effects. To identify this purely seasonal pattern in how birth timing relates to attained heights, we employ ordinary least squares (OLS) regressions using the following base specification:Yi = β0 + β1m birthmonthim + δiZi + ui,where Yi indicates the HAZ score of a child aged 12–59 months, birthmonthim is a vector of dummy variables equal to 1 when m = the birth month of that child and zero otherwise, and Zi is a vector of control variables at the child-, maternal-, household-, and district levels. District fixed effects are also included, and standard errors are clustered by birth year and district. To account for the effects of a child’s age on measured HAZ, as in Cummins (2015), we use two different specifications: age and age squared in columns 1, 3, and 5, and a linear year of birth term in columns 2, 4, and 6. The two types of age control yield a similar month of birth effects, so for the remaining tests, we use only the simple age in months and months-squared specification, and focus attention on the differences between boys and girls in regard to what birth timing might be relatively unfavorable. Results in Table 3 reveal that the worst month for all children to be born is April, but boys are also disadvantaged by being born in June and September relative to the omitted month, January. As shown in Fig. 3 .1, April is one of the lowest-NDVI months, while September the highest. A few other months have some correlation with attained heights, and the magnitude of these seasonality effects are very large: for all children, being born in April has an effect size similar to 1.9 quintiles of wealth, and for boys, being born in September has an effect size similar to 2.4 quintiles of wealth. Empirical methods ~~~~~~~~~~~~~~~~~ To investigate causal mechanisms and potential remedies for the effect of birth timing on attained height, we turn to variation in NDVI at each stage of child development during pregnancy and the first four three-month periods after birth. Our focus is on heterogeneity and effect modifiers, so we divide the sample into boys and girls to identify sex-specific vulnerability at each stage of early child development, and then test for the protective effects of sanitation and food markets by dividing the sample into households with or without toilets and districts with above- or below-median reliance on local production as opposed to food purchases or gifts. This split-sample approach, made possible by our large number of observations, permits variation among subsamples in all coefficients while avoiding a proliferation of collinear interaction terms. In this design, exposure to climate variation is a natural experiment to which children are exposed under conditions that might or might not be protective. Testing for heterogeneity in treatment response can help identify causal mechanisms, target services to those most at risk, and help spread desirable effect modifiers. In this case, we seek to identify the protective effects of toilets and food markets against adverse “treatments” that cannot be experimentally assigned, but are experienced randomly by children in utero and during the first year after birth. In each subsample, we employ OLS with the following base specification:Yi = β0 + β1t NDVIit + δiZi + ui Our notation is the same as for Eq. (1), except that the vector of NDVIit conditions is the average value of NDVI in each three- month period t, from the first trimester of pregnancy through the first year after birth when the child is aged 0–2 months, 3–5 months, 6–8 months, and 9–11 months. All Zi control variables are the same as in our exploratory regressions, first using both types of age controls, and then using only the more flexible quadratic approach to controlling for the child’s age at the time of measurement. Our hypothesis is that the β1t coefficients of significance in the whole sample (Table 4) become insignificant among households with toilets (Table 5) and in districts with more food market activity (Table 6). We conduct this difference-in-differences test using a split sample approach to gain maximum flexibility for the coefficients on each regressor. The test is powered by a large sample size achieved through merging two DHS rounds, and a large magnitude of the baseline β1 effects which could be brought to zero by sanitation and food markets. The results of Table 4 are similar using either the flexible specification in columns 1, 3, and 5, or the linear specifications in 2, 4 and 6. In both cases, higher NDVI in the second trimester of pregnancy is associated with greater attained heights, but only for boys. Girls have lower attained heights when NDVI is higher in the first three months after birth. Boys also have lower attained heights when NDVI is higher in months 3–5 after birth, but only in the less flexible functional form for age at measurement used in column 4. The magnitudes of effects are quite large. Coefficients shown in Table 4 are scaled per 1000 points of NDVI, so based on our preferred specifications in columns 3 and 5, each 100-point change in NDVI experienced by boys in midgestation is associated with an 0.088 difference in HAZ—almost as large as the 0.107 difference associated with each quintile of household wealth. For girls, each 100-point change in NDVI experienced during the first three months after birth is associated with a 0.054 difference in HAZ, which is almost half of the 0.112 difference associated with each quintile of wealth. Coefficients on wealth and other control variables are similar in direction and significance for both boys and girls, except that the magnitude of coefficients is somewhat larger for boys, suggesting greater susceptibility to variables such as altitude and maternal age. Turning to the effects of NDVI on households with and without toilets, from Table 5 we observe the predicted heterogeneity in effect sizes and significance. In households without toilets, fluctuations in NDVI significantly affect males only in utero (column 3) and significantly affect females only in the first three months after birth (column 5). Households with toilets are largely immune to the effects of NDVI fluctuations at any stage of early child development. Effect sizes for the subsample without toilets are larger than for the country as a whole: a 100-point change in NDVI during the second trimester of pregnancy affects boys’ heights as much as the equivalent of 1.2 quintiles of wealth, while a 100-point change in NDVI during the first three months after birth affects girls’ heights as much as the equivalent of 1.6 quintiles of wealth. However, those effects completely disappear in households with toilets. The effects of NDVI on children’s height in districts with higher or lower levels of food market activity (shown in Table 6) reveal a murkier picture. In the low-market-use districts, girls are susceptible to NDVI fluctuations only in their first three months after birth, and boys are susceptible to NDVI fluctuation in their second trimester of gestation, but these subsamples also show a significant correlation with NDVI during the period of complementary feeding (six to eight months of age). This particular correlation holds true for boys in districts with low food market use, and for girls in districts with high food market use. This could be a spurious correlation arising by chance in this sample, in part because these two periods in a child’s life occur in the same season in successive years, so the two NDVIs are highly collinear. Effect sizes for the main results are roughly similar here as in the previous table. For boys in districts with low food market participation, a 100-point change in NDVI in midgestation has an effect size similar to 1.5 quintiles of wealth; for girls, a 100-point change in NDVI in the three months after birth has a smaller effect, similar to about 0.5 quintiles of wealth. Robustness checks ~~~~~~~~~~~~~~~~~ To test the robustness of our results, we look for artifacts of the method by checking whether our regressions generate more statistically significant results than would arise by chance. These placebo regressions use the same data as our main results in Tables 5 and 6, with dependent variables that were determined long before the NDVI fluctuations occurred, so no causal effect is possible. These variables consist primarily of maternal characteristics (columns 1–6) plus household wealth and urban or rural location (columns 7 and 8). The base rate of entirely spurious placebo effects is proportional to p-values; for example, one tenth of the effects appears significant at a p-value of 10 percent. Table 7 reveals that, among the 7 × 8 = 56 distinct placebo “treatments” tested on each age-group, only five effects arise with a p-value < 0.1, four effects with a p-value < 0.05, and one with a p-value < 0.01. The number and pattern of significant coefficients in Table 7 correspond to the frequency of significant correlations that would be expected to occur by chance alone, and there is no time-related pattern to these correlations, as each occurs at a different stage of child development. Because our placebo regression tests used the same data and model structure as Tables 5 and 6, the clear pattern that we observe for NDVI treatment effects on child height is very likely to be a robust result of causal mechanisms.","The objective of this study is to identify the influence of early-life agroclimatic conditions on children’s attained heights in Nepal, and to test for the protective effects of household sanitation and food markets. In so doing, we find clear heterogeneity in vulnerability to changes in NDVI: boys are most affected by agroclimatic conditions during their second trimester of gestation, whereas girls are most vulnerable in the three months after birth. These findings are consistent with biomedical studies of sex-specific fetal development and socioeconomic studies of gender bias in child care. Both kinds of vulnerability are eliminated in households with toilets, and greatly reduced in districts that have more active use of food markets. The magnitude of fluctuations against which sanitation and food markets are protective is very large: on average, in households without toilets, for boys each 100-point variation in NDVI during the second trimester of gestation is associated with as much difference in attained height as conferred by 1.2 quintiles of household wealth, and for girls that same variation during the first three months after birth is associated with a difference in height equivalent to 1.6 quintiles of household wealth. The average seasonal change in NDVI from peak to trough in Nepal during the 2000–2011 period is approximately 200 points in the Mountains and Hills regions and 300 points in the Terai (lowland) region. The resulting difference in attained height between the most and least favorable times of year depends on location but is similar to the differences associated with about one to three quintiles of wealth. To test the internal validity of these findings, we conducted a number of placebo regressions that revealed no significant artifactual correlations that could have resulted from selection effects in birth timing. Heterogeneity in response to agroclimatic conditions provides important clues as to the causal mechanisms involved, and valuable guidance regarding the targeting of interventions. Our results are consistent with the premise that household sanitation protects children against fecal–oral transmission of disease, while food purchases made possible by more robust food markets protect children against fluctuations in local food production. Thus, policy interventions to promote sanitation and strengthen food markets could help protect children against adverse conditions in the future, helping them to grow up healthy despite the effects of climate change. Our research design exploited random exposure to varying agroclimatic circumstances in relation to birth timing. To detect selection effects, we test for the effects of socioeconomic status on month of birth and found no correlations. We also use a series of placebo regressions to test whether our research design generates spurious correlations, and find only the expected base rate of noncausal statistical significance. These empirical tests increase confidence in the robustness of our results for the Nepali context. The main limitation of our study concerns the protective effects of sanitation and food markets, as household toilets and market activity were not randomly assigned. Our regressions control for observable influences on child height, such as household wealth, parental education, maternal BMI, and altitude as well as district fixed effects. Violations of the parallel- trends assumption behind our difference-in-differences design would involve other factors that make households less vulnerable to agroclimatic conditions, such as cultural differences. For example, perhaps households with toilets and in districts with stronger food markets have access to more effective social insurance than the average household with similar levels of observable control variables. While the large protective effects found in our data are plausibly explained by differences in fecal–oral disease transmission and food consumption smoothing, cultural differences or other unobserved factors cannot be ruled out without conducting large-scale, randomized trials. Future work could extend our results using other natural experiments in observational data, perhaps over longer periods and more countries, or refining the measurement of agroclimatic conditions and the timing of exposure. This study confirms that exploiting natural experiments can be of great value in identifying the periods in utero and during infancy when child development is most at risk, and extends previous results to reveal the socioeconomic conditions that are most able to protect children against those risks. These results point to the power of investments in maternal health during pregnancy to protect boys, and in immediate postnatal care to protect girls, as well as the value of sanitation and food markets for building resilience against variation in agroclimatic conditions.","This project was supported by the United States Agency for International Development through the Feed the Future Innovation Lab for Nutrition [grant number AID-O-AA-1-1-00005, for work by WAM, PM and GES], and by the Bill & Melinda Gates Foundation through the International Food Policy Research Institute [project number 301052.001.001.515.01.01, for SAB and WAM]."],["We study the effects of credit shocks in a model with heterogeneous entrepreneurs, financing constraints, and a realistic firm-size distribution. As entrepreneurial firms can grow only slowly and rely heavily on retained earnings to expand the size of their business, we show that, by reducing entrepreneurial firm size and earnings, negative shocks have a very persistent effect on real activity. In determining the speed of recovery from an adverse economic shock, the most important factor is the extent to which the shock erodes entrepreneurial wealth. --------------------------------------------------------------------------------","The recent turmoil in financial markets has had deep consequences for the allocation of credit within the economy. Access to credit is particularly important for nascent and growing firms, for which it is much more difficult to rely only on retained earnings as a source of financing. In this paper, we study the effects of various types of financial shocks in a model with two nonfinancial sectors: a corporate sector, primarily composed of mature firms, and an entrepreneurial sector, whose leverage is limited by its inability to fully commit to repay debts. The constraints generate a large, and realistic, dispersion in firm size, and limit the rate at which entrepreneurial firms can grow. We build on the entrepreneurship model of Quadrini (2000) and Cagetti and De Nardi (2006, 2009), and introduce a financial intermediation sector that channels resources from savers to users of capital. Both entrepreneurs and corporate firms require access to intermediated funds. The reliance of entrepreneurs on intermediaries is one of the parameters used in our calibration and is most associated with matching the ratio of the wealth of entrepreneurs to that of workers. We calibrate directly the reliance of corporate firms on outside funding to match data from the flow of funds. Our main experiment considers the effects of an increase in the cost of channeling funds through intermediaries, which increases the cost of borrowing and, in general equilibrium, also depresses the rate of return earned by savers. This shock can be the result of either a negative productivity shock in the financial intermediation sector, or the destruction of capital specific to this sector (e.g., the loss in value of mortgage-backed securities). For the parameters that best match our target moments, we find that entrepreneurial firms are affected to a deeper extent than corporate firms. To the extent that entrepreneurial firms tend to be smaller, this is in line with the empirical findings of Gertler and Gilchrist (1994). When intermediation costs return to their steady-state levels, both entrepreneurs and corporate firms stage an initial rebound, but the path to a full recovery is then slow. The wealth accumulation of the entrepreneurs is affected in a very persistent way. Negative credit shocks reduce firm size, and, because entrepreneurial firms can grow only slowly, limit the speed at which firms return to their previous scale when the shocks subside. This slow transition is characterized by more capital misallocation and, hence, lower output than in steady state. A key prediction of our model is that the effects of adverse shocks are more persistent on small businesses. There is some evidence showing that this is the case. Credit flows to small and large businesses after the financial crisis have behaved in very different ways. According to the Financial Accounts of the Unites States, credit flows to both corporate and noncorporate firms (the latter mostly small businesses) dropped sharply during the financial crisis. However, credit flows to corporations resumed relatively early in 2010 and went back to healthy levels by 2011, as corporate firms could access various credit markets (such as the bond market). Credit to noncorporate businesses continued to decline through 2010 and started rising only slowly thereafter. The same pattern emerges markedly for the recession of 1990–91, and, to a lesser extent, for most recessions in the past 50 years, as documented in Fig. 1. In turn, credit availability and credit flows have an impact on firms' growth. Chodorow-Reich (2014) shows that firms that had banking relations with less healthy banks not only had more difficulty obtaining credit after the crisis, but saw larger declines in employment. Interestingly, this effect is larger for smaller firms. Similarly, according to surveys conducted by the National Federation of Independent Business (NFIB), the crisis had a long-lasting impact on small firms' investment plans. The net percentage of firms that planned capital outlays and that of firms that anticipated business expansions plummeted during the recession and has increased only slightly in recent years, a much smaller recovery than that seen in aggregate data on investment.1 A similar pattern appears in the previous two recessions. These indexes dropped sharply during the recession, but returned to their pre-recession levels more slowly than aggregate data on investment would suggest. Our model is also consistent with the differential behavior of employment at firms of different sizes during the latest recession and recovery, documented also by Siemer (2013). Employment at smaller firms fell more sharply than at larger firms during the recession, and, since then, the ratio of employment at small firms (relative to large firms) has remained depressed and hasn't returned yet to pre-recession levels. In our model, an increase in intermediation costs also generates an endogenous tightening of borrowing constraints, as entrepreneurial activity becomes less profitable and the outside option of absconding part of the capital becomes comparatively more attractive; this channel accounts for about 50% of the drop in entrepreneurial firm size. Government policy interacts with the financial disruption. We study two aspects of this interaction. First, the recession initiated by the financial shock creates a shortfall in the government budget. If income taxes are raised to finance this shortfall, they constitute a new, independent drain on entrepreneurial profits; this drain can be even bigger than the financial shock itself, leading to an even longer recovery. Second, we analyze the effects of a government-targeted intervention in financial markets that drives a wedge in the cost of funds across different classes of borrowers. Our experiment is closest in spirit to the U.S. Treasury's guarantee of money market mutual funds (and implicitly of the underlying commercial paper): We consider a case in which the government is able to completely and costlessly insulate the corporate sector from the shock.2 This guarantee is helpful in reducing the depth of the recession, but it does nothing to improve the recovery, as it concentrates the shock in the sector that is most vulnerable in the long run. We contrast the effects of our baseline shock with alternative scenarios, such as a collateral shock, that makes it harder for entrepreneurs to pledge future repayment of debt, similar to Jermann and Quadrini (2007), or a traditional shock to total factor productivity (TFP). We find that the response to these shocks may be quite similar, to the extent that the balance sheet of entrepreneurs is hit in a similar way; the evolution of this balance sheet is the key element that affects the speed of the recovery. In our set-up, all these shocks have a very persistent effect on real activity.","Many works incorporate credit-market frictions in macroeconomic models but, rather than studying the direct effect of shocks to these frictions, they focus on how these frictions affect aggregate investment and help generate and amplify business fluctuations. Among the earliest and most influential contributions, Bernanke and Gertler (1989) introduce agency problems such as costly state verification in a dynamic general equilibrium set-up, and Kiyotaki and Moore (1997) further illustrate the impact of collateral constraints and their interaction with asset prices and firms' net worth. In both papers, credit imperfections link investment decisions to the firms' balance sheets and generate a “financial accelerator” that amplifies and propagates shocks to the macroeconomy. The recent financial crisis has given further impetus to this literature, highlighting both the many channels through which credit market imperfections can affect real activity and the possible effects of government interventions to improve the functioning of credit markets and the flow of funds between borrowers and lenders. For a review of this literature, see Bernanke et al. (1999) for earlier contributions and Gertler and Kiyotaki (2010), Brunnermeier and Sannikov (2014), and Krishnamurthy (2010) for more recent ones. Here, we only mention a few of the papers most related to our work. We model several types of financial frictions. Financial intermediation (and more in general frictions in credit markets) introduces a wedge between the returns to lenders and the cost of capital to borrowers, a wedge related to the spread between liquid and easily intermediated securities such as Treasuries and corporate bonds. These credit spreads vary over time and their level and variation have been shown to be empirically correlated to and potentially key to understanding output fluctuations (for instance, Gilchrist et al., 2010; Christiano et al., 2010; Adrian and Shin, 2010). Their role has been highlighted, among others, by Hall (2011), who shows that in a simple representative-agent economy, credit spreads (including those for households) are powerful determinants of economic activity and can generate fluctuations of the magnitude of those seen in the recent crisis, and by Curdia and Woodford (2010), who study how monetary policy rules should respond to shocks to credit spreads. We also find that spreads have a significant impact on aggregate output during a credit crisis; by themselves, spreads have a fairly short-lived effect in our model economy. It is a different source of frictions that propagates the effect of spreads and generates a very persistent drop in output. Among borrowers, we explicitly distinguish corporate and entrepreneurial firms, since these two types of firms react to shocks in different ways in the data, and credit disruptions are likely to have a differential impact on the two (see, for example, Quadrini, 1999). We model credit frictions to entrepreneurs as endogenous borrowing constraints arising from imperfect enforceability of debt contracts (as in Kehoe and Levine, 1993; Alvarez and Jermann, 2000). In this set-up, credit availability to entrepreneurs depends on their balance sheet and their available collateral. The presence of limited commitment slows the growth of nascent firms and links it to the entrepreneurs' cash flow. It is this channel that propagates the initial financial shock in our model and is responsible for our main results. Our paper is thus also closely related to Khan and Thomas (2013), who examine the effect of capital misallocation that results from a collateral requirement shock in a real business-cycle model with heterogeneous firms and capital rigidities. Compared to their work, we focus more closely on small entrepreneurs that own their firm and thus face a consolidated resource constraint to finance investment and consumption. For these entrepreneurs, rebuilding the firm's balance sheet comes directly at the expense of their own consumption, which slows down the adjustment more than in the case of a corporate firm owned by diversified shareholders. The differential impact of credit frictions on businesses of different sizes has been shown to be useful in explaining a variety of phenomena, for instance firm-size distribution (Athreya and Akyol, 2009; Monge, 2009), firm dynamics (Albuquerque and Hopenhayn, 2004), macroeconomic fluctuations (Cooley et al., 2004; Jermann and Quadrini, 2012), and growth (Buera and Shin, 2013). More broadly, the interaction between frictions, entrepreneurship, and inequality is crucial to understanding the response to macroeconomic shocks (Jermann and Quadrini, 2007), the effect of certain government policies (Cagetti and De Nardi, 2009; Meh, 2005; Kitao, 2008), and asset pricing (Heaton and Lucas, 2000; Roussanov, 2010; Covas and Fujita, 2011). Shourideh and Zetlin-Jones (2012) report that as much as 90% of investment by private businesses is financed by outside funding. In our economy too, a substantial fraction of small-business investment is undertaken with borrowed funds, although this number is between 50% and 75%. Compared to their work, we have different definitions of businesses: We look at self-employed entrepreneurs who actively manage their business, whereas they focus on the universe of private businesses. For our class of entrepreneurs, the distinction between personal assets and business assets is tenuous as, for example, private assets are frequently used as collateral for the business. Credit ~~~~~~ External financing to both entrepreneurs and non-entrepreneurial firms is provided by competitive financial intermediaries. The intermediaries borrow funds from workers (and possibly entrepreneurs, though in equilibrium almost all entrepreneurs will be credit constrained and will invest all their wealth in their own firm). The entrepreneurial demand for borrowed funds arises endogenously in the model. As in Kehoe and Levine (1993), entrepreneurs are subject to borrowing constraints that are endogenously determined in equilibrium and stem from the assumptions that contracts are imperfectly enforceable. Government and taxation ~~~~~~~~~~~~~~~~~~~~~~~ We model progressive taxation of total income as in Cagetti and De Nardi (2009) and use their parameter estimates. As a first pass, we abstract from the tax implications of corporate finance decisions by assuming that corporate income taxes are zero and that capital gains are taxed as regular income.4 Households ~~~~~~~~~~ An old entrepreneur who is still able to run a business can decide to keep the activity going or retire, while a retiree cannot start a new entrepreneurial activity.","In this section, we describe the parameters taken from the literature or estimated outside of the model (Table 1), and the moments we use to calibrate the remaining parameters (Tables 2 and 3). Non-calibrated parameters ~~~~~~~~~~~~~~~~~~~~~~~~~ The coefficient of relative risk aversion σ, the capital share in the non-entrepreneurial Cobb–Douglas production function α, and the depreciation rate δ are set to values commonly used in the literature (for instance, respectively, Attanasio et al., 1999; Stokey and Rebelo, 1995; Gollin, 2002). The probabilities of aging and of dying are such that the average length of working life is 45 years and that of retirement is 11 years. We assume that the logarithm of the workers' income is an AR(1) process and approximate it as a 5-point Markov chain using the method in Tauchen and Hussey (1991). The autocorrelation coefficient and the variance of the error term for the AR(1) process are chosen to obtain a correlation coefficient of 0.95 (among others, Lillard and Willis, 1978) and a Gini coefficient for earnings of 0.38 (Huggett, 1996; De Nardi, 2004). The Social Security replacement rate is 40% of average gross income (see Kotlikoff et al., 1999). The steady- state ratio of government spending to GDP is set to 18.7%, and the tax rate on consumption is 11%. All of these parameter choices are discussed in Cagetti and De Nardi (2009). We set the steady-state financial intermediation cost to obtain a 1.5% spread between the interest rate paid by borrowers and that received by lenders. This is calibrated to the historical average of the spread between Baa-rated companies and Treasuries. In our model, both public and private debt are risk free, and the spread is entirely due to the special liquidity role of Treasuries, which are assumed not to require any intermediation. For this reason, we choose to match our private borrowing rate to an empirical counterpart that features low default risk but is also unlikely to carry any liquidity premium (see Krishnamurthy and Vissing-Jorgensen, 2012, for more discussion). As a comparison with other securities, the average spread between AAA-rated corporate bonds and Treasuries is about 1.25% and that between BBB-rated bonds and Treasuries is about 2.2%. The parameter ξ, the constraint on debt financing, is the average ratio of total corporate debt to the value of corporate tangible assets, a ratio equal to about 0.34 in the Flow of Funds Accounts. Corporate debt includes commercial paper, corporate bonds, mortgages, and other loans; tangible assets include equipment and software (at replacement cost), structures (at market value), and inventories. The ratio has generally been increasing since the beginning of the data in 1950, from below 0.25 during the 1950's to values near or above 0.5 in recent years. This ratio jumped to 0.58 immediately after the financial crisis, as the crash in commercial real estate prices sharply reduced the value of tangible assets, but since 2009 the ratio has been trending down toward pre-recession norms. For the other parameters, we take a ratio of government expenditures to GDP of 18.7% (NIPA data), a consumption tax of about 11% (Altig et al., 2001), and a level of government debt that, given the equilibrium interest rate, yields an average ratio of total interest payments to GDP of 3% (Altig et al., 2001). Calibration targets ~~~~~~~~~~~~~~~~~~~ In previous work, Cagetti and De Nardi (2009) have discussed the relevant empirical counterpart to the concept of entrepreneur in this model. Our entrepreneurs are the self- employed business owners that actively manage their own firm(s). We identify them in the SCF as those that declare that they are self-employed, that they own a business, and that they actively manage it. In total, we calibrate nine parameters. We use the first seven parameters to target the following moments: the capital–output ratio, the fraction of entrepreneurs in the population, the fraction of entrepreneurs exiting entrepreneurship during each period, the fraction of workers becoming entrepreneurs during each period,7 the ratio of median net worth of entrepreneurs to that of workers, the fraction of people with zero wealth, and the fraction of entrepreneurs hiring workers on the labor market. We choose the other two parameters to match the revenue from estate and gift taxes and the fraction of the estates that pay estate taxes. Table 2 reports the target values from the data and the values generated from our model; Table 3 reports the parameter values used in our calibration. For the capital–output ratio, we use the Federal Reserve Board Flow of Funds Accounts. We define capital as tangible assets excluding consumer durables, federal, state, and local government assets. The measure thus includes equipment and software, structures, both residential and nonresidential, and inventories. Equipment and software is measured at replacement cost; structures are measured at market value (except nonresidential structures owned by the financial sector, for which market value information is not recorded). With this definition, the ratio for the available years (1960–2009) is 2.96. Excluding the data after 2000, years that experienced first a large increase and then a drop in house values, the ratio is slightly smaller, about 2.89. The fraction of entrepreneurs in the population, the probability of entering and exiting entrepreneurship, the ratio of the median wealth of an entrepreneur to that of a non- entrepreneur, and the fraction of entrepreneurs hiring workers are computed from the 1989 Survey of Consumer Finances (other waves give similar results). We compute the transition matrix between entrepreneurship and non-entrepreneurship by looking at households that are present in two consecutive surveys. The fraction of the population at zero wealth, also computed from the SCF, is somewhat sensitive to the exact cutoff point (whether exactly zero, or some positive but small amount such as $100). This fraction varies from roughly 7% to 13%. The percentage of entrepreneurs hiring workers (besides themselves and, possibly, their spouse) is also computed from the SCF. We do not use the statutory exemption and tax schedule to model the estate tax. As explained in Cagetti and De Nardi (2009) and in the references therein, the effective tax rate can differ substantially from the statutory one. We thus calibrate the exemption level and the (flat) tax rate above the exemption to match the percentage of estates that pay an estate tax (2%) and the total amount of revenues of estate and gift taxes (about 0.2–0.3% of GDP).","As shown in Cagetti and De Nardi (2006, 2009), our model of entrepreneurship, although simple, matches very well the wealth distributions of both entrepreneurs and workers. In the presence of borrowing constraints, this is very important to determine the response to financial shocks for both the whole distribution of entrepreneurs and for the important macroeconomic aggregates. Figs. 2 and 3 compare the distribution of net worth for workers and entrepreneurs generated by the model and in the actual SCF data and confirm that the model generates the long upper tail of the wealth distribution that is observed in the data for the whole population, as well as the large wealth holdings concentrated in the hands of a few entrepreneurs. In our calibration, the entrepreneurial sector employs about 57% of the capital, a little over the value reported by the Small Business Administration (which is about 50%).8 It also employs 33% of the efficiency units of labor in the economy; in the data, the Small Business Administration reports that the entrepreneurial sector employs almost 50% of the workers in the economy; in the data, larger (corporate) firms tend to pay more, which helps in closing this gap. Table 4 displays the distribution of labor hiring by entrepreneurial firms. The first line is computed from the 2007 SCF data, which is the last survey year before the crisis (the numbers from previous years are very similar). The question asked in the SCF is how many workers the entrepreneurs hire in their firm (we exclude the entrepreneur and his or her spouse). To compare it with the model, we assume that the average employee of the entrepreneurial firm (up to the 95% quantile) lies at either the 33rd percentile (second line) or the median of the efficiency distribution (third line) and that he works full time. Given that some of the employees in the SCF data will be part time, we conclude that the distribution of hiring by entrepreneurial firms in the model matches the one in the data reasonably well. In the data, entrepreneurs are much richer than workers, and their saving rate does not quickly decline with wealth. To match these facts, the calibration implies borrowing limits that are tight compared with the optimal firm size, hence the growth process of entrepreneurial firms is slow. This plays an important role for the response of the economy to various shocks, to which we now turn.","Throughout the experiments below, we start the economy at the deterministic steady state in period 1. We trace the response to a temporary change in the technology or credit parameters (a “shock”) that lasts for three years, from period 2 to period 4. This path is known by all agents at the beginning of period 2, before any decision is made. Negative technology shock in the intermediation sector ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We consider the effect of a shock that increases ϕ from 1.5% to 3.5% for three years.10 This is a way of capturing either of two alternative shocks: More monitoring is necessary to ensure loan performance due to the financial turmoil. ϕ stands in as payments to a factor that is fixed in the short run and that is temporarily depleted. As an example, suppose that banks face capital requirements and that some initial losses wipe some of the capital out, constraining the banks' ability to offer additional intermediation services. In this case, the increase in ϕ would reflect the additional reward for the scarcer banking capital.11 Fig. 5 shows the effect of the financial intermediation shock on the number of entrepreneurs and their average firm size. The fraction of entrepreneurs drops, particularly when taxes adjust, but this margin is not very persistent: when intermediation costs and taxes are back to normal, entrepreneurs quickly reenter the market. This is because the minimum firm size that makes entry profitable is small, and potential entrepreneurs can save to reach that point quickly. Average firm size also drops; this effect is bigger and much more persistent. The intermediation cost reduces the entrepreneurs' cash flow and their ability to retain earnings to foster their business' growth. Since both the wealth distribution and the distribution of assets across firms that we match are very spread out, our model implies a very gradual growth of firms, with almost no entrepreneur attaining sufficient wealth that borrowing constraints cease to bind. It follows that any negative shock has almost a permanent effect on each entrepreneur and its aggregate impact vanishes fully only when each entrepreneur loses his ability and closes the firm. As soon as the shock is over, firm growth resumes, but at a slow pace dictated by the tight borrowing limits. The alternative ways in which taxes, spending, and interest rates adjust across the three experiments reveal important differences. Consider first the cases in which taxes are held fixed, and government spending acts as a residual. Fig. 6 plots the borrowing and lending rates for the case in which the lending rate is held fixed (a small open economy) and that in which the capital market clears. During the periods of the shock, borrowing rates spike higher when savers have the opportunity of earning a fixed rate abroad. As a consequence, in Fig. 5, average firm size drops more when lending rates are held constant (solid line) than when the effect of the shock is spread between borrowers and savers, as in general equilibrium (dashed line). The difference between partial and general equilibrium reverses after the shock is over. The shock triggers a reduction in aggregate capital; in general equilibrium, the resulting higher interest rates impair the entrepreneurs' ability to rebuild their balance sheet and lead to a slower recovery in firm size. The differences between the solid and dashed lines in Fig. 5 are minor compared with the differences between either of those lines and the dotted line, which represents the case in which the government balances its budget by increasing taxes rather than cutting government spending. To balance the budget, the government needs an increase in the tax rate of about 1.5% for nine years. The government imbalance does not have a large impact on the depth of the initial recession, but it causes a prolonged slump once the fiscal adjustment takes place. In our model, taxes work very similarly to the intermediation shock that is the original disturbance: Both changes deprive entrepreneurs of the cash flow which is the source of growth for their firms. As a consequence, when taxes increase, the average firm size of entrepreneurs suffers from a gradual contraction whose cumulative effect is several times larger than the original, shorter-lived shock. The gradual erosion of entrepreneurial wealth has a corresponding effect on GDP, both because it depresses capital accumulation (through lower interest rates) and because it shifts capital and labor towards less productive uses in the corporate sector. In the short run, investment takes a greater hit than consumption. However, as the period of high taxes drags on, a smaller firm size implies greater marginal product of capital, which acts as a countervailing force and prevents further falls in investment. In contrast, the impact on consumption, which is initially more muted, compounds over time, so that, by the end of the period of high taxes, consumption and investment have dropped approximately by the same amount relative to their steady-state level. The recovery from the double shock of an increase in the intermediation cost and the subsequent response in taxes is delayed compared with what happens when government spending is cut, and it starts from a weaker position. Fig. 7 compares value added in the entrepreneurial vs. the corporate sector of the economy. Since both sectors use capital intermediated by the financial sector, their value added drops when the cost of accessing the intermediaries' services increases. For our calibration, entrepreneurs are more reliant on financial intermediation, and the drop in their value added is twice as large as that for the corporate sector. In period five the borrowing-cost shock is over and the effect reverses. In this period, the value added in the entrepreneurial sector grows faster than in the corporate sector; nonetheless, entrepreneurs do not recover fully, while the corporate sector stages a full recovery. The difference across the two sectors is due to the nature of the credit frictions faced by the two types of firms. Entrepreneurs are primarily constrained by their net worth, which can only be rebuilt slowly, whereas corporate firms curtail their investment only because of the additional cost of borrowing in Eq. (8), a period-by-period cost that returns almost to normal as soon as intermediation costs revert to their steady-state level. From period five, the two economies in Fig. 7 diverge. When (wasteful) government spending acts as the residual (left panel), no further shock perturbs the entrepreneurs' wealth accumulation, and the economy immediately starts on a path of slow convergence back to the steady state. In contrast, when taxes adjust, their effect on the entrepreneurs' balance sheet tightens the constraint on entrepreneurial firm size and results in a reallocation of resources from entrepreneurs to corporate firms. Due to computational limitations related to the endogenous borrowing constraints, our model features an inelastic labor supply, and thus it cannot capture the decline in labor occurring during the downturn. However, we can analyze the relative allocation of labor across the two sectors, which we show in Fig. 8. This picture mirrors what we observed for output: The recession caused by the financial shock shrinks the share of employment at entrepreneurial firms, in line with Gertler and Gilchrist (1994), who observe that small firms are more sensitive to the business cycle. Our results are also in line with what Moscarini and Postel-Vinay (2012) find, that small firms grow comparatively faster early in the recovery; this is particularly clear in the left panel of Fig. 8, where the prolonged slump due to the increase in taxes is absent. The comparatively fast growth in employment at entrepreneurial firms is partially compensating for the losses occurred during the recession, but it is not enough to close the gap in levels: In our model, employment in the sector remains depressed for a long time.12 Our results about relative employment levels mimic the Business Employment Dynamics data on employment by firm size published by the Bureau of Labor Statistics and the results by Siemer (2013). The dataset does not identify entrepreneurial firms per se, but, as a proxy, we can look at firms by number of employees. Fig. 9 shows that the ratio of employment at firms with fewer than a certain number of employees relative to employment at firms with more than the same threshold number fell sharply during the 2007–2008 recession, indicating that small firms shrank more than larger ones. In addition, after the initial fall, the ratio has moved up only slowly, and remains below its pre-recession level even several years into the recovery, indicating that the effect of the shock is particularly persistent for small firms. Having analyzed the forces that drive the behavior of our economy, we now turn to their aggregate implications. Fig. 10 plots aggregate GDP. The increase in intermediation costs (an intermediate input in our economy) depresses TFP and output during the financial shock. This is particularly true when lending rates are fixed, since in this case capital moves out of the economy.13 In general equilibrium, the output drop on impact is mostly driven by the TFP effect of the shock: The drop in output is close to 3%, with 0.3% being due to the misallocation of factors. As the net worth of entrepreneurs is eroded, the misallocation becomes a more prominent force; after the shock is over, the entire difference between the solid line and the steady state is due to this misallocation, whereas in general equilibrium (dashed and dotted line) the decrease in capital accumulation plays a role. When taxes hit the entrepreneurs' ability to accumulate wealth and grow their own business, the economy fares much worse. This second dip would be less pronounced with a longer transition period over which taxes are raised; but, of course, in this case the policy response to the shock would have even more persistent effects. Fig. 11 plots aggregate consumption and investment.14 As in most business-cycle models, investment bears the brunt of the shock at first. Nonetheless, consumption drops too. Unlike a pure tightening of borrowing constraints, a shock to financial intermediation entails real output costs that reduce total available resources from the outset. In general equilibrium, the behavior of aggregate investment contributes to a slow recovery: After the initial drop in the periods of the shock, investment never overshoots its steady-state level. When government spending adjusts, investment merely returns close to steady state; when taxes further depress wealth accumulation, investment remains 2% below its steady state several years after taxes have returned to their steady-state level. It has often been remarked that, during the crisis, credit standards were extremely tight and businesses found it difficult to access credit at any price (see, for example, the Quarterly Senior Loan Officer Opinion Survey conducted by the Federal Reserve Board). A similar effect arises in our model: The increase in intermediation costs is accompanied by an endogenous tightening of the constraints, because entrepreneurship becomes less profitable and thus the temptation of terminating the business and absconding with part of the capital becomes stronger. To illustrate this mechanism, we run an alternative experiment, where intermediation costs increase by the same amount and duration (2% for three years), but borrowing constraints are held fixed exogenously at their steady-state values. Fig. 12 compares the effect of the intermediation shock for exogenous and endogenous borrowing constraints with government spending acting as a residual. It shows that the adjustments of the extensive margin are almost exclusively driven by the tightening of the borrowing limits, which forces potential entrepreneurs to accumulate more wealth before entry becomes worthwhile. The profit loss from the intermediation shock leads to smaller firms even with fixed borrowing constraints, but the drop in average firm size is about half as large. After the shock, the speed of recovery is similar, but the level from which the economy has to recover is much lower with endogenous borrowing constraints, so that it takes longer to return to the same level of output. The behavior of GDP with endogenous vs. fixed borrowing limits (Fig. 13) mirrors the one for the average size of firms, but quantitatively the tightening of borrowing constraints has a more muted impact in the periods of the shock. When the shock is active, both entrepreneurs and corporate firms are subject to it, and holding borrowing limits fixed only benefits entrepreneurs. After the shock is over, the gap between the two lines widens, because the persistent effect of the shock is dictated by the evolution of entrepreneurial wealth, which is less severely impacted when borrowing constraints are held fixed. Figs. 14 and 15 compare the behavior of entrepreneurial firms and GDP with endogenous vs. fixed borrowing constraints in the case in which the government raises taxes to balance its budget. The differences are starker here, because an increase in taxes drains the profitability of entrepreneurs and generates its own credit crunch if borrowing constraints are allowed to adjust endogenously. Negative technology shock in the intermediation sector, only for entrepreneurs ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the wake of the financial crisis of 2008, the government took several actions aimed at restoring calm in several financial markets. Among the actions that were most successful ex-post was a blanket guarantee of money market mutual funds, and thus, indirectly, of the commercial paper of corporate industrial firms that those funds purchased. More generally, companies with direct access to markets seemed better able to cope than those that were forced to go through the banking sector.15 Fig. 16 studies the differences between this experiment and the case of a pure shock to ϕ for the entrepreneurial sector, in the case of general equilibrium and fixed taxes (with government spending adjusting as a residual). Since corporate firms are insulated from the shock in this new experiment, entrepreneurs face stiffer competition in the factor markets, which thins their ranks and leads them to shrink their firm size more. The recovery is affected by two opposite forces. The greater hit taken by entrepreneurs slows the return to the steady state. However, aggregate investment (Fig. 17) drops less when only one sector is hit by the shock, and the additional capital is beneficial to the recovery. The first force dominates in the short run, but about five years after the shock the two experiments become quite similar. Fig. 18 displays the value added in the two sectors in response to the shock to ϕ and the contemporaneous offset through ξ. When the corporate sector is completely insulated, its size actually expands during the financial disruption, as it poaches workers and capital from the entrepreneurs. Fig. 19 shows how all of these effects combine to determine aggregate GDP. Even in the best-case scenario, in which the government intervention entails no cost, it is successful at reducing the severity of the recession but it has almost no impact on the recovery. By helping the corporate sector, the government exacerbates the misallocation of resources due to financial frictions. A shock to required collateral ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Here, we consider a shock that increases the collateral that the entrepreneurs need to secure their loans. Specifically, we raise the fraction of capital with which entrepreneurs can abscond (f) from 75% to 80%. We calibrate this shock to have an effect on aggregate output during the credit crunch that is of similar magnitude of the drop that we obtained considering a shock to ϕ for the cases of general equilibrium. This can be seen in Fig. 20.17 Since this shock only affects the entrepreneurial sector, matching the output drop on impact requires a much deeper contraction in the number of entrepreneurs and firm size when f increases than the baseline case in which borrowing costs increase for both entrepreneurs and corporate firms. This can be seen in Fig. 21. It might seem surprising that the deeper contraction in entrepreneurial firms does not bear bigger implications for the entrepreneurs' wealth in the recovery phase. The reason for this result is that a shock to f hits only the marginal profits of the firm: It forces entrepreneurs to shrink their scale, but it has no effect on their profits for a given scale of operations. In contrast, an increase in ϕ raises the rental rate of capital paid by entrepreneurs. This effect applies to all of the capital that they rent and has a negative effect on their profits even conditioning on their scale of operations. A TFP shock ~~~~~~~~~~~ We finally contrast a credit shock to a TFP shock that hits both the corporate sector and the entrepreneurial sector. In this case, TFP drops by 2.5% for three years, and subsequently reverts to steady state. Once again, the magnitude of the TFP drop is chosen so as to obtain a similar GDP drop on impact in general equilibrium. As Figs. 22 and 23 show, the evolution of the economy under this shock is fairly similar to that of a shock to ϕ.","From the experiments that we ran, we learn three lessons. First and foremost, we find that it is not the source of the disturbance that determines our economy's speed of recovery, but rather the way in which the shock affects the profitability of credit-constrained entrepreneurs. Recovery is slowest in the case of an increase in borrowing rates from which the corporate sector is shielded (our experiment of Section 6.2). Among our experiments, this one has the shallowest recession; and yet during the recovery output is at a similar level as that of the others, in which the economy needs to make up for deeper drops. When losses are concentrated in the entrepreneurial sector, it takes more time for entrepreneurs to rebuild their balance sheet. Second, the way public finances adjust in response to the shortfalls caused by a recession is important. Income taxes are a further drain on the cash flow available for successful business owners to grow and represent a further significant drag on the economy. From an efficiency perspective, entrepreneurship subsidies would contribute to increasing output. It should be noted that this does not necessarily imply that subsidizing entrepreneurs is an optimal policy. Even if it were easy to identify the exact counterpart to credit-constrained, highly productive entrepreneurs, this policy would require taxing workers, who are on average far poorer in our economy, as in the data, to subsidize comparatively richer business owners, raising equity considerations. Finally, in an environment with endogenous borrowing constraints, financial shocks that increase interest costs have two effects. The interest rate increase represents a direct drain on firms' profits. The indirect effect is that higher borrowing rates trigger a tightening of credit limits. Hence, for a given contraction in credit, financial shocks that affect borrowing rates have potentially more severe implications than pure credit rationing."],["We study vector autoregressions that impose equality and/or inequality restrictions to set-identify the dynamic responses to a single structural shock. We make three contributions. First, we present an algorithm to compute the largest and smallest value that an impulse-response coefficient can attain over its identified set. Second, we provide conditions under which these largest and smallest values are directionally differentiable functions of the model's reduced-form parameters. Third, we propose a delta-method approach to conduct inference about the structural impulse-response coefficients. We use our results to assess the effects of the announcement of the Quantitative Easing program in August 2010. --------------------------------------------------------------------------------","An increasingly popular practice in empirical macroeconomics is to set-identify the parameters of a structural vector autoregression [SVAR] by means of exclusion and/or sign restrictions. Most studies working with this type of models have relied on Bayesian methods to construct posterior credible sets for the structural parameters of interest (for example, Inoue and Kilian, 2013; Arias et al., 2017; Baumeister and Hamilton 2015). A practical concern with Bayesian analysis in set-identified SVARs is that posterior inference continues to be influenced by prior beliefs even if the sample size is infinite (Poirier, 1998; Gustafson, 2009; Moon and Schorfheide, 2012). This observation has motivated the study of alternative approaches to inference that dispense with the specification of a prior distribution over structural parameters that are only set- identified. There are two existing proposals that characterize the estimation uncertainty of set-identified structural responses, without postulating a specific prior for the parameters of the structural model. On the one hand, Granziera et al. (2017) [GMS17] have proposed a frequentist confidence interval for structural impulse-response coefficients based on a moment-inequality-minimum-distance framework. On the other hand, Giacomini and Kitagawa (2015) [GK15] have proposed a robust Bayes credible interval that achieves a given credibility level regardless of the prior specified over the model’s set-identified structural parameters. We contribute to the analysis of set-identified SVARs by proposing a novel delta-method interval for the coefficients of the impulse-response function [IRF]. We show that our delta-method interval is point-wise consistent in level and, under certain regularity conditions, has asymptotic robust Bayesian credibility of at least the nominal level. Thus, our inference approach can be interpreted both from a frequentist and a robust Bayes perspective. We also argue that the computational cost of our procedure compares favorably with GMS17 and GK15. Broadly speaking, our approach is based on a closed-form characterization of the endpoints of the identified set and their directional derivatives. Our delta-method interval – which may be viewed as a generalization of the pioneering work of Lütkepohl (1990) on delta-method inference for point-identified VARs – takes the form of a plug-in estimator for the identified set plus/minus standard errors. The main limitation of our approach is that the delta-method interval is only defined for SVAR models that impose equality and inequality restrictions on a single structural shock (e.g., a monetary policy shock). Admittedly, this is problematic, as some popular applications of set-identified SVARs feature restrictions on multiple structural innovations.1 In spite of this observation, single-shock set-identified models have been applied in several empirical studies: for example, to study the effects of monetary policy on output (Uhlig, 2005), the impact of monetary policy on the housing market (Vargas- Silva, 2008), the effects of labor market shocks on worker flows (Fujita, 2011), the effects of exchange rates on aggregate prices (An and Wang, 2012), and the effect of optimism shocks on business cycles fluctuations (Beaudry et al., 2011). Thus, we think there is room for our results to have an impact on empirical work. To illustrate the usefulness of our main results, we estimate a monetary structural vector autoregression using monthly U.S. data from July 1979 to December 2007 (a sample that deliberately ends a half-year before the financial crisis begins). The goal of our exercise is to use pre- crisis data to learn about the responses of macroeconomic variables to shocks that have effects similar to the ‘unconventional’ monetary policy interventions implemented after the crisis. We set-identify an unconventional monetary policy [UMP] shock as an innovation that decreases the two-year government bond rate upon impact, but has no effect over the nominal federal funds rate.2 We consider two additional sign restrictions on the contemporaneous responses of inflation and output. Namely, we assume that – upon impact – neither inflation nor output can respond negatively to a UMP shock. Since the model is only set-identified, our analysis effectively captures the effects of any historical economic shock that affected the economy in the same way as an UMP shock. We apply our delta-method approach to construct a confidence interval for the dynamic responses of industrial production, inflation, the two-year government bond rate, and the nominal federal funds rate. We use our delta-method intervals to assess the effects of the announcement of the second part of the so-called Quantitative Easing program (QE2) in August 2010. Pre-crisis data turns out to be extremely useful to learn about the post- crisis response of macroeconomic aggregates to unconventional monetary policy. The remainder of the paper is organized as follows. Section 2 presents an overview of the main methodological results in this paper. Section 3 introduces our empirical application, which is used as a running example throughout the paper. Section 4.1 presents our algorithm to evaluate the endpoints of the identified set. Section 4.2 establishes the differentiability properties of the endpoints. Section 4.3 presents our delta-method approach and establishes its asymptotic frequentist validity as well as its asymptotic robust Bayesian credibility. Section 5 presents the delta-method intervals for the dynamic responses to the QE2 program. Section 6 concludes. All of our proofs are collected in Appendix A. Additional figures and implementation details of different procedures are collected in Appendix B.","This section presents the baseline SVAR model, discusses the class of set-identifying restrictions that we consider, and provides an overview of our main methodological results. Set-identifying restrictions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A common practice in empirical macroeconomics is to use equality and inequality restrictions to set-identify the structural IRFs in (2.2). An example of an equality restriction in a monetary VAR is that prices do not react contemporaneously to monetary policy shocks. An example of an inequality restriction is that a contractionary monetary policy shock cannot increase prices. The simple formulation in (2.4) allows the researcher to incorporate the following identifying restrictions: Overview of the main results ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our delta-method approach is supported by the three results described in the abstract, which can be summarized as follows:","This section introduces our empirical application, which will be used as a running example to illustrate our assumptions and results. Baumeister and Benati (2013) study a related identification scheme. They consider a Bayesian SVAR to study an analogous ‘spread’ monetary policy shock that leaves the short-term nominal rate unchanged, but affects the spread between the ten-year Treasury-bond yield and the policy rate. Outline for the rest of our paper: We have already presented an overview of our main results and described our running example. In the remaining part of the paper, we formalize Theorems 1–3 and use them to conduct inference about the responses to an unconventional monetary policy shock. Assumptions We make two assumptions on the sign and zero restrictions allowed in the model: This assumption simply requires that the identifying restrictions do not contradict each other. Using the algorithm in the UMP example There are at least two other ways of evaluating the maximum and minimum response (although only our algorithm is guaranteed to provide a global solution in a finite number of steps). One approach is to simply use a numerical solver (such as Matlab’s fmincon) to get the value of the non-linear, non-convex program in (2.5). The result in Theorem 1 allows us to avoid the specification of the standard tuning parameters for numerical optimization routines (such as initial conditions, algorithms for the solver, tolerance levels for the solutions, and number of iterations). Assumptions In order to establish our differentiability result we need an additional regularity condition. Our key assumption is as follows: Directional differentiability We now state the definition of directional differentiability and present our second theorem. Theorem 3 ~~~~~~~~~ We now describe the main large-sample assumptions used to establish the frequentist coverage and the robust Bayesian credibility of our delta-method interval. Assumptions Asymptotic Normality of μ̂T Bernstein–von Mises Theorem","In August 2010 the Federal Open Market Committee announced: “The Committee will keep constant the Federal Reserve’s holdings of securities at their current level by reinvesting principal payments from agency debt and agency mortgage-backed securities in longer-term Treasury securities.” This announcement was an important prelude for the second part of the Quantitative Easing program (QE2) (see p. 244 in Krishnamurthy and Vissing-Jorgensen (2011) for a detailed discussion). In addition, this announcement generated a drop in the intraday yield for two- and ten-year treasury bond. In fact, from the end of July 2010 to the end of August 2010 the 2 year Treasury bond rate fell by 10 basis points. Fig. 4 uses our delta-method approach to construct confidence bands for the evolution of the levels of the four variables in the monetary SVAR. We fix all the variables at their level on July 2010 and we trace their evolution (over a 12-month window) according to the confidence set for their cumulative responses. The motivation for this exercise is as follows. Suppose that – back in August 2010 – an econometrician is asked to provide confidence bands for the evolution of IP, CPI, 2YTB, and FF after the August 2010 announcement of the Federal Open Market Committee (FOMC). The econometrician observes the realization of the macroeconomic variables from July 1979 until August 2010, but decides to deliberately ignore the two years of data after the crisis (to avoid introducing structural changes, stochastic volatility, or any other feature that will complicate the estimation of the VAR). The econometrician uses the data until December 2007 – one semester before the financial crisis – to conduct delta-method inference on the cumulative responses to a one standard deviation unconventional monetary policy shock. The econometrician then uses these cumulative responses to get a rough idea of the evolution of the variables (in levels) following the announcement of the Federal Reserve in August 2010. The econometrician assumes there is a linear trend for CPI/IP, and ignores sampling uncertainty coming from the trend estimation in reporting the bands. An ex-post evaluation of this exercise (over a window of 12 months) is reported in Fig. 4.17 We note that the observed dynamics for CPI, IP, GS2, and FFR from August 2010 to July 2011 fall within the bounds motivated by our delta-method interval. We also note that our delta-method interval misses the observed value at most three out of 12 months, which means that our 68% confidence set covers each of these variables at least 75% of the time. We also report the 68% Bayesian credible sets. Computational Cost: We close this section with some comments regarding the computational cost of our delta-method procedure. Most of the work to compute the endpoints of the identified set and its derivatives is analytical. Consequently, practitioners can expect the computational burden of our procedure to be low. We note that the implementation of our delta-method interval in the running example takes only around .15 s (using a standard Laptop @2.4 GHz IntelCore i7). Comparison with the Projection Approach: Fig. 6 in Appendix B.1 presents a comparison between the delta- method approach and the projection approach recently proposed by Gafarov et al. (2016) [GMM16]. The projection approach has two theoretical properties that we were not able to verify for the delta-method. First, projection is consistent in level uniformly over a reasonable class of data generating processes. Second, projection yields valid simultaneous inference; that is, it covers the whole impulse-response function (across different horizons and different variables) and not only its scalar coefficients.18 We note that in our application the projection confidence interval (which is wider than the delta-method bands) contains the realized value of IP, CPI, 2YTB, and FF for every horizon under consideration. Comparison with GK Robust Approach: Fig. 7 in the Appendix reports the robust-Bayesian credible set in Giacomini and Kitagawa (2015). The implementation of the robust-Bayes credible set (based on 10,000 posterior draws and using our algorithm to evaluate the endpoints) took around 9106 s.19 Comparison with GSM: Fig. 9 in the Appendix reports the 68% Bonferroni confidence set of Granziera et al. (2017).20 Appendix A.7.1 describes the algorithm and related computational issues. The computational cost is approximately 4,000 s on a single core machine for 10,000 grid points. It is hard to provide a general theoretical comparison of the length of the Bonferroni CS and the delta method. The efficiency ranking of the two procedures is likely depend on the particular DGP. One can see that, in our illustrative example, the 68% delta method CS is tighter than the corresponding Bonferroni CS with the same nominal level for almost all combinations of the horizons and time series. One possible explanation behind the larger length of Granziera et al. (2017) is that their procedure is uniformly consistent in level over the class of GDPs for which the reduced form impulse response functions converge to a normal distribution. We note that our delta-method is not guaranteed to have this property.","This paper focused on set-identified structural VAR models that impose equality and inequality restrictions to set-identify only one structural shock. For this class of models, the endpoints of the identified set have special properties that allow an intuitive and computationally simple approach to conduct frequentist and (asymptotic) robust Bayes inference. Specifically, the paper made three contributions: (i) We presented an algorithm to compute – for each horizon, each variable, a fixed vector of reduced-form parameters, and a given collection of equality and/or inequality restrictions – thelargest and smallest value of the coefficients of the structural IRF (see Theorem 1). Our algorithm did not require random sampling from the space of orthogonal matrices or unit vectors. Instead, we treated the bounds of the identified set as the maximum and minimum value of a mathematical program whose solutions we were able to characterize analytically. Our algorithm can be used outside our delta-method framework (for example, in computing the maximum and minimum response for the (Giacomini and Kitagawa, 2015) robust Bayes approach). (ii) We provided sufficient conditions under which the largest and smallest value of the structural parameters are directionally differentiable functions of the reduced-form parameters (see Theorem 2). This result also seems to be of interest in its own right and could be used to explore the frequentist properties of the robust-Bayesian procedure in Giacomini and Kitagawa (2015). (iii) Finally, we proposed a computationally convenient delta-method approach to conduct inference for the set-identified coefficients of the structural IRF. We presented sufficient conditions to guarantee the point-wise consistency in level and asymptotic robust Bayes credibility of our suggested inference approach. We note that the delta-method in this paper exploited the structure of the directional derivative. We illustrated our results by set-identifying the responses of different U.S. macroeconomic variables to an unconventional monetary policy shock. We used the theory and methods developed in this paper to assess the effects of the announcement of the second part of the Quantitative Easing program in August 2010."],["Child height is an important indicator of human capital and human development, in large part because early life health and net nutrition shape both child height and adult economic productivity and health. Between 2005 and 2010, the average height of children under 5 in Cambodia significantly increased. What contributed to this improvement? Recent evidence suggests that exposure to poor sanitation – and specifically to widespread open defecation – can pose a critical threat to child growth. We closely analyze the sanitation height gradient in Cambodia in these two years. Decomposition analysis, in the spirit of Blinder-Oaxaca, suggests that the reduction in children's exposure to open defecation can statistically account for much or all of the increase in average child height between 2005 and 2010. In particular, we see evidence of externalities, indicating an important role for public policy: it is the sanitation behavior of a child's neighbors that matters more for child height rather than the household's sanitation behavior by itself. Moving from an area in which 100% of households defecate in the open to an area in which no households defecate in the open is associated with an average increase in height-for-age z-score of between 0.3 and 0.5. Our estimates are quantitatively robust and comparable with other estimates in the literature. --------------------------------------------------------------------------------","Child height is an important indicator of human capital and human development, in large part because of its importance for adult economic productivity and health (Currie, 2009; Vogl, 2014). This is chiefly because height is determined by health and net nutrition in the first few years of life, a critical period for cognitive development (Case and Paxson, 2008). In poor countries, the disease environment to which children are exposed and income are important indicators for adult height (Bozzoli et al., 2009). The importance for height and child development of early-life health, relative to genetics, is even greater in these countries compared to richer countries (Martorell et al., 1977; Spears, 2012b). Recent econometric evidence suggests that exposure to germs from open defecation is an important determinant of child height in developing countries (World Bank, 2008; Spears, 2013), and epidemiological evidence suggests that potential mechanisms for this relationship include diarrhea, intestinal parasites, and environmental enteropathy, a disease of the small intestine. It is therefore important to better understand the relationship among sanitation, the early-life disease environment, and subsequent child health and human capital outcomes, especially in countries where practicing open defecation is widespread. We study the relationship between open defecation and child height in Cambodia, where 77% of households defecated in the open in 2005 and 63% in 2010. The primary contribution of this paper is to document what accounts for Cambodian children growing taller between 2005 and 2010. We show that much of the increase in child height over this period of time can be statistically accounted for by the increase in sanitation coverage over the same period. In studying Cambodia, this paper makes three important contributions to the literature. First, in Cambodia open defecation is particularly common, representing an enduring development challenge and an unusually threatening disease environment for children. Second, although the country remains far from eliminating open defecation or child stunting, Cambodia saw an improvement in child height from 2005 to 2010, coupled with a decrease in open defecation. This improvement, which was unusually rapid among developing countries, gives us the opportunity to study any association that may exist between the two.1 Third, we use decomposition techniques to examine whether the change in sanitation can statistically account for the improvement in child height over this period of time. The empirical analysis in this paper is in two parts. The main question we seek to answer is how much of the increase in child height from 2005 to 2010 can be statistically accounted for by the reduction in open defecation. We apply three complementary decomposition techniques: regression analysis to examine whether controlling for local open defecation eliminates the statistical importance of the indicator for survey year, a standard linear Blinder-Oaxaca decomposition, and a non- parametric decomposition. Although open defecation remains common in Cambodia, we find that its decline over the period studied can account for most or all of the increase in child height. Before computing decompositions, we first explore the association between exposure to open defecation and child height over time by combining the two most recent Demographic and Health Surveys in Cambodia and using panel data methods. This analysis provides support for the association between sanitation and stunting on which we rely for computing decompositions. Using urban and rural province part fixed effects, which isolates the variation within province parts, we find that the geographic areas in which open defecation decreased by more saw a greater improvement in child height, on average. In particular, we document negative externalities of open defecation: the rate of open defecation in the child’s locality is more important for child height than the sanitation practices of the child’s own household, indicating an important role for policy as households’ own private demand for latrines may be too low. This preliminary result is robust to a range of specifications. We test the mechanisms through which we expect open defecation to affect child height. We show that the association between open defecation and child height is steeper in urban areas, which is consistent with greater exposure to fecal pathogens when open defecation occurs in higher density areas, and that there is also an association between open defecation and child weight, consistent with mechanisms that affect child growth. We also perform two robustness checks: estimating our model by isolating the variation within regions in a particular year, and including data from the Demographic and Health Survey conducted in 2000. Open defecation and child stunting ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ According to WHO and UNICEF Joint Monitoring Program (JMP) statistics for 2015, 13% of people in the world defecate in the open. Of these roughly one billion people, about 7.5 million live in Cambodia, representing 48% of the country’s population.2 Among developing countries, open defecation is particularly common in Cambodia, notably more common than in the rest of Southeast Asia, where 10% defecate in the open, sub-Saharan Africa, where 23% do, and South Asia, where 34% do. Poor sanitation has important implications for the health and nutritional status of children, and a sizeable body of evidence in the fields of medicine and epidemiology demonstrates this link through a combination of three or more possible mechanisms, the importance of each of which may differ across contexts. These mechanisms include diarrhea, intestinal parasites, and environmental enteropathy. Checkley et al. (2004) used a cohort study in Peru to show that safe water and sanitation practices, which reduced fecal-oral contamination, were associated with fewer diarrheal episodes and better nutritional outcomes, as measured by height-for-age, in children. A meta-analysis conducted by Esrey (1996) has shown an effect of sanitation on intestinal parasites. More recently, researchers have investigated the role of environmental enteropathy (EE) as another, and perhaps more important, mechanism linking fecal-oral contamination to malnutrition (Humphrey, 2009). EE is a largely subclinical condition that is demonstrated by damage to the walls of the small intestine thereby reducing its absorptive capacity. There is substantial evidence linking markers of EE to lower height- for-age z-scores (Kosek et al., 2013; Goto et al., 2009; Campbell et al., 2003; Lunn et al., 1991), and there is now a growing literature linking the sanitation environment to markers of EE. One recent observational study finds that indicators of EE and malnutrition are higher among children who live in “dirtier” households, where they are exposed to more fecal pathogens (Lin et al., 2013). EE may even be at play when clinical conditions like diarrhea are absent. A recent study in Mali found an effect on child height, but not on diarrhea, of a randomly assigned Community-Led Total Sanitation program (Pickering et al., 2015). This paper joins a growing econometric literature documenting effects of open defecation on child height. Gertler et al. (2015) use experimentally-induced variation in open defecation from four randomized controlled trials (RCT) of independently conducted sanitation programs in different countries to find a causal relationship between village open defecation and child height. An RCT in Indonesia, one of the experiments studied in Gertler et al.’s meta analysis, shows that a Total Sanitation and Sanitation Marketing project increased average height of children living in households without access to sanitation at baseline (Cameron et al., 2013). Another RCT, conducted in Maharashtra, finds that improvements in sanitation brought about by the Indian government sanitation program increased average child height (Hammer and Spears, 2015). In India in particular, econometric studies of a government sanitation program document a link between sanitation and infant mortality (Spears, 2012a), child height (ibid.), cognitive achievement (Spears and Lamba, 2015), and adult wages (Lawson and Spears, 2016). In the vein of the analysis we conduct in this paper, Headey (2015) applies econometric methods to Demographic and Health Surveys in Ethiopia to identify an effect of improved sanitation on child height, and Spears (2013) investigates the difference in child height between India and Africa and documents that cross-country variation in sanitation can statistically explain a large fraction of international height differences. Several studies of sanitation programs, however, document no health impacts. Clasen et al. (2014), for instance, find no significant health impact in an RCT studying a government sanitation program in Orissa. Patil et al. (2014) similarly find no impact in a sanitation study conducted in Madhya Pradesh. The authors of both studies note, though, that the absence of an impact on health may have been because latrine use remained low despite large increases in latrine coverage.3 Sanitation in Cambodia ~~~~~~~~~~~~~~~~~~~~~~ Cambodia is a Southeast Asian country of 14 million people. Three-fourths of the population – and 90% of Cambodia’s poor – live in rural areas (Sobrado et al., 2014). The country has a moderate population density of about 75 people per square kilometer. Over the five-year period we study from 2005 to 2010 the fraction of the population living on less than $1.25 a day fell from around 35% to under 20%, and GDP per capita, in purchasing power parity terms, increased from $1962 to $2513 (2011 dollars), according to World Bank World Development Indicator Statistics. Open defecation has historically been high in Cambodia. We study the change over the period from 2005 to 2010, when exposure to open defecation of the average child under five fell by about 14 percentage points from 74% of the average child’s community in 2005. Sanitation coverage also varies substantially across geographic areas within Cambodia (Robinson, 2007), another dimension of heterogeneity that this paper exploits. Unsurprisingly, given Cambodia’s high rates of open defecation relative to Southeast Asia, government, NGOs, and other partners have played an important role in development activity in Cambodia. In the period in which we study, a number of sanitation programs were active in Cambodia, mostly following a supply- driven approach to sanitation. The private sector in Cambodia has played a significant role in the provision of the majority of latrines, accounting for almost 80% of all latrines built in the country (Rosenboom et al., 2011). Although there has been some diversity of approach and method across programs and over time during this period, much of the improvement in sanitation that this paper studies reflects new latrines that were largely financed by households themselves, complemented by some subsidized provision through development programs.","Demographic and Health Surveys (DHS) are large, nationally representative surveys conducted in poor and middle-income countries. We use data on the heights of children under five years old in the two most recent DHS in Cambodia, conducted in 2005 and 2010. Our dependent variable of interest is child height-for-age, which is a z-score of a child’s distance in standard deviations from the average height of healthy children in a reference population of the same sex and age in months. We compute z-scores using the WHO’s 2006 international reference population, and follow their recommendation of omitting children beyond six standard deviations from the mean. Our key independent variable is local area open defecation. Each household is classified as defecating in the open or not according to its report of where members “usually” defecate in the DHS questionnaire.4 However, infectious diseases often involve negative externalities (Gersovitz and Hammer, 2004), and intestinal disease resulting from open defecation is no exception: children are exposed to fecal pathogens from neighboring households. Therefore, we compute the fraction of households in a child’s survey Primary Sampling Unit (PSU) who defecate in the open, a continuous variable from zero to one, as a measure of “local area open defecation.”5 Because all households potentially contribute feces to the environment but not all households have children under five years old, we compute PSU averages from the DHS household recode. All other variables are also taken from the DHS, except measures of province mean consumption, which are from Knowles (2012), and population density, which are computed from Cambodian census data. A concern in observational studies such as these is that sanitation improvements may have been endogenously correlated with other improvements, for instance in consumption or wealth. For this reason, we focus on conducting an exercise in statistical accounting which provides an estimate of the fraction of the change in child height that can be statistically accounted for by the simultaneous change in exposure to open defecation. We thus do not attempt to estimate a causal effect of open defecation on child height, but provide regression estimates as supporting evidence of a relationship between open defecation and child height. In order to minimize the concern that our regression results may be driven by other factors that were simultaneously improving across Cambodia during the period of study, we employ several different strategies. First, we use geographic fixed effects. Each province is split into two parts classified as urban and rural. Urban and rural province part fixed effects, henceforth called region fixed effects, control for factors that differ across geographic areas that are correlated both with sanitation and child height. For instance, international organizations have targeted development programs in certain provinces in Cambodia. These programs may have led to sanitation improvements, but also could have improved other services or infrastructures that influence child nutrition. Region fixed effects control for such variation between regions. We also use time fixed effects to control for secular changes in child height over time. Second, it may be the case that open defecation in the child’s locality is correlated with other PSU-level variables such as village infrastructure or wealth. Thus, for comparison as placebo independent variables and as controls, we compute the fraction of households in a PSU with electricity, radio, television, refrigerator, bicycle, motorcycle, and car, as alternative measures of local living standards and infrastructure development. Third, we use a very extensive set of household controls that address maternal nutrition, household socioeconomic status, household education, and access to health care. These are discussed in more detail in Section 2.1. Fixed effects identification strategy ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The dependent variable, z, is a child’s height-for-age z-score. Local area open defecation is a fraction zero to one, and household open defecation is a binary indicator. Mother’s height, BMI, and age at birth, are included to control for heterogeneity in maternal nutrition, and to account for any possible direct effect of mother’s size (Ounsted et al., 1986). Standard errors are clustered by survey PSU, the level of heterogeneity in the independent variable of interest; 610 are more than enough for asymptotic clustered standard errors (Cameron et al., 2008). The fixed effects identification strategy differences out any fixed heterogeneity across regions within Cambodia, as well as the secular trend in child height. We include eight further sets of control variables in stages, in order to demonstrate robustness of our regression specification and the stability of the estimates of interest: Bilpt Birth characteristics: 13 indicators for birth order, 11 indicators for month of birth (Doblhammer and Vaupel, 2001), and whether the birth occurred in an institutional facility. Ppt Province characteristics: Province- level measures of average consumption and population density for 2005 and 2010. PSUlpt PSU characteristics: Fraction, from zero to one, of households within the PSU with electricity, radio, television, refrigerator, bicycle, motorcycle, and car.7 These controls serve as both placebo independent variables and as alternative controls for PSU welfare and infrastructure. Hilpt Household characteristics: seven binary indicators for whether the household has electricity, and owns a radio, television, refrigerator, bicycle, motorcycle, or car; ten indicators for floor material; 18 indicators for household size, ten indicators for type of cooking fuel; and 14 indicators for water source during the dry season and wet season, separately. These variables serve as additional controls for household wealth and socio-economic status. Eilpt Education: the number of years of education completed by the father and a binary indicator for mother’s literacy. Vilpt Vaccinations: an indicator for the child having a health and vaccination card, and eight binary indicators for the child receiving three rounds of the polio vaccine, three Diphtheria, Pertussis, and Tetanus (DPT) shots, and Bacillus Calmette–Guérin (BCG) and measles vaccines. Filpt Breastfeeding: a binary indicator for the child being breastfed immediately. Milpt Milk consumption: a binary indicator for the child being fed tinned, powdered, or fresh milk (de Beer, 2012; Baten, 2009; Baten and Blum, 2014) the previous day and/or night. This variable is only available for the youngest child under five.8 All specifications include 120 age-in-months times sex dummies Ailpt to non-parametrically control for the correlation between height-for-age z-score and age at measurement (Cummins, 2013). We test the mechanisms through which we believe open defecation causes stunting using two methods. If the biological mechanisms we assume are indeed occurring, we would expect the negative impact of open defecation on child height to be greater in areas where people live nearer together and are thus more exposed to others’ fecal pathogens. We test this mechanism using our data by introducing an interaction between our open defecation variable and an indicator for whether the PSU is urban. We would also expect there to be an association between open defecation and weight- for-age if open defecation affects child nutrition by causing intestinal disease, and we test for this as well. Finally, we perform two robustness checks. We use region by time fixed effects in order to control for time-variant regional characteristics. While this paper is primarily interested in exploring the extent to which changes in open defecation can account for changes in height over time in Cambodia, this robustness check nevertheless rules out any coincidental differences between regions over time from driving our main result. We also include data from the 2000 DHS conducted in Cambodia. Cambodia experienced a more modest decline in open defecation between 2000 and 2005 as compared to the subsequent five years. However, inclusion of data from 2000 presents an opportunity to check whether the main results of the analysis hold. Decomposition of change between 2005 and 2010 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ If open defecation is associated with child stunting, and if open defecation became less common between 2005 and 2010, then how much of the increase in child height over this period can be statistically accounted for by the reduction in open defecation? In a separate analysis from the fixed effects estimates of the association between open defecation and child height, we approach this question with three complementary decomposition methods in Section 5. First, in the course of the regression analysis, we see that controlling for local open defecation eliminates the statistical importance of the indicator for survey year. Comparing this estimate to the actual change in child height between 2005 and 2010 gives the fraction accounted for by the change in open defecation. A final, most flexible, decomposition is non-parametric reweighting (DiNardo et al., 1996).9 We generate a counterfactual estimate of 2005 child height by reweighting children in 2005 to match the 2010 distribution of open defecation exposure. Comparing this counterfactual estimate of the height of children in 2005 to actual child height in 2010 provides another estimate of the fraction of the difference in child height that can be accounted for by the difference in open defecation. This function allows us to change the distribution of exposure to open defecation for 2005 children so that it matches the distribution for 2010 children.","Children in our sample are on average almost two standard deviations shorter than the healthy international reference population, and they live in poor households with parents that have low levels of education. Table 1 presents sample means of many of the variables used in our analysis. Note that these summary statistics, like all estimates in this paper, are representative of children under five, and not of all Cambodians. The first and second columns show averages for 2005 and 2010, and the third column reports a test that these are different. Over the period we study, the height of children under five significantly increased relative to the international reference population while the fraction of open defecation in the average child’s community, and by individual households, significantly decreased. Standards of living also detectably improved in Cambodia between 2005 and 2010: households got richer, levels of education rose, and institutional deliveries and early initiation of breastfeeding increased. Height is associated with open defecation: non-parametric descriptive regressions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Fig. 1 presents local polynomial regressions that document that child height is associated with open defecation. Panel (a) plots child height as a function of age in months, replicating the well-known fact that most stunting occurs in the first two years of life, with height-for-age generally flat thereafter. These curves are plotted separately for children who live in local areas (PSUs) where no one surveyed defecates in the open (seven percent of children), where everyone surveyed defecates in the open (18% of children), and the rest, living in areas with an intermediate sanitation profile. Children in all three groups start off too short at birth, but the lines separate as stunting unfolds over the first two years. Children exposed to the most open defecation are more than a standard deviation shorter than children exposed to no open defecation, on average. However, open defecation is not the only cause of child growth defects: even children exposed to no open defecation are more than a standard deviation shorter than the reference population. One reason why children exposed to better sanitation are taller is because they are also richer. They are more likely to live in households that have toilets, while poorer households are more likely to defecate in the open. However, open defecation imposes negative externalities on everyone in the vicinity, and fecal pathogens in the environment transmit disease to neighboring children. Panel (b) of Fig. 1 plots average child height- for-age of children under five as a function of local area (PSU) open defecation. Households that do and do not defecate in the open are plotted separately. Unsurprisingly, children who live in households that use toilets are taller, on average, than children who do not. However, the key feature of the graph is that both lines slant downwards. Whether or not a child’s own household defecates in the open, she is shorter, on average, if more of the households in her community do. What is the association between sanitation and child height over time in Cambodia? Are the geographic areas that experienced a decrease in open defecation the same areas that experienced an increase in child height? Panel (a) of Fig. 2 visually explores this question. Each line in the figure connects one region’s 2005 averages of child height and rate of open defecation in the child’s locality to its 2010 averages. Each region represents an urban or rural part of a Cambodian province. Of the 38 lines, 25 slope downwards, indicating that in 25 regions of Cambodia, average open defecation in a child’s locality reduced and average child height increased. As a preliminary non-parametric statistical significance test, a binomial distribution reports that there is only a four percent chance of seeing at least 25 of 38 lines slope downwards if these slopes are independent of one another and equally as likely to slope up or down, indicating no association at all. The main contribution of this paper is to explore what can account for the increase in child height in Cambodia between 2005 and 2010. Panel (b) of Fig. 2 presents a visual depiction of the extent to which the change in open defecation can account for the change in child height. Within the two years 2005 and 2010, local polynomial regression lines plot average child height-for-age against open defecation rates in the locality. The lines are relatively close to one another and quite nearly pass through both of the overall year average points. Note that this is not mechanically determined: although each year’s average point must be on or near its own line, the two lines could be vertically far apart. If the lines were vertically separated, this would indicate differences in average child height for the two years even at the same level of open defecation in the locality. However, the fact that the lines are close together indicates that the association between height and sanitation in 2005 is similar to the association in 2010. Since the points are on similar lines, it appears that the within- year association between height and sanitation can statistically account for the between- years change in height. This figure is a visual representation of the results of our decomposition techniques, which we will discuss in further detail in Section 5 of this paper. Before we do so, however, we will first explore the relationship between exposure to open defecation and child height in greater detail using regression analysis.","This section explores whether the regions that experienced a greater decrease in open defecation also experienced a greater increase in child height between 2005 and 2010. Table 2 reports our regression results. Panel (a) shows estimates from OLS regressions without fixed effects, while Panel (b) displays results from regressions with region fixed effects. The OLS regressions identify the variation both between and within regions, while the fixed effects regressions isolate the variation occurring within regions, indicating that the relationship is not driven by coincidental differences between regions. An increase from zero to one in the rate of open defecation in the child’s locality is linearly associated with a decrease in children’s height by between 0.3 and 0.5 standard deviations, using fixed effects and varying sets of control variables. Regression results ~~~~~~~~~~~~~~~~~~ Column 1 simply reports the average improvement in child height from 2005 to 2010. Could the reduction in open defecation account for this overall increase in child height? Notably, when we introduce open defecation in the child’s locality in Column 2, the 2010 dummy variable becomes statistically and practically insignificant, indicating that the change in sanitation can statistically explain the average change in height over time. As we progress from Column 3 to Column 6, we progressively add more control variables. Column 3 adds household open defecation, mother’s anthropometry and age at birth, birth order, month of birth, and whether the birth occurred in a facility.10 While measures of maternal anthropometry are unsurprisingly predictive of child height, they do not diminish the role of sanitation. In Column 4, we further add province-level consumption and density, and PSU averages for electricity coverage and radio, television, refrigerator, bicycle, motorcycle, and car ownership. Inclusion of these variables helps control for PSU-level infrastructure and wealth. Column 5 adds controls for household wealth and socio-economic status by including household-level indicators for ownership of the assets listed above, floor material, household size, type of cooking fuel, water source, father’s educational attainment, and mother’s literacy. It also includes vaccination data for polio, DPT, BCG, and measles, availability of an immunization card, and early initiation of breastfeeding. In Column 6, an indicator for milk consumption is included. This model has fewer observations because this variable is only available for the youngest child under five, rather than for all children under five. For this reason, Column 5 represents the authors’ preferred specification. All specifications include 120 age-in-month dummies, separately for boys and girls.11,12,13,14 Three conclusions emerge from the table. The first is the robustness of the coefficient on open defecation in the child’s locality. The variable remains highly statistically significant in all specifications. If instead of clustering standard errors at the PSU level, we cluster at the more conservative level of 38 regions, the t-statistic on open defecation in the child’s locality in the most controlled specification, Column 5 of Panel (b), becomes −2.75, although this may be too few clusters for asymptotic results (Cameron et al., 2008). Secondly, the clear similarity in the size of coefficients from the OLS and fixed effects regressions suggests that the results from the OLS regressions are not a spurious artifact of heterogeneity across regions. Fixed effects are well-known to risk attenuation bias. However, that appears to be unlikely in this case due to the small differences between corresponding models of Panel (a) and Panel (b).15 Finally, an important result for policy is that the coefficient on household open defecation is not statistically significant.16 This is unsurprising because open defecation, like other sources of infectious disease, involves important negative externalities. In Table 3, we split the sample by whether or not the household in which the child lives defecates in the open. Open defecation in the community predicts child height, regardless of whether the household defecates in the open. This result corroborates the importance of negative externalities (Geruso and Spears, 2015). Because such externalities are a classic economic rationale for public action, they point to the importance of a policy response to open defecation. Mechanism check: steeper slope in urban areas ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The importance of open defecation in the child’s locality, rather than the household’s own open defecation, indicates a key role for externalities of disease. If so, then because children are more exposed to others’ fecal pathogens where people live nearer together, we would further expect that open defecation should have a steeper association with child height in urban areas, where population density is particularly high (Hathi et al., 2014; Bateman et al., 1993; Bateman and Smith, 1991). We test this by introducing an interaction between prevalence of open defecation in the locality and an urban dummy to the specification in Column 2 of Table 2. Does open defecation indeed have a steeper association with child height in urban parts of Cambodia than in rural parts? The fraction of open defecation in the PSU interacts with an indicator for urban place, with and without fixed effects. An increase in the rate of open defecation in the child’s locality from zero to one is associated with a 0.34 (s.e. 0.16) standard deviation greater decrease in average child height in urban areas compared to rural areas using a standard OLS model, and a 0.41 (s.e. 0.22) standard deviations greater decrease using a model with region fixed effects. These coefficients are statistically significant at the two-sided ten percent level. This finding is consistent with Spears (2012a), Spears (2013), and Hathi et al. (2014). Mechanism check: weight-for-age ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ If open defecation does indeed affect child height by causing intestinal diseases that lead to undernutrition, then we may simultaneously expect to see an effect of open defecation on weight, another measure of nutritional status. Weight-for-age is associated with recent diarrheal episodes (Schmidt et al., 2010) and environmental enteropathy (Lin et al., 2013; Humphrey, 2009). As a measure of nutrition, weight-for-age is more responsive to recent changes in diet, care practices, or the disease environment, while height-for-age is a measure of net nutrition in the first two years of life (Waterlow 1972; Black et al., 2008). Open defecation in the child’s locality statistically significantly predicts weight-for-age, with and without fixed effects. Replacing height- for-age with weight-for-age in Column 3 of Table 2 shows that an increase in open defecation in the child’s locality from zero to one is associated with a reduction in weight of 0.56 (s.e. 0.083) standard deviations in a standard OLS model, and 0.23 (s.e. 0.074) standard deviations in a model with region fixed effects. These coefficients are significant at the two-sided one percent level. Robustness checks ~~~~~~~~~~~~~~~~~ We test the robustness of our results using two methods: including region by time fixed effects in order to isolate the variation occurring within regions in a particular year, and adding data from the DHS conducted in Cambodia in 2000. Table 4 reports the results of our robustness checks. In Column 1, we include region by time fixed effects. This more restrictive specification changes the coefficient on open defecation in the child’s locality only very slightly: using the same set of controls as in Column 5 in Table 2, the associated reduction in child height arising from an increase in open defecation in the child’s locality from zero to one is 0.35 standard deviations using region fixed effects (Table 4, Column 1) and 0.29 using region by time fixed effects (Table 4, Column 2). Column 3 of Table 4 includes data from the 2000 Cambodian DHS. Between 2000 and 2005, Cambodia experienced a much more modest decline in open defecation of only eight percentage points, compared to the subsequent five years in which the decline was 14 percentage points. Nevertheless, including data from the year 2000 only supports our main result.","How much of the increase in child height between 2005 and 2010 can be explained by the decrease in average exposure to open defecation? Econometric decompositions ask how much of the difference in the outcome variable across two groups can be accounted for by observable differences in input variables (Fortin et al., 2011). Although the canonical use of decompositions in labor economics is to analyze differences in economic outcomes (such as wages) between two groups (such as black and white people in the United States), here we will be asking how much of the difference in child height in Cambodia between 2005 and 2010 can be accounted for by the difference in the level of open defecation in a child’s locality. In general, econometric decompositions of observational data are tools of statistical accounting that may or may not have a causal interpretation depending on the details of the data and the source of heterogeneity studied. We thus interpret decomposition results conservatively as accounting for differences. In Section 3.1, we discussed Panel (b) of Fig. 2, which presents a visual depiction of the extent to which the change in open defecation can explain the change in child height. Each line plots local polynomial regressions of the sanitation height gradient for each year. The relative closeness of the lines indicates that the gradient is similar in both years, and the fact that the overall year averages for both years are almost on these lines indicates that the within-year association between height and sanitation appears to statistically account for the between-years change in height. Various methods of econometric decomposition are available, and we study three. The first and simplest was already presented in the difference between Columns 1 and 2 of Table 2. Adding a linear control for open defecation in a child’s locality (PSU mean open defecation) eliminates a statistically significant difference in child height between the two DHS rounds. Using OLS and fixed effects estimation strategies, controlling for open defecation in the child’s locality statistically accounts for 88% and 86% of the difference in child height from 2005 to 2010, respectively (see Table 5). The following two sub-sections will consider a Blinder (1973) – Oaxaca (1973) decomposition, and will apply a non-parametric reweighting technique. Blinder-Oaxaca decomposition ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Blinder-Oaxaca decomposition calculates the explained change in child height between 2005 and 2010 using an estimate of the height-sanitation gradient. In this case, two regressions separately estimate β12005 and β12010 and then average them to create a counterfactual within-year height sanitation slope. Multiplying this by the difference in the average exposure to open defecation between 2005 and 2010 provides an estimate of the change in child height that can be accounted for by the change in open defecation. As Table 5 shows, this approach finds that the reduction in open defecation to which the average child was exposed can statistically account for 0.12 of the 0.13 standard deviation difference in child height. Thus, 92% of the difference in height can be accounted for by the difference in sanitation. Non-parametric reweighting decomposition ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ This method creates a counterfactual average height for children in 2005, reweighted to match the 2010 distribution of open defecation. In particular, the sample is split into 12 bins of PSU open defecation levels: ten deciles with extra categories for open defecation of zero and open defecation of one. These are crossed with an indicator for own household open defecation to create 22 overall open defecation bins (there are 22 instead of 24 because there are no children in households that openly defecate who also live in PSUs where nobody sampled defecates in the open, and vice versa). Then, within each of the 22 bins, the total sample weight is computed separately for 2010 and 2005. A new set of weights is computed for children in 2005 by multiplying their sampling weight by the ratio of their bin’s 2010 total sampling weight to their bin’s 2005 total sampling weight. Finally, the new weighted average represents the counterfactual height of children in 2005 if they had been exposed to the same levels of open defecation as children in 2010. This result is compared with the other decomposition methods in Table 5. The true sample mean height for age was −1.77 in 2005 and −1.64 in 2010. When the 2005 sample is reweighted to match the 2010 sanitation distribution, the counterfactual mean height-for-age is −1.63, essentially the same as the true 2010 average (the difference is not statistically significant, t = 0.29). Therefore, all three approaches to decomposing the change over time in child height reach similar conclusions. Using simple pooled regression, a Blinder- Oaxaca decomposition, or non-parametric reweighting, the decline in exposure to open defecation can statistically account for almost all of the approximately 0.13 standard deviation increase in height-for-age.","Child height is an important economic variable predicting adult human capital, cognitive achievement, and health. The average child under five in Cambodia was 0.13 standard deviations taller in 2010 than in 2005. Decomposition analysis finds that much of the increase in child height between 2005 and 2010 can be accounted for by the simultaneous reduction in open defecation. At the same time, regression analysis finds a robust and large association between exposure to open defecation and child height. The point estimates computed in this analysis are consistent with other studies. A meta-analysis combining data from three large-scale randomized interventions conducted independently in India, Indonesia, and Mali finds that eliminating open defecation in a village in which everyone practices open defecation is associated with a 0.4 standard deviation increase in height (Gertler et al., 2015). This is similar to the point estimates found in this paper of between 0.3 and 0.5 standard deviations, with region fixed effects and controls. The change in child height over this period of time represents an important difference: Spears’ (2012b) estimates of the height-cognitive achievement gradient for Indian children suggest that a 0.13 standard deviation increase in child height would be associated with a 1–4 percentage point increase in the probability of being able to read words or paragraphs among 8–11 year-olds. This difference is also quantitatively similar to the India-Africa height gap (Spears, 2013). These results indicate that widespread open defecation could be a critical constraint for human development. Moreover, we have seen various indicators of the role of negative externalities in propagating fecal pathogens. The health benefits of better sanitation are significant, and in Cambodia, the cost associated with constructing a latrine can be as low as $25 (Rosenboom et al., 2011). Lawson and Spears (2016) find a robust relationship between adult wages and the disease environment during childhood in India, and the fiscal implications indicate that public investment in sanitation infrastructure may come at very low net present cost. Interventions that are long-term are more likely to lead to sustainable improvements in nutritional indicators for children (Reiger and Wagner, 2015). Thus, if latrine adoption is durable, it can have a substantial impact on child height. Between 2005 and 2010, open defecation decreased, and child height increased, but open defecation is still common in Cambodia and the mean child was still 1.64 standard deviations below the healthy reference population in 2010. In any country where this is the case, spillovers of poor sanitation indicate that reducing open defecation must be a policy priority."],["This paper, which reexamines the Poyago-Theotoky model, provides additional investigation that was conducted under a corrected environmental damage parameter. As new findings, we obtain the following. First, social welfare under a time-consistent emission tax (emission subsidy) policy is always welfare-enhancing rather than the case of laissez-faire. Second, if the environmental damage parameter is sufficiently small, then the equilibrium emission tax rate is invariably negative. It is therefore an emission subsidy. Moreover, total emissions under the emission subsidy scenario become less than those under laissez-faire if the damage parameter is sufficiently small, and if the R&D cost is low. However, total emissions under the emission subsidy become greater than those under laissez-faire if the damage parameter is sufficiently small, and if the R&D cost is high. © 2013 The Authors. --------------------------------------------------------------------------------","In an industrialized economy in which giant firms dominate markets, perpetual innovation drives the rate of growth in living standards continuously. Particularly, it is true that the petroleum and chemical industries have both been oligopolistic, routinizing corporate and environmental innovations, and contributing to economic growth with some environmental friendliness. These industries typically have developed pollution abatement instruments such as desulfurization equipment and denitrification equipment, which are categorized into end-of-pipe technologies. Emission regulation has received great amounts of attention from many researchers. Studies of environmental regulation and emission-reducing R&D in an oligopolistic market have been conducted by Chiou and Hu (2001), Puller (2006), Riveiro (2008), Celik and Orbay (2011), Yakita and Yamauchi (2011), Pal (2012), and others. Particularly, Jaffe et al. (2002) and Requate (2005) provide excellent surveys in that field. The strategic behavior of polluting oligopolists has a strong impact on social welfare. Therefore, it is highly necessary for policy designers to accumulate the welfare performances and implications of various environmental policy instruments in oligopolistic markets. In environmental policy studies, the timing of policy variables is an important topic. Poyago-Theotoky (2007) presents pioneering work on emission tax policy and competition policy related to quantity-setting duopolists with end-of-pipe technology. Particularly, that investigation examines whether polluting Cournot duopolists' coordination of behavior in environmental R&D is socially allowable or not when the government has no precommitment capability with respect to emission taxes. The study finds important policy implications for socially desirable R&D formation under time-consistent emission taxes. Furthermore, Poyago-Theotoky (2010) announces a corrigendum showing that the negative emission tax (i.e., emission subsidy) might be partially justified if it considerably improves the market inefficiency caused by Cournot duopolists. Indeed, this is a surprising result, but the fundamental question remains. That is “When does an emission subsidy reduce (or increase) emissions?” The answer is fervently sought by policy designers. Nevertheless, little attention has been devoted to that question. In real society, many developed countries confront obligations related to greenhouse gas reduction, along with international commitments such as the Kyoto Protocol. In addition, most emerging countries seek both industrialization and environmental improvement. Therefore, it is necessary for social planners to investigate the regulatory circumstances under which a time-consistent emission subsidy reduces (or increases) emissions. On the other hand, in the existing literature on environmental regulation, negative emission taxes (emission subsidy) in oligopolistic markets are discussed by Requate (1993a, 1993b), Petrakis and Xepapadeas (1999, 2003), David (2005), Fujiwara (2009), Ben Youssef and Dinar (2011), and others.1 Those studies point out that negative emission tax rate can exist in equilibrium when environmental damage is sufficiently small. However, investigations of negative emission taxes are utterly inadequate. This paper carefully analyzes the emission-reducing effects of negative emission taxes. Arguments presented in this paper proceed as follows. Section 2 introduces the Poyago-Theotoky (2007) model and equilibrium outcomes. Section 3 presents an examination of the sign of the equilibrium emission tax rate and effects on total emissions. Section 4 presents policy implications and conclusions.","This section presents the model and its equilibrium outcomes. Market structure Considering an industry comprising two homogeneous firms, firm i and firm j, engaging in quantity competition with the same cost structure and emissions- reducing technology, qi is assumed to denote firm i's output. Inverse demand is given as p(qi,qj) = a − (qi + qj), (i, j = 1, 2; i ≠ j), where a(> 0) is a market size parameter. Environmental R&D and cost structure The value of each firm's emissions per unit output is assumed to be one. Firm i's environmental R&D effort is denoted as zi. Both firms use end-of-pipe technology for pollution abatement. Although this abatement technology is insufficient to reduce emissions per unit output, it mitigates emissions by adsorbing emissions at the end of the production process. Firm i receives benefits not only from its own environmental R&D efforts but also from the efforts of its rival. When firm i's production level is qi, then the R&D expenditures (γ/2)zi2, (γ > 0) enable firm i to abate its emissions from qi to ei(qi,zi) ≡ qi − zi − βzj.2 A lower value of γ implies higher efficiency of the environmental R&D cost. Symmetric parameter β ∈ [0,1] denotes the spillover effects of R&D. Firm i's positive externality from rival's R&D efforts is denoted as βzj. No fixed costs for pollution abatement are necessary. In addition, firm i's total cost function is additively separable with respect to production costs and R&D expenditures: C(qi,zi) = cqi + (γ/2)zi2, (c > 0, A ≡ a − c > 0). Timing The regulator has no precommitment ability for emission tax rate t. The time structure is the following: Stage 1: Firm i determines zi to maximize its own profit (πi) or joint profits (πi + πj). Stage 2: The regulator determines emission tax rate (t) to maximize social welfare. Stage 3: Firm i determines output level (qi) noncooperatively to maximize its own profit. Equilibrium outcomes ~~~~~~~~~~~~~~~~~~~~ Poyago-Theotoky (2007) examines two environmental R&D scenarios (R&D competition and R&D cartelization) and derives the subgame-perfect Nash equilibrium (SPNE) under a time- consistent emission tax.4 To begin, we explore the case of environmental R&D competition. In stage 3, firm i's profit is πi(qi,qj) = {a − (qi + qj)}qi − cqi − t{qi − zi − βzj} − (γ/2)zi2. Each firm chooses an output level noncooperatively and simultaneously to maximize its own profit. From the first-order conditions, the symmetric equilibrium output is calculated as q(t) = (A − t)/3. In the first stage, firm i's profit is πi(zi,zj) = [q(t(zi,zj))]2 + t(zi,zj){zi + βzj} − (γ/2)zi2. Each firm determines its environmental R&D efforts noncooperatively and simultaneously. From the first-order conditions ∂πi(zi,zj)/∂zi = 0, (i, j = 1, 2; i ≠ j), we obtain the equilibrium R&D efforts zN and the equilibrium values of other variables. The results are presented in Table 1.5 Environmental R&D cartelization implies that each firm determines its environmental R&D effort collusively to maximize joint profits (πi(zi,zj) + πj(zi,zj)) during the first stage. Each equilibrium value under R&D cartelization is also reported in Table 1.","This section presents examination of the sign of the equilibrium emission tax rate and effects on total emissions. Sign of emission tax rate ~~~~~~~~~~~~~~~~~~~~~~~~~ Poyago-Theotoky (2007) proves that each firm always has some incentive for R&D cooperation, that is πC ≥ πN, and that SWC > SWN if 1/2 < d < 3/2.6 Furthermore, that research shows that social welfare under environmental R&D cartelization is higher than the case of noncooperative environmental R&D, except in the case of large damage and lower R&D cost. The exceptional case is presented as Region IV in Fig. 1. In Region IV, SWC < SWN. On the other hand, in Regions I, II, and III in Fig. 1, SWC > SWN. Precisely speaking, if d ≥ 3/2, then SWC ≥ (<) SWN for all γ ≥ (<)γφ.7 In this model, a negative emission tax rate (i.e., emission subsidy) is fundamentally equivalent to the policy mix of a production subsidy and an abatement tax because each firm's emission function is assumed as pollution generated by production minus net abatements. A production subsidy has two effects. One is a damage-increasing effect. The other is the decreasing effect of market inefficiency. When d is sufficiently small, the increasing effect of environmental damage is dominated by the improvement effect on market inefficiency. This is the economic intuition underlying the negative emission tax rate.10 Particularly, part (i) of Proposition 1 is newly obtained under the parameter range corrected by Ouchida and Goto (2011), whereas Poyago-Theotoky (2010) points out the existence of a negative emission tax in the case of part (ii) of Proposition 1. As shown by Requate (1993a, 1993b), Petrakis and Xepapadeas (1999, 2003), Fujiwara (2009), Ben Youssef and Dinar (2011), and others, a negative emission tax can be justified socially when environmental damage is sufficiently slight.11 The Poyago-Theotoky (2007) model is related with these studies, and is also somewhat different in the following four points: timing of emission taxation, market structure, spillover effect, and the type of emission function. Despite such differences, Proposition 1 of the present paper and Poyago-Theotoky (2010) show that a negative emission tax is realized in equilibrium when environmental damage is sufficiently small. This fact suggests that the necessary condition for negative emission tax is still robust in the case of a time-consistent emission tax policy toward Cournot duopolists. Negative emission tax and total emissions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Readily inferred from Proposition 2 (iv). □ In Fig. 1, γE is described as the borderline between Region I and Region II. In both Regions I and II, the sign of the equilibrium emission tax rate is negative: the emission subsidy is realized in the equilibrium. Therefore, we can identify that in Region I, the emission subsidy yields greater total emissions than those under laissez-faire. In Region II, however, it yields less total emissions than those under laissez-faire, i.e., the emission subsidy reduces total emissions. Among reports of existing studies of the literature, except for this research, no study has ever explicitly examined emission-reducing and emission-increasing effects under emission subsidy. We now explore the intuition underlying the existence of Regions I and II. As environmental damage d decreases, the equilibrium emission tax rate tC, R&D effort zC, and accordingly R&D effort with a spillover effect (1 + β)zC respectively denote decreases. Additionally, as the R&D cost parameter γ becomes large, each value of zC and (1 + β)zC decreases greatly. Ultimately, if the value of d is sufficiently small, then the sign of tC becomes negative. However, as the value of d becomes small, the equilibrium production level per firm qC becomes greater. Consequently, as shown by Region I, total emissions under emission subsidies become greater than that under laissez-faire if the damage parameter is sufficiently small, and if R&D cost is not low. In Region II, although the sign of equilibrium emission tax rate is negative, total emissions under an emission subsidy become smaller than the case of laissez-faire because the emission- increasing effect is dominated by the large abatement effect.12 From the definitions of γCt and γE in Eqs. (2) and (3), we obtain the following proposition with respect to Region II. The greater the spillover effect becomes, the broader the regulatory circumstances in which emission subsidies reduce total emissions. Proposition 3 states that Region II is increasing in β. Fig. 1 presents this result graphically. The value of β can be regarded as the level of intellectual property rights (IPR) protection. Therefore, a stronger IPR protection level generates smaller Region II. When the spillover is perfect (β = 1), Region I becomes the smallest.","In this section, we derive two policy implications from our analysis explained above: one is derived from part (i) of Proposition 2 (i.e., Region I in Fig. 1); the other is derived from part (ii) of Proposition 2 (i.e., Region II). The latter is more meaningful. The petroleum refining industry and the petrochemical industry are examples of polluting industries using desulfurization equipment and denitrification equipment. Those industries, which are directly applicable to this model, emit plenty of greenhouse gasses. This study provides fundamental results that are expected to be useful in avoiding policy failure when the regulator's policy variable is a time-consistent emission tax. Part (i) of Proposition 2 implies that if the environmental damage parameter is sufficiently small, and if R&D cost parameter is sufficiently high, then the government will choose a time- consistent emission subsidy that yields “greater emissions” than those under laissez-faire in Cournot duopoly. Now, let us assume that the government faces the obligation of emission reduction. In this case, a policy contradiction is unavoidable. In reality, governments of the developed countries committed to reduction of greenhouse gas emissions. The regulator of those countries should investigate the real value of R&D spillover. On the other hand, part (ii) of Proposition 2 implies that if the environmental damage parameter is sufficiently small, then the government will choose a time-consistent emission subsidy that yields “less emissions” than those under laissez-faire in Cournot duopoly. That is to say, an emission subsidy can reduce the total emissions.13 Previous studies have devoted little attention to such circumstances (i.e., Region II). As described in this paper, our new findings present two points of particular significance.16 One is derivation of incremental policy implications. Another is to give a further theoretical foundation for time-consistent emission tax policy in quantity-setting duopoly. This research plays an indispensable complementary role for contributions by Poyago-Theotoky (2007, 2010)."],["In this study, we estimate the effect of fast food environment surrounding schools on childhood body mass index (BMI). We use two methods that arrive at a similar conclusion, but with different implications. Using school distance from the nearest federal highway to instrument for restaurant location, we find the surrounding restaurants to only marginally affect a student's BMI measure. The effect size also decreases with increasing radial distances from school, 0.016 standard deviations at one-third of a mile and 0.0032 standard deviations at a mile radial distance. This indicates the decreasing influence of restaurants on a child's BMI as its distance from school increases. On a subset of students who were exogenously assigned to different school food environment, we find no effect of the fast food restaurants. An important contextual aspect is that nearly all schools in this sample observed closed campus policy, which does not allow students to leave campus during lunch hours. --------------------------------------------------------------------------------","Rising rates in childhood obesity have become an important worldwide health and public policy issue. Even though this has received more attention in the United States, childhood obesity is a growing problem in other countries, including those in Europe and Asia. Because the majority of children attend public schools, there is much interest in regulating the food environment in schools to promote health objectives (Story et al., 2009; Sharma et al., 2009). France, for example, banned food promotion in schools, whereas Mexico and India have banned sales of sodas and certain unhealthy foods in schools (Villanueva, 2011; Barquera et al., 2013; Khandelwal and Reddy, 2013). In the US, regulation has mostly focused on school meals and those foods stocked in school vending machines1 . Addressing childhood obesity has become a policy priority, especially because obesity during childhood may continue into adulthood (Serdula et al., 1993). Most public policy efforts have focused on school lunches and other food products sold within the school. There is a growing concern, however, that fast food restaurants around schools might increase caloric intake among children and thereby increase obesity (Davis and Carpenter, 2009). There are incentives for fast-food restaurants, especially national chains, to target children for increasing current sales and creating brand loyalty. This is more apparent among hamburger fast food restaurants, sandwich places, and pizzerias (Austin et al., 2005). Collectively, we will refer to these three restaurant types as fast food restaurants. An important question of public policy relevance is to know whether the presence of fast food restaurants around schools have any effect on childhood obesity. We seek to identify the effect of fast food restaurants surrounding schools on a child’s BMI. Estimating the causal effect of fast food restaurants surrounding schools, however, is challenging due to: a) self-selection of families into neighborhoods; b) omitted variable bias; and c) selection of fast food restaurants in into neighborhoods according to fast- food demand. Self-selection of parents due to occupation, education or other reasons could create neighborhoods that could be similar in their attitude towards health and well-being of their children. Omitted variables in the determination of a child’s BMI could include features of the school food environment. We address biases in the estimation using an instrumental variables approach, where the instrument is the proximity of the school to US highways that create a demand for fast food restaurants to serve highway travelers. This instrument has been used in several studies (for example, Dunn, 2010; Anderson and Matsa, 2011). In addition to this method, we estimate individual-level fixed effects for a subset of students who were exogenously reassigned to another school. This strategy follows that used by Asirvatham et al. (2018) in their analysis of peer effects. Both these methods arrive at a similar conclusion – fast food restaurants surrounding schools have a negligible effect on BMI outcomes of school children. The analytic sample is based on measured BMI records from an ongoing BMI screening program within the Arkansas public school system. There are several ways that fast food restaurants surrounding schools could influence BMI outcomes. First, fast food restaurants and pizzerias offer calorie-dense foods that play a direct role in obesity. Even though most schools in this study observe closed campuses, a high density of restaurants around schools will generally increase the likelihood that students visit restaurants before or after school hours either unsupervised with friends or with a parent or caregiver. Second, restaurants located nearby reduces travel cost to obtain food. It could also be that proximity disincentivizes traveling longer distances to purchase healthier foods. Third, exposure to the sight of restaurants might increase the likelihood that children would choose fast foods after being exposed to fast food restaurants en route from school to home. Fast food signage could reinforce marketing messages aimed at children and increase likelihood of fast food requests on food-away-from-home meal occasions. School meals can be served as early as 10:30 am and many children will be hungry at the end of the school day. The presence of fast foods in school environment could thus increase desire for fast foods on other dining occasions regardless of whether the child consumes fast food on the way to or from school. Finally, teachers who leave campus for lunch could bring in items from restaurants (e.g., drink cups) and thereby model fast food as a reasonable meal choice. This is similar to the argument that advertisements targeting children might increase fast food consumption (Jashinsky et al., 2017). Study participants in Cambridgeshire, UK, who faced a greater exposure to fast food outlets showed higher intake of takeaway food (Burgoine et al., 2014).","Much of the work in this line of research provides some evidence of a positive association or correlation of obesity rates or BMI with restaurant availability or proximity (Williams et al., 2014). Most of the studies are correlational and therefore it is difficult to imply any causality. In this section, we discuss few of the studies that have conducted a more rigorous analysis going beyond correlation or association. These studies in general find a small effect. Three studies used an instrumental variables strategy based on proximity to highways (or highway on/off ramps), to estimate the impact of fast food density on body weight outcomes among adults. In a county-level study, Dunn (2010) used the number of exits on a highway within a county to instrument for the number of restaurants in a county and finds restaurant availability to effect only females and non- whites in medium-density counties in 11 states. Such a relationship was not found in rural or low-density counties. In a similar study on residents of central Texas, Dunn et al. (2012) reported only non-whites to exhibit higher obesity rates in response to fast food exposure. These authors also used distance to the nearest major highway as an instrument for the fast-food restaurant. Anderson and Matsa (2011) find no effect of fast food restaurants. They also exploit the variation in travel costs of local residents to restaurants along interstate highways, which primarily serves travelers. Using related data, they observed that obese individuals complement eating outside by consuming nutritionally deficient or “junk foods.” They conclude that policies focusing solely on regulating fast food restaurants may not achieve any significant reduction in BMI. The above studies examine adult populations. Two studies analyze fast food restaurants around schools and find fast food restaurants to significantly affect childhood obesity rates. Currie et al. (2010) uses a cross-sectional sample of students from grade 9 in public schools in California. In their study, the effect is identified by comparing groups of individuals at only slightly different distances to a restaurant (i.e., of only one-tenth of a mile). They measure changes in exposure, with obesity being measured as the fraction of obese students at the grade-level in a school. In their study, they use changes in obesity prevalence in the same grade (grade 9) but across different years. Alviola et al. (2014) study the public school sample in Arkansas. Following the earlier studies on adult populations, their instrumental variable was the distance of school from the nearest highway. Using school-level cross-sectional data, they report coefficients that measure the difference in obesity rates across schools that face different restaurant counts. Their study estimates a 1.23 percentage point increase in school obesity rates in response to an additional restaurant within a one-mile radius from a school. In this article, we also focus on Arkansas public schoolchildren. However, unlike Alviola et al.’s (2014) school-level analysis, this study uses individual-level analysis of BMI z-scores over time; i.e, we use student-level panel data in contrast to the school-level cross-sectional data they used. In addition to an identification strategy based on an IV estimation using distance of schools to the nearest highway as an instrument, we also employ an additional identification strategy based on a plausibly exogenous reorganization of public schools that caused some children to be reassigned to a new school zone. Hence, these re-assigned students face a new food environment in their newly assigned school. The present article extends these studies in several ways. One, we have a unique panel dataset, which gives us the ability to measure changes within a student and estimate fixed effects at the individual level. Several of the previous studies use obesity rates (proportion of obese adults or children) or obesity status (a binary variable). We are able to estimate the impact on precise changes in BMI in terms of standard deviations of the child’s BMI z-score. Two, in contrast to several past studies, the BMI data used here are measured by trained personnel, as opposed to being self-reported. Three, we control for the commercial food environment near a child’s residence with precise geographic information. It is important to control for other sources of calories because this could influence the quality and quantity of food consumed. Four, as mentioned above, this study uses two methods to identify the effect of fast food availability around schools on children’s BMI z-scores: IV and the exogenous assignment of students to nearby schools as a result of a court-mandated restructuring program. These are discussed in detail in the section Identification Strategies.","The primary aim of this study is to estimate the effect of fast food restaurants surrounding schools on childhood obesity. The variation in a student’s exposure to fast food restaurants comes in two ways. One, change in the count of restaurants surrounding schools. Two, students moving to different schools either because of a natural progression through the public school system or by just relocating within the state and, thereby, to another school. In this dataset, about 42 percent of the students relocated but remained in the public schools in Arkansas. A child’s BMI could be influenced by several environmental factors inside schools, outside of schools, within the family, and within the community, some of which could be related to the density of fast food restaurants. Our data allow us to control for several observable factors. However, note that the coefficient on the number of fast food restaurants takes into account only the covariation between the restaurant count around the school and BMI z-score. Food intake from home might bias the coefficient of interest, but we include variables that capture the commercial food environment around a student’s residence, including distance to nearest grocery, dollar store, convenience store, fast food, pizzeria and sandwich place. To the extent household behaviors towards food consumption persist with time, the inclusion of student-level fixed effects can be advantageous as they difference out unobserved time- invariant factors that might be correlated with the variable of interest. In terms of the features of the food environment within schools, all schools were subject to Act 1220 of 2003. Act 1220 required nutrition standards to be applied to all foods and beverages sold or made available to maintain a healthy school environment. For example, Act 1220 requires 50% vended beverages to be a healthy choice such as water, 100% fruits juice and low- fat/fat-free milk. Act 1220 was in effect across Arkansas during the entire study period. Thus, we do not expect wide variation in within school environment across schools in terms of the food and physical education/activity environment. Other factors within schools could play some role. For example, peers may influence diet and physical activity choices that might then influence body weight. One challenge with including peers’ variable is that it brings with it into the model a host of other biases. As described in Asirvatham et al. (2018), these biases could further complicate the model. We, therefore, do not include peers’ weight variable but include school-level percentages of different race groups and meal status. Both of these are correlated with obesity and, therefore, to some extent account for peers influence (Boyd et al., 2011). Nevertheless, we acknowledge that the methods used here can only partially addresses the biases created by unobserved changes within the school environment during the study period. There could also be common factors at the community-level that might drive restaurant counts. Consider a town with a population that does not demand high calorie foods. Such a population might also have lower restaurant counts in the near vicinity of schools. One implication of such demand is that the lower income population might demand lower-cost food products, which are often provided by fast food restaurants. To address all of these biases, we use an instrument that drives the fast food restaurants’ location choice.","We estimate the effect on two samples, one on the whole student population and the other on a subsample of students who were exogenously reassigned to other schools in response to a court-mandated reorganization of schools and districts. This allows us to test for possible mechanisms of the effect and propose some recommendation to reduce its effect on BMI outcomes. Although complementary, the analyses we conducted using these two samples provide different insights. We utilized the instrumental variables approach on the larger student population for whom all information was available, and then exploited a natural experiment of exogenous school assignment of a subset of students in response to a court- mandated school restructuring. Below we discuss both methods in detail. Instrumental variable (IV) method ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The instrument we use is the distance of the school to the nearest interstate or US highway. This has been well established in the literature. Dunn (2010); Anderson and Matsa (2011), Dunn et al. (2012), and Alviola et al. (2014) use some measure of highway proximity as an instrumental variable for fast food restaurants. Fast food restaurants locate in places where there is demand for the products they offer. By locating closer to highways, fast food restaurants are able to profit from the demand from travelers, but this could exogenously increases exposure to fast-food restaurants for those living near major highways, or in our case for schools located close to major highways. The identifying assumption is that nearness to highway is a factor in restaurant location decisions. That is, nearness to highway is a factor in restaurant location decisions because proximity to interstate and US highways might bring in firms supplying services and products to travelers. Using County Business Patterns data, Dunn (2010) shows that interstate exits drive the location of fast food establishments in way that is different from other businesses. He also shows that number of interstate exits is not correlated with other healthy behaviors, such as fruit and vegetable consumption and physical activities. Dunn (2010) discusses both aspects of demand and supply that could bias an estimate. For example, higher demand for fast food could lead to more restaurants. On the other hand, residents with preference for health could enforce restrictions on location and number of fast food restaurants. He concludes that the direction of the bias is generally positive. In our context, schools factoring in distance from state and federal highways might violate the exclusion restriction. This violation would occur because a school’s decision and a restaurant’s decision are influenced by the same explanatory variable. Even though some states do stipulate school construction to be at a distance from highways, Arkansas only stipulates distance from any source of sound that produces 65 decibels sustained and 75 decibels peak2 . Thus, we argue that the distance of schools from highway in the context of policies in place in Arkansas does not violate the exclusion restriction. During our study period, the locations of highways are fixed. However, students change schools as part of a natural progression through the public- school system and this creates time series variation in the distance between the child’s school and highway. The instrument used in the literature, in general, considers only interstate highway, but due to the sparseness of interstate highways in Arkansas, as shown by Alviola et al. (2014), we consider both interstate and US highways for our instrument.5 The standard way to include the IV in an equation is to use it as a single variable. In this study, we include it in two different ways. Firstly, as a single continuous variable. Secondly, to control for a non-linear relationship, the IV is split into segments based on different radial distances from the school to the highway. Since we estimate the effect at one-third, two-thirds, and a mile, the instrument is split into four parts, namely one- third mile, two-thirds mile, a mile, and more than a mile. A school that is a third of a mile away from the nearest federal highway will have the specific distance in the sub- variable for one-third mile, with zeros in other sub-variables. This is tantamount to relaxing the assumption that the instrument has a unique linear relationship with the endogenous variable. Exogenous assignment to schools ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In addition to the instrumental variable method, we analyze a subsample of students who were exogenously assigned to different schools. This subsample of 2739 observations constitutes a small fraction of the total 1,352,696 observations in the main analysis sample, but provides a complementary analysis to the IV methods explained above. The strategy is similar to that pursued by Asirvatham et al. (2018) in their analysis of peer effects using the Arkansas BMI data. As explained in this earlier study, the exogenous reassignment was created by a court mandated restructuring of public schools. In 1992, the Lake View School District and other plaintiffs claimed that the school funding system was unconstitutional. In the Lake View School District No. 25 v. Huckabee case, the Arkansas Supreme Court ruled that the state educational funding was unconstitutional. To meet the court mandate, the State passed the Public Education Reorganization Act, Act 60, during the Second Extraordinary Session of 2003. School reorganization thus occurred to overhaul the public school funding system. Since the primary motivation was not to restructure the schools to improve students’ health and because the legislation was passed in a special session of the state legislature in response to the court ruling, we posit that this restructuring created an exogenous school assignment for the students it affected. There were two ways the State sought to comply with the court decision: either through consolidation or annexation for all districts with fewer than 350 students. Consolidation involved cutting administrative overhead by bringing schools and/or school districts under fewer management personnel. Annexation required physically closing school locations and have the children attend a nearby school that met the criterion set by the legislature. Thus, annexation created a change in school of attendance that was plausibly exogenous. An exogenous reassignment to other schools more directly addresses the issue of self- selection into schools. This reassignment, we argue, is largely uncorrelated with observables and unobservables that determine BMI outcomes. This is especially valid because the students are not in their original school attendance zone, which then largely rules out common factors or linkages between student BMI outcomes and restaurant location around their re-assigned school. In the results section, we discuss differences between annexed schools and reassigned schools.","In an effort to combat high rates of childhood obesity, the Arkansas General Assembly passed the Act 1220 of 2003. Among other things, this legislation mandated that public schoolchildren be assessed for BMI beginning in the 2003–2004 school year. BMI screenings have been ongoing since that time. We use these BMI screenings as the main data source in this study. The Arkansas Center for Health Improvement (ACHI) led the development and implementation of the state-wide BMI assessment process.13 ACHI developed a statewide protocol for standardized measurements across the state. Height and weight measurements are measured by trained personnel in schools and are reported to ACHI. The dataset we use includes BMI z-scores, race, gender and participation in free or reduced lunches. These data also include the weight status of each child, which based on reference growth charts from the Centers for Disease Control and Prevention (CDC). Only those students with at least two BMI observations are included in our analysis sample. Medical professionals and the Centers for Disease Control & Prevention (CDC) use BMI z-scores as a general indicator to track body weight progress among children. As a child grows, so does his/her body weight and height. The BMI z-score takes into account this growth and development, which also differs by sex. The BMI z-score provides a uniform measure of body weight across our sample that contains children that differ by age and sex. Our empirics are based on a panel dataset covering the years 2004–2010. One problem we confronted in assembling the data set is that state policy relating to the frequency of BMI measurement changed during our study period. From 2004 to 2007, the BMI of school children was measured annually for all grades. Thereafter, BMI was measured and reported only for children in even-numbered grades inclusive of kindergarten. Thus, we have BMI prevalence rates for students from all grades from 2004 through 2007, but only for students in even grades after 2007. The non- reporting of obesity prevalence in odd grades after 2007 should not bias our estimates, since the decision to stop measuring the BMI of children in odd-numbered grades was exogenous in that it was not made by the child, the child’s family, or the child’s school. However, this change in reporting does affect our ability to take into account BMI changes in a consistent fashion over time. Another source of data we bring into the analysis is location of food businesses. These are based on Dun & Bradstreet business lists and include restaurants, grocers, and other food stores. Details on the construction of the food environment are provided in the Appendix. Using Geographic Information System (GIS) software, we create measures of the commercial food environment around schools and residences. These variables measure the number and type of restaurants at varying radial distances from schools in the increments of a third of a mile up to a mile. Specifically, we measure the number of fast food restaurants, sandwich places and pizzerias. Measures of the food environment around a student’s residence include distance to the nearest grocery store, dollar store, convenience store, fast food restaurant, pizzeria, and sandwich place. Fast food restaurant counts may reflect local demand for fast food. To the best of the authors’ knowledge, all but one school had a “closed campus” policy during the study period. Seniors in this high school had the option of eating out during the lunch period with school permission. Students in the sample would not generally have sanctioned access to fast foods during the school day suggesting that access would be primarily before and after school hours. We also restricted the sample to those less than 18 years old. Annexed sample ~~~~~~~~~~~~~~ Table 1 presents the summary statistics of the entire population of students included in the analysis (called the general sample) and also compares it to the students who were reassigned to a different school (called the annexed sample). Students in the general sample differ from those students who were reassigned to a different school in certain characteristics. The BMI z-scores in the annexed sample are 0.05 standard deviations larger than the overall sample. There is a negligible difference across gender proportion. Annexed schools, however, had 27 percentage points more African American students and 21 percentage points less Caucasian students. The annexed sample also had a higher percentage of students qualifying for free lunches, 65% compared to 43% among all schools. In terms of the commercial food environment, students in annexed schools were further away from all of the measured food establishments. This sample also had slightly more rural students. The summary statistics generally reflect the communities that were affected by this school restructuring process. For instance, this reform closed smaller schools and integrated them into larger schools. Smaller schools are predominantly in the rural areas. Overall, differences in school-level demographic characteristics between the general sample and the annexed schools are markedly different in terms of race, household income, urbanization and access to some of the most important food establishments. One concern in the restructuring process is if students were reassigned to schools with similar characteristics or in similar neighborhoods. Asirvatham et al. (2018) assessed differences in the schools that closed and the schools to which students from the closed schools were reassigned. Data on children from the sending schools, i.e., annexed schools, constitute the second analytic sample we used in our analysis. This exogenous assignment addresses the issue of self-selection because, had it not been for the annexation, these students would not otherwise be attending a school outside of their original attendance zone. Asirvatham et al. (2018) found significant differences in percent free and reduced lunch between sending schools and receiving schools indicating that the schools differed in socioeconomic characteristics. However, by the nature of the school restructuring process and difference in student population characteristics, we can affirm that a student’s current BMI is unlikely to be correlated with unobserved factors at the new school.","Our basic model to estimate the effect of fast food restaurants surrounding schools on student BMI is:Yit = β0 + β1Rk(i)t + β2Xit + β3Xk(i)t + β4Fit + Uikt,where Yit is the BMI z-score of the ith student in time t; Rk(i)t is the restaurant counts around the kth school of the ith student at time t; Xit is a vector of student i’s characteristics; Xk(i)t is a vector depicting demographic student characteristics at the school-level in school k where student i attends, including proportion of students of different race categories and gender and percent qualifying for free and reduced lunch; Fit is the vector of commercial food environment near the residence of student i which includes distance to nearest grocery, dollar store, convenience store, fast food, pizzeria and sandwich place; and Ui is the error term which equals μi + εit, where μi is the unobserved time invariant component and εit is the spherical error term. Variables measuring different aspects of the commercial food environment around student’s school and residence were constructed using GIS software and geolocation of the schools, food establishments and student residence. To tease out the effect of distance of the fast food restaurants from schools, we developed three variables measuring the number of fast food restaurants within different radial distances from school, namely one-third of a mile, two-thirds of a mile, and a mile. Besides the number of restaurants at different radial distances, we also created two sets of variables with different combinations of restaurants. One included only fast food restaurants and the other included fast food, sandwiches and pizzerias. The panel nature of the data and the amount of information on students, schools and food environment allow us to control for individual and school-level characteristics, and also to use student-level fixed effects that further reduces the endogeneity bias due to omitted variables in the estimate. Since there could be year-to-year changes in a student’s BMI that if not accounted for might bias the estimates, we also estimate a two- way FE model by adding binary variables for different years in the data period. Thus, the two-way FE models include both student-level fixed effects and year fixed effects. Note that the observations are annual so year fixed effects would capture time-invariant year to year changes in the BMI that are not captured by the included right-hand side variables.","As discussed above, our research objective is to estimate the impact of fast food restaurants surrounding the schools on a child’s BMI z-score, and check if there are any differences in the effects by radial distance of restaurants from school. In this section, we first present results from the pooled OLS (top panel, Table 2) and student fixed effects models (bottom panel, Table 2). The pooled OLS shows that the effect, though likely biased, is very small or zero in magnitude and that the sign of the coefficient is always negative, which is inconsistent with our prior expectation that fast food exposure should be positively related to BMI. Importantly, healthier parents choosing healthier neighborhoods might create positive bias. On the supply side, restaurants locating in neighborhoods with a larger proportion of lower income families might at least show some restaurant effect. Pooled OLS, however, assumes all observations are independent. Our data have multiple observations on each student and therefore we are able to obtain within estimates by controlling for the student-level fixed effects. These results are in the bottom panel of Table 2. Some of the estimates are positive in sign but are very small in magnitude. Moreover, results are robust to inclusion of grade effects or additional controls for food stores around children’s residences. Controlling for the commercial food environment surrounding the residence of each student had little impact on the estimates once student-level time-invariant unobserved factors are accounted for. Even though we control for unobserved time-invariant factors at the student-level beyond the residential food environment, we cannot entirely rule out the self-selection bias. This could potentially bias the estimate because parents choose where they live. For example, an occupation requiring frequent long-distance travel might increase the likelihood of choosing a residence near highways. As previously discussed, our first identification strategy is to instrument for the restaurant counts. Results from fixed effects IV method are presented in Tables 3 and 4. The heteroskedasticity-robust first stage regression results in Table 3 indicate that average distance of schools from federal highways (the instrument) is a very significant and strong predictor of restaurant counts surrounding schools. The negative coefficient indicates that the number of restaurants decrease as the distance of the school from the highway increases. Large F-statistics are indicative of high correlation between restaurants counts and the instrument. On the smallness of the magnitude, it should be noted that more than 50% of the schools did not have restaurants within the certain defined radii at some point during the study period. The sign suggests that the farther the average distance of a school from nearest highway, the lesser is the number of restaurants locating near schools. The reduced form fixed effects regression also yielded significant estimates of the instrument. Despite the absence of fast food restaurants for a large number of schools, the magnitude at one-third radial distance suggests a 2.1 percent increase in BMI z-score, 0.0147 standard deviations from the mean 0.704 SD. At a two-thirds mile radial distance, the estimate decreased by a third and at a mile radius, the estimate further decreases by a fifth. The instrumental variables estimates show interesting results. First, the general result is that restaurants around any defined radii have at least some effect on student BMI. Second, within the same group of restaurants, the effect decreases as the radial distance increases. This reflects the fact that as a restaurant locates further away from a school, its influence on BMI begins to wane. This could indicate the reduced effect on calorie intake as distance from a school increases or that distance from school does matter – even though the effect is small. Thirdly, among different groups in general, the estimate is higher when the fast food measure excludes sandwich places and pizzerias. At a third of a mile radius, fast food restaurants alone show an increase of 0.024 standard deviations in BMI z-score compared to 0.015. This seems counterintuitive, but keep in mind that only about 50% of the schools had fast food restaurants. This relatively smaller magnitude could indicate that the effect of sandwich places and pizzerias might be much lower than the fast foods itself. Thus, the IV method finds a very modest effect of the restaurants surrounding schools on a student’s BMI. Our complementary analysis of the annexed sample is consistent with the findings above in that there is no evidence of a large fast-food effect on BMI (Table 5). It is much smaller. Both OLS and fixed effects estimates show no significant effect at any radial distance from the school. The observations on reassigned students constituted less than a quarter percent of the total observations used in the analysis above. The regression model used to the obtain the results in Table 3 include similar control variables to the those used in the panel IV models reported above. As discussed above, this sample is different from the general sample analyzed in several respects. Those differences could partly explain why no effect is observed in the annexed sample compared to only a marginal effect in the general sample. Compared to the general sample students in the annexed sample had to travel further away because the school in their attendance zone closed which may have prevented access to restaurants before and after school-hours. The OLS and the fixed effects estimates show high standard errors. The results presented so far analyze all students and include several covariates that might influence a student’s BMI. In the Appendix, we ran additional regressions to examine if the restaurant effect might vary by age or by geographic areas based on city centers or population densities. Specifically, specifications based on age group present estimates for elementary, middle, and high schools (Tables A1–A3, respectively). Estimates from regressions for rural and urban schools are presented in Tables A4 and A5, respectively. The results show some differences by age but nothing of economic significance. The main conclusions of the paper appear to hold across the different age categories. We also do not find any difference across rural and urban areas.","In this study, we estimate the effect of the number of fast food restaurants surrounding schools on the body mass index (BMI) z-score of school children. The study uses individual-level panel data from students attending public schools in Arkansas. The results from the two methods on two samples drawn from the same student population are similar. First, an instrumental variable strategy is employed using an instrument that has been validated by previous studies. A school’s distance from the nearest highway is used to instrument for the endogenous variable of interest, the restaurant count surrounding the school at specified radial distances. The identifying assumption is that fast food restaurants located by the highways cater to the highway travelers. We use the same instrument in two ways for efficient estimation. In the first specification, we use a single variable. In the other specification, we separate the instrumental variable into parts based on its distance from the highways. The same instrument is used in two different specifications, and both arrive at similar results. Second, we exploit a natural experiment that resulted when a number of Arkansas schools were reorganized in response to a state Supreme Court decision on school funding. This created an exogenous reassignment of students to schools other than the one in their attendance zone. Students in this exogenously assigned sample have different exposure to restaurants surrounding schools. In fact, given their distance and time to travel to a school that is located further away, it is unlikely that they would have had the opportunity to visit a fast food restaurant while taking the school bus. Thus, the two estimates, although complementary, have different implications. This study finds the effect of restaurants surrounding schools on childhood obesity to be negligible. Our point estimate, while very small, is highly significant. The smaller estimate might cultivate a laissez-faire attitude towards regulating restaurants around schools, but the findings reported here must be interpreted within the context of the study. Most Arkansas schools to had closed campus policies during the study period that meant children would only have access to outside restaurant food before or after school hours. Hence, replicating our study in other contexts to test the robustness of our findings would be warranted.","This research was funded by the Agriculture and Food Research Initiative of the USDA National Institute of Food and Agriculture, grant number 2011-68001-30014. This work was also partly supported by the National Research Foundation of Korea (NRF-2014S1A3A2044459); Research Council of Norway Grant #233800; the Arkansas Biosciences Institute (ABI); and the National Institute of General Medical Sciences of the National Institutes of Health under Award Number P20GM109096. The research reported in this study was approved by the Institutional Review Board, University of Arkansas, Fayetteville, USA under protocols # 10-11-235 and # 14-07-026. The procedures of this study are in accordance with ethical standards for human research."],["We examine effects of protein and energy intakes on height and weight growth for children between 6 and 24 months old in Guatemala and the Philippines. Using instrumental variables to control for endogeneity and estimating multiple specifications, we find that protein intake plays an important and positive role in height and weight growth in the 6-24 month period. Energy from other macronutrients, however, does not have a robust relation with these two anthropometric measures. Our estimates indicate that in contexts with substantial child undernutrition, increases in protein-rich food intake in the first 24 months can have important growth effects, which previous studies indicate are related significantly to a range of outcomes over the life cycle. --------------------------------------------------------------------------------","Inadequate child growth and weight gain are of paramount concern. Approximately 165 million children under five years old in developing countries are stunted and 100 million are underweight (Black et al. (2013)). Growing evidence indicates that early-life undernutrition is associated with, and likely in part causes, reduced education, adult cognitive skills, and wages (Grantham-McGregor et al., 2007; Engle et al., 2007, 2011; Victora et al., 2008; Hoddinott et al., 2008, 2013; Behrman et al., 2009; Maluccio et al., 2009). Despite widespread concern about early-life undernutrition there is limited systematic knowledge about production technologies for key outcomes, particularly height and weight, needed to inform more-effective program and policy design. This gap is partially due to inherent difficulties in modeling these complex biological and behavioral processes—often strong assumptions are required for estimation, so that it is difficult to make definitive conclusions. A major challenge in estimating production functions for height and weight is that inputs reflect behavioral choices. Using data from the same Philippine study analyzed in this paper, Akin et al. (1992) and Liu et al. (2009) find that families allocate nutrients to compensate for prior poor health. Where allocations reflect compensatory behaviors that are not controlled for in the estimation, the estimated effect of nutrients on growth can be biased. Another challenge is measurement error in inputs. Using related data from Guatemala, Griffen (2016) finds that estimates of energy effects on height are substantially larger using instrumental variables (IV) than with ordinary least squares (OLS) probably in part due to measurement error. In this paper, we examine relations between energy intake and: (1) linear growth and (2) weight gain. We use longitudinal data from Guatemala and the Philippines that includes detailed information on anthropometric outcomes, nutrition and other inputs collected at intervals of two-three months to estimate height and weight production functions for children in the critical age range 6–24 months. In our specifications, height and weight depend on lagged height and weight, energy intakes, breastfeeding, diarrhea, and individual fixed endowments. We combine individual fixed-effects (FE) with instrumental variables (IV) to control for both endogeneity and measurement error. This paper presents three important methodological contributions. First, we estimate production functions for two countries, Guatemala and the Philippines, and for two anthropometric measures, height and weight, which allows us to compare the robustness of our findings across different settings and anthropometric outcomes. Second, we improve on previous IV literature on growth by providing details of instrument selection and an assessment of how the results are robust to changes in the instrument set. We present estimates for numerous instrument combinations, putting emphasis on those judged more reliable based on over-identification and weak instrument tests. Third, in addition to considering total energy intake, which is the nutritional input usually considered in the economics literature, we disaggregate energy intake into two components: proteins and (all) other macronutrients (which we refer to as “non-proteins”, meaning fat and carbohydrates). This emphasis on dietary quality, highlighted by Arimond and Ruel (2004), is especially relevant because it may help design interventions that better reduce stunting and underweight. We find robust and positive effects of proteins on height and weight growth. Energy from other macronutrient consumption (non-proteins), is not systematically related to these anthropometric measures, which suggests that protein-rich foods are particularly important for growth of undernourished children. Input selection ~~~~~~~~~~~~~~~ Our choice of inputs is guided by Black et al. (2008) who argue that inadequate diet and disease are the main immediate causes of stunting and wasting. With respect to diet, two energy sources have been identified as being especially important for child growth: proteins and non-protein energy from other macronutrients. Infants require certain minimum amounts of energy and proteins to maintain long-term good health but these requirements are heterogeneous and depend on several factors including weight and whether the child is breastfed (FAO, 2001; WHO, 2007). Children's energy requirements are partly driven by energy costs of linear growth, which has two components: (1) energy needed to synthesize growing tissues and (2) energy stored in these tissues (FAO, 2001). These comprise approximately one-third of total energy requirements during the first three months of life, but despite increasing in absolute terms they decline to only 3% by age 24 months, in part because overall energy requirements increase substantially with body size. Proteins are needed to balance nitrogen loss, maintain the body's muscle mass, and fulfill needs related to tissue deposition (WHO, 2007). There is also evidence from research on animals that protein provides anabolic drive for linear bone growth (WHO, 2007).2 To study the relative importance of protein and non-protein sources, we first examine the relationship between total energy and height and weight and then consider the potential for separate roles of the two at once in a single growth model. The comparison of proteins with non-proteins highlights the relative importance of proteins in children's diets and informs what types of interventions might have greater impact on height and weight.3 There is a limited literature focused on the distinction between total energy and protein energy. Pucilowska et al. (1993) find that high-protein supplementation in Bangladeshi children with shigellosis, a severe bacterial disease, increased weight compared to normal protein diets. A randomized evaluation for children up to 2 years of age in several European countries demonstrated that receiving baby formula with high protein content (% calories from protein) increased weight, but not height (Koletzko et al. (2009)). Both of these study populations, however, are different from the ones we examine. The Bangladeshi sample is restricted to children recovering from shigellosis while the European sample had not experienced the same nutritional deficiencies found in our samples. Using a sample more similar to ours, Moradi (2010) finds that access to high-quality protein, such as from livestock farming, better predicts height in some African countries than other energy sources. Similarly, Baten and Blum (2014), using global information for the first part of the twentieth century, that includes Guatemala and the Philippines, also find that local availability of cattle, milk and meat were an important predictor of adult height.4 A related issue is protein quality. Proteins are composed of amino acids with specific cell functions, and amino acid content defines protein quality. For instance, plant-based proteins lack essential amino acids unlike animal-based proteins (Dewey, 2013). In addition, plant-based diets have high levels of phytic acid, which might inhibit zinc absorption (Gibson, 2006), and zinc plays a key role in cellular growth and differentiation (Imdad and Bhutta, 2011). For animal-based protein, Mølgaard et al. (2011) argue that dairy intake has positive impacts on child growth. Although the mechanism is not entirely clear, this may be due to the stimulating effect on plasma insulin-like growth factor (IGF-1) (Michaelsen, 2013). Breastfeeding is another critically important source of nutrition in early life (Black et al., 2013). In this paper, we have data on breastfeeding status but not on the amount of breast milk consumed. Thus, our energy intake measures exclude energy from breastmilk requiring us to control for breastfeeding status in the models. Among diseases that affect growth, Walker et al. (2011) suggest that persistent diarrhea and other diseases can have long-lasting effects on children's physical development. Therefore, in our analyses, we incorporate diarrhea as an input, as it is considered a major contributor to stunting, wasting and child mortality (Black et al., 2013). Height and weight production functions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The main challenges for estimating height and weight production functions include the endogeneity of inputs and measurement error (Behrman and Deolalikar, 1988). To overcome these, we follow the general approach developed in recent research on production function estimation for cognitive and non-cognitive skills (Todd and Wolpin, 2003, 2007; Cunha and Heckman, 2007). Because they include individual fixed effects and the entire input history, Eqs. (1) and (2) are difficult to estimate. For example, if inputs are treated as endogenous and an IV approach were used, it would be necessary to have at least one instrument for each period in the entire input history. Thus, instead of directly estimating these two equations, we make two further assumptions that allow less demanding specifications in terms of data and instrument requirements, while remaining more flexible than previous specifications in the literature. Effects of past inputs follow a monotonic (likely decreasing) pattern at a constant rate γ for each period.5 That is: βt−j = γβt−1−j and δt−j = γδt−1−j. The coefficients on inputs in the height function are the same as those in the weight function, up to a multiplicative constant δt−1−j = ((1 + σ)/α)βt−1−j. Together, these assumptions reduce the set of endogenous variables to a tractable number, thereby reducing the number of required instrumental variables. Under these assumptions, height growth can be expressed as a function of current inputs, past height and weight, and an error involving current (t) and previous period (t − 1) shocks. Current inputs enter directly; the full history of past inputs enter indirectly through the lagged height and weight. As with the change-in-height Eq. (3), the change-in-weight Eq. (4) depends on current inputs, past height and weight, and an error including current and previous period shocks.6 This framework forms the core of our approach to estimating production functions for height and weight. Estimation of Eqs. (3) and (4) allow recovering β0 from Eq. (1) and δ0 from (2). Estimation and identification ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Although differencing removes individual-level fixed effects and thus controls for important sources of potential bias (unobserved persistent heterogeneity including, e.g., genetic endowments and fixed parental and household characteristics), to consistently estimate the parameters in the relations for change in height (Eq. (3)) and change in weight (Eq. (4)), we still need to overcome several endogeneity problems. First, by construction previous height and weight are correlated with the error terms of Eqs. (3) and (4) (see Eqs. (1) and (2)). Moreover, if we assume that the household responds to past shocks as is likely and for which there is evidence for the Philippines (Akin et al., 1992; Liu et al., 2009), current inputs may be correlated with the error terms. We address potential endogeneity by using IV, which also addresses bias due to random measurement error in x under the assumption that the instruments are uncorrelated with that measurement error. The set of candidate instruments we use differs by country but draws on plausibly exogenous factors including a randomized intervention in Guatemala and prices of common foods in both countries. We treat market prices as exogenous to households (as in Liu et al. (2009)). Using prices as instruments for inputs is a well-established approach in the estimation of production functions (Todd and Wolpin (2003)). We also include past height and weight measures, hi,t−2 and wi,t−2 as instruments to help identify the effects of lagged height and weight. (Instruments are described in further detail in Section 3.3.) Using the available instruments, we endogenize protein and non-protein intakes, as well as lagged height and weight. However, we do not have access to instruments in both countries that also would allow us to control for the potential endogeneity of breastfeeding or diarrhea.7 Controlling for individual-level fixed effects is an important aspect of our approach, however, and goes part way toward addressing their potential endogeneity. For example, fixed effects control for the possibility that certain children have a pre- disposition for diarrhea, or live in particularly unsanitary households. However, if households change breastfeeding practices when health shocks affect their children's health or change sanitary conditions to reduce the diarrhea prevalence, the estimated effects of breastfeeding and diarrhea could be downward-biased. For instance, households that have increased breastfeeding could be compensating for negative health shocks, suggesting a negative relationship between growth and breastfeeding, while correcting for endogeneity could show a positive relationship (and similarly for diarrhea). Because our principal objective is to study the roles of proteins and non-proteins in the production functions, however, we do not emphasize the coefficients for diarrhea and breastfeeding but instead make clear the assumptions under which our primary coefficients of interest are consistently estimated even if breastfeeding or diarrhea are endogenous in the model. Our estimation approach is consistent provided the instruments are not correlated with the error term in the production function, conditional on breastfeeding and diarrhea as well as other covariates mentioned below. This is plausible for the same reason that the instruments are exogenous in relation to the energy inputs, e.g., that they are not correlated with individual-level time-varying health shocks.8 In principle, there also could be interactions among inputs in the production function, such as between nutrient intakes and diarrhea, or between breastfeeding and other nutrient intakes but a specification incorporating such interactions would be even more challenging to estimate, requiring additional instruments. Given that there are already four variables that we treat as endogenous in our main models (protein, non-protein, lagged height, and lagged weight), we do not estimate models with such potential interactions; instead, we studied possible interactions by splitting the sample. For instance, to examine whether diarrhea or breastfeeding interacts with diets, we estimated specifications for the sample that is breastfed and compare the results with the sample that is not breastfed. We carried out a similar exercise for diarrhea. Our results indicate that coefficients are not affected when we separate the sample by breastfeeding types. For diarrhea, there was some evidence of interaction effects, where diarrhea lowers the effects of macronutrients, but because most of the specifications suffer from problems of weak instruments, we are unable to draw strong conclusions. The estimation of the growth equations also includes an indicator for whether the child was female, number of days since the previous measurement, and age and age squared at time t. Our methods permit us to improve upon the previous literature that investigates the effects of total energy on anthropometrics. Since we do not have a single set of preferred instruments, we are able to robustly study effects of total energy on height and weight across two settings. We do this estimating the changes in height and weight, first using total energy intakes and then separating protein and energy from other macronutrient intakes to examine their relative partial effects in each model.","Estimation of (5) and (6) requires high-frequency longitudinal data in early life that contain information on the outcomes (height9 and weight) and inputs (proteins and other macronutrients, breastfeeding, and diarrhea), as well as plausibly exogenous instruments. We now describe the data and contexts for two unique studies that fulfill these substantial requirements relatively well, one in Guatemala from the 1970s and the other in the Philippines from the 1980s. Guatemala ~~~~~~~~~ We use data from The Institute of Nutrition of Central America and Panama (INCAP) 1969–1977 nutritional supplementation trial. Four rural villages from eastern Guatemala were selected, one relatively large pair (∼900 residents) and one smaller pair (∼500 residents). At the outset, the villages were similar in terms of child nutritional status, measured as height at age three years, and were highly malnourished with over 50% of children severely stunted, i.e., with height-for-age z-score <−3. One large and one small village were randomly selected to receive a high-protein supplement (Atole); the others received an alternative supplement devoid of protein (Fresco). A 180 ml serving of Atole contained 11.5 grams of protein and 163 kcal. Fresco had no protein and a 180 ml serving had 59 kcal. The main hypothesis was that increased protein would accelerate mental development; additionally, it was expected that the high-protein nutritional supplement would affect physical growth. The nutritional supplements were distributed in centrally- located feeding centers in each village (Habicht et al., 1995). Virtually all (>98%) families participated (Martorell et al. (1995)). From 1969 to 1977, anthropometric measures (height and weight) were taken every three months for all children 24 months of age or under (including newborns entering the study) in the four villages. This yields a maximum usable sample for our analyses of 878 children measured at least twice by the age of 24 months. The amount of supplement intake was recorded daily in all villages. Home dietary information was collected every three months, including the types and amounts (except for breastmilk) of all foods and liquids consumed. These dietary histories were based on a 24-h recall period in the larger villages and a 72-h period in the smaller villages (from which we construct daily averages), and permit calculation of protein and non-protein intakes for the 24-h period by summing the nutritional content for each food item. The survey recorded the total months a child was breastfed. Nutrients from breastfeeding were not included in the nutritional intake calculations. Retrospective information on illness, specifically the length in days of episodes of diarrhea and fever, was collected semi-monthly. The Philippines ~~~~~~~~~~~~~~~ We use the Cebu Longitudinal Health and Nutritional Survey, a survey of Filipino children born between May 1983 and April 1984 in 33 rural and urban communities (barangays) in Metropolitan Cebu. The baseline survey included 3327 women sampled at a median of 30 weeks of gestation, and yielded a sample of 3080 singleton live births. This sample also exhibits high levels of undernutrition; at age 24 months, 62% of the children were stunted and 32% underweight. During the first two years of each child's life, data were collected every two months. This included anthropometric measurements, 24-h dietary recall of types and amounts (except breast milk) of all foods and liquids eaten, breastfeeding, and recent illness history. For breastfed children, the survey also collected the frequency and length of time spent breastfeeding. Total protein and energy intakes were calculated from foods consumed the previous day (24-h recall method). At each survey, mothers reported whether the child had diarrhea in the past 24 h, and if so, when the episode began, and the number of days the child had diarrhea during the previous week (Adair et al., 2011). The maximum usable sample of children between 6 and 24 months of age for the Philippines is 2713. Variable construction ~~~~~~~~~~~~~~~~~~~~~ Linear growth and weight gain are calculated as the difference between consecutive measurements. Although measurements were scheduled at specified intervals (every three months in Guatemala, every two in the Philippines), there were deviations including instances where a scheduled measurement did not occur. Because children experience high growth and growth spurts during the first two years of life, even differences of several days can be associated with significant differences in growth. We account for this by controlling for the exact number of days between measurements. Ideal data for this analysis would have information on protein and non-protein intakes over the entire period between measurements, but even in these uniquely comprehensive studies such detailed information is not available. Therefore, we approximate intakes over the entire period by using the average of the 24-h intakes calculated from the dietary recall information at the beginning and end of each period (which decreases measurement error relative to using only one point in time) multiplied by the exact number of days between measurements. For Guatemala, we add to this figure the intakes from the supplement (which were measured daily throughout the period) to obtain total protein and other intakes (as well as their sum, measured as total energy).10 For breastfeeding, we create a dummy indicator for whether the child was breastfed in the month previous to measurement at time t. While this does not fully exploit the detailed information available for the Philippines, it is done to have similar specifications across countries. The final input we include is diarrhea. For Guatemala, the protocol was to collect information every 15 days, so it is possible to construct the number of days experiencing diarrhea for the complete periods between anthropometric measurements.11 For the Philippines, it is only possible to construct the number of days with diarrhea during the week previous to each bimonthly anthropometric measurement. To extrapolate this to the full period between measurements, we estimate a count model for number of days with diarrhea for each two-month period with the Guatemalan data and use the estimated parameters from that model to predict number of days each Filipino child had diarrhea in each two-month period.12 As outlined in Section 2.3, in our main specifications we instrument for protein, other macronutrient intakes, and lagged height and weight. We now describe in detail the other instruments besides twice lagged height and weight. In both countries we use unit prices for various food items, selected with emphasis on foods with high protein content and/or important in the local diet. For Guatemala, prices are averages of national-level prices measured during December each year. We use lagged prices of eggs, chicken, pork, beef, dry beans, corn, and rice. Unit price variables for Guatemala are deflated and measured over the eight-year study period. For the Philippines, we use community-specific prices collected as part of the broader study. Between January 1983 and May 1986, enumerators visited two stores in each community, every other month, and collected prices (and quantity units) for a list of items. Not all items, however, were sold at each store at each visit. Consequently, there is not a complete set of prices for each item from each store (or even from each community in instances where no price was available from either store) in each measurement period. We selected as instruments the prices of dried fish, eggs, corn and tomatoes since these are the ones with the highest frequency in the sample.13 We use both current and lagged prices of those selected food items. By estimating a large set of instrument combinations, our approach does not depend on any one particular price, avoiding subjective instrument selection. For Guatemala, we also exploit the experimental variation resulting from the randomized allocation. We use a dummy variable that indicates whether the village had a feeding center that provided the high-protein supplement. We also interact this indicator with the distance from the home of the child to that feeding center. While the presence of a randomized allocation of a high-protein supplement provides an important source of exogenous variation, since there are four endogenous variables, additional instrumental variables also are used, i.e., twice lagged anthropometrics and food prices. For the Philippines we rely on price variation, which, unlike the annual Guatemalan food price data, varies both within-years and spatially, with information on these food items for the majority of measurement periods and each of the 33 communities. Descriptive statistics ~~~~~~~~~~~~~~~~~~~~~~ Over the period from ages 6 to 24 months, each Guatemalan child is observed an average of 4.3 times and each Filipino child 9.1 times. The sample we describe includes all observations (measurements of children at different ages) with complete information for the following variables: change in height between consecutive measurement periods (linear growth), change in weight between consecutive periods (weight gain), total energy, energy from protein, energy from non-protein, breastfeeding indicator, and days with diarrhea.14 The final number of observations used in each specification varies depending on the availability of the instrumental variables used in that specification, since instruments for some observations are missing. Table 1 compares the main variables for both samples. On average and at all ages, the Filipino children in the early 1980s were taller than the Guatemalan children in the 1970s. For example, at 12 months of age, Filipino children were on average 70.7 cm tall, while their Guatemalan counterparts were 1.8 cm shorter. In terms of average weight, however, there were no significant differences between countries—at 24 months, children from both countries averaged 9.8 kg. 44% of the Guatemalan children were stunted, and 27% underweight. The corresponding levels were lower, 25% and 11%, for Filipino children. In 2011 for low- and middle-income countries, average levels of stunting were 28% and of underweight 17%, and 36% and 18% in Africa (Black et al., 2013). With broadly similar levels of stunting and underweight, thus, our historical samples remain relevant to understanding undernutrition in many countries and regions. Table 2 shows that Guatemalan children appear more likely to have been breastfed at all ages. In both countries, breastfeeding declines with age. At six months, 99% of Guatemalan children were breastfed, while at 24 months only 18% were; the proportions were 76% and 14% for Filipino children. Patterns between diarrhea and age are less clear. In Guatemala, average number of days with diarrhea (per 3-month measurement period) increases with age to 15 months, after which it declines. Levels are relatively lower in the Philippines, fluctuating between about 2 and 6 days (per 2-month period), with no clear age pattern. For Guatemala, information is complete on all of the instruments except the distance to the feeding center, which is missing for ∼5% of observations. For the Philippines, on the other hand, incomplete price availability leads to larger reductions in the sample size. The potential sample has 24,820 child-age observations; the lagged price of corn, which is the most complete, has 18,710 observations and the lagged price of tomatoes, the least complete, has 16,084 observations. Overview ~~~~~~~~ We estimate height and production functions for children 6–24 months, the period widely considered to be a critical window for post-birth nutritional investment.15 We use Generalized Method of Moments (GMM) for exactly-identified models and Limited Information Maximum Likelihood (LIML) for over-identified models because the latter allows for smaller finite-sample bias (Stock and Yogo, 2005). As noted, we cluster error terms at the individual level to take into account correlation of individual error terms and serial correlation (Baum et al., 2007).16 We first estimate height and weight production functions using only total energy (i.e., the sum of calories from protein and other sources), then we analyze separately the roles of proteins and non-proteins. In all specifications, the energy intakes, lagged height, and lagged weight are treated as endogenous, and we control for breastfeeding, number of days without diarrhea since the previous measurement, child sex, number of days since the previous measurement, and age and age squared. Because there are many potential instrument combinations, to establish general results that do not depend on one specific instrument combination, we estimated large subsets of all possible combinations. For Guatemala we first restricted the instrument sets to combinations that always had the Atole experiment indicator. Then, we systematically varied inclusion of distance interactions with Atole indicator, second lags of height, second lags of weight, and from two to four of the seven food prices (eggs, chicken, pork, beef, rice, beans and corn). For the Philippines, we systematically varied inclusion of second lags of height, second lags of weight, and from two to six of the eight (four current and four lagged) food prices (eggs, fish, tomatoes and corn). A summary of our instrument combinations is found in the Data Appendix Section 6. For Guatemala, there are 546 specifications (i.e., each with a different instrument set) for the version of the model with total energy (Eq. (5)) and 525 when proteins and non- proteins are included separately (Eq. (6)).17 The total number of specifications estimated for the Philippines is 602 for both models. For each specification, we calculate the robust versions of the Hansen-J (HJ) over-identification test, the Anderson–Rubin under- identification test (Anderson and Rubin, 1949), and the Wald F-statistic (robust Cragg–Donald or CD statistic) to detect weak instruments. Since our main models have four endogenous variables and we estimate them assuming heterokedasticity, it is not possible to compare CD statistics with critical values from Stock and Yogo (2005). The robust versions of these tests were developed in Kleibergen and Paap (2006). We also calculate for each endogenous variable Angrist and Pischke's (AP) partial F (Angrist and Pischke, 2009), which are informative about the presence of weak instruments. Finally, for all over-identified models we calculate the Hausman test of equality of OLS and IV estimates. Since each production function is estimated multiple times, we explore distributions of estimated coefficients rather than a single or small set of “preferred” specifications, allowing us to draw more general conclusions. We do not choose or define a preferred specification because there are no obvious criteria for doing so and because of the concern that any potential preferred specification would not be robust to changes in the set of instruments. Although a priori the instruments we propose are plausibly exogenous and strong, we put relatively more confidence in those instrument sets that better satisfy over-identification and weak instrument tests. The results of each type of specification are presented in Tables 3–6 and Figs. 1–3. In Tables 3 and 5, and Fig. 1, we present the estimated overall energy coefficients. In Tables 4 and 6 (Panels A and B), and Fig. 2, we present the estimated protein coefficients, and in Tables 4 and 6 (Panels C and D), and Fig. 3, the estimated non-protein coefficients. Each table presents the 25th, 50th and 75th percentiles of the estimated coefficient distributions and, in the final two columns, the percentages of the coefficient estimates that are significantly (p < 0.05) positive or negative. For each Panel in each table, the first row reports distributions for all estimated specifications and, in subsequent rows, for specifications that are over- identified, and for those that have HJ P-values > 0.05 and CD statistics > 1, 3, or 7 (provided there are more than 10 such specifications in each case).19 These sets of specifications focus on results for which relatively strong and exogenous instruments are available. Figs. 1–3 present point estimates (and associated 95% confidence intervals) for all specifications that have HJ P-values > 0.05 and CD > 1 (corresponding to the third rows in Tables 3–6). The scale of the x-axis corresponds to the natural logarithm of CD statistics and the y-axis the coefficient values.20 To facilitate interpretation of the coefficient magnitudes, we simulate changes in height and weight when energy intakes increase ceteris paribus For this exercise, we use the most restrictive specifications with CD > 7 (or CD > 3 if there are fewer than ten specifications with CD > 7) and HJ P-values > 0.05. Within that set of specifications, we select the median coefficient and simulate effects of increasing energy intakes by 300 kcal per day, protein intakes by 10 g per day, or non-protein intakes by 250 kcal per day. Each of these is approximately one SD of respective intakes of 18-month old infants in both countries. This hypothetical daily increase is then multiplied by 90 in Guatemala and by 60 in the Philippines to approximate total intakes for a given measurement period, and then multiplied by corresponding coefficients to obtain anthropometric changes. We call this exercise median prediction. Guatemala ~~~~~~~~~ Table 3 summarizes for Guatemala distributions of coefficient estimates on total energy in the height and weight equations, and Fig. 1A and B show the coefficients and confidence intervals for the corresponding specifications with CD > 1. Total energy positively affects height and weight changes. These positive relationships are most evident for specifications with relatively stronger and more exogenous instruments. Our findings are consistent with previous literature that uses stronger identification assumptions estimating similar relationships from the same data sources (Habicht et al., 1995; Griffen, 2016). For height in Guatemala, estimated coefficients on total energy are positive in the vast majority of cases, positive and significant (p < 0.05) in 35% of cases, and never negative and significant. The positive relationship is more robust when we consider specifications with relatively stronger and more exogenous instruments, according to the tests. Restricting to over-identified specifications in which HJ P-values > 0.05 and CD > 3, total energy coefficient estimates are positive and significant 57% of the time. To provide further interpretation of the magnitude of the coefficients, we calculate the median prediction (Section 4.1), taking the median coefficient of the specifications with CD > 3; we calculate the effect of increasing energy per day by 300 kcal. For Guatemala, this implies a 0.62 cm predicted change in height. For weight production functions, estimated coefficients on total energy are positive and significant for 36% of specifications, and are never significantly negative. Specifications with higher CD statistics have larger proportions of positive significant coefficient estimates. Fig. 1B shows that while there are fewer specifications with higher CD statistic levels compared to the height model, for those with stronger instruments, the estimates are generally positive. The median prediction exercise indicates increasing energy intake by 300 kcal per day yields a predicted 620 g change in weight. Next, we consider the roles of protein and non-protein energy separately in the growth model. Proteins robustly and positively affect growth in height and weight in Guatemala, but the relationship of non-proteins (after controlling for protein) with these anthropometric measures is non-positive. Panel A of Table 4 (and Fig. 2A) shows that for 53% of all specifications, protein coefficient estimates are positive and significant. In specifications with CD > 3, the estimates are always positive and significant. In specifications with stronger instruments, the estimated coefficient dispersion (i.e., the distance between the 25th and 75th percentiles) decreases; for specifications with CD > 1 the ratio of the coefficients in the 75th and 25th percentiles is 1.3, while for the specifications with CD > 3 the ratio is 1.06. Our median prediction exercise indicates that if protein were to increase by 10 g per day, the predicted change in height is 0.39 cm. For weight change (Panel B of Table 4 and Fig. 2B), we find an even more robust pattern for proteins. In nearly all specifications (92%), protein coefficient estimates are positive and significant, and for specifications with CD > 1, they are always positive and significant. For all specifications, the estimate at the 75th percentile is only 1.2 times larger than that at the 25th percentile. This pattern of stability and significance of coefficient estimates also can be seen in Fig. 2B where the dispersion of the estimated coefficients is small, and there is a clear pattern of positive and significant effects of protein intake on weight growth. An increment in protein intake of 10 g per day results in a predicted 195 g change in weight. By contrast, there is little evidence that energy from non-proteins affects changes in height and weight. Panel C in Table 4 and Fig. 3A show that for Guatemala, in nearly all cases (98%) the estimated coefficient is insignificant in the height model. For the weight production function (Panel D of Table 4 and Fig. 3B), the point estimates are never significant. Philippines ~~~~~~~~~~~ Table 5 shows the distribution of the total energy coefficient estimates for the Philippines and Fig. 1A and B the corresponding coefficients and confidence intervals for specifications with CD > 1. As in Guatemala, positive relations are most evident for specifications with relatively stronger and more exogenous instruments. The positive impacts of total energy on height and weight are consistent with those found under somewhat stronger identification assumptions and using the same data, by Liu et al. (2009) and de Cao (2015). Across all specifications summarized in the Panel A of Table 5, 13% have positive and significant coefficient estimates (p < 0.05), while none have negative and statistically significant estimates. Restricting results to the 45 specifications with HJ test P-values > 0.05 and CD > 7, 64% of estimated total energy coefficients are positive and significant. Specifications with higher CD statistics tend to have more concentrated coefficient estimate distributions. If daily energy intake increases by 300 kcal the predicted change in height is 0.18 cm. For weight, evidence is similar regarding the role of total energy. The bottom panel of Table 5 indicates that for 15% of all the specifications in the Philippines, the estimated coefficient on total energy is positive and significant and never negative and significant. Specifications with the highest CD statistics tend to have larger shares of positive and significant coefficient estimates. Our median prediction results in a predicted change in weight of 37 g. Panel A of Table 6 (and Fig. 2A) shows that for 39% of all specifications, protein coefficient estimates are positive and significant. While there are fewer specifications with strong instruments than in Guatemala, for specifications with CD > 3, 100% of the coefficient estimates are positive and significant. In specifications with stronger instruments, the estimated coefficients dispersion decreases. Increasing protein consumption by 10 g per day is predicted to result in a 2.24 cm change in height. For all specifications (Panel B of Table 6 and Fig. 2B), 48% of estimated coefficients on protein for weight are positive and significant – 100% in specifications with CD > 3. Similar to Guatemala, coefficient estimate dispersion decreases with stronger instruments. Increasing protein consumption by 10 g per day results in a predicted 703 g change in weight. Somewhat surprisingly, non- protein intakes are generally negatively related to both height and weight gain. For height, Panel C of Table 6 reports that 88% of the specifications with the strongest instruments (CD > 3) yield negative and significant estimated coefficients. For weight, 100% of estimates in specifications with the strongest instruments are negative and significant. These findings for non-protein energy for the Philippines are somewhat counter-intuitive, because they suggest that such energy intakes are detrimental to growth. Most individual foods (including those consumed in these regions during the study periods), however, include both proteins and non-proteins and virtually all diets do. Consequently, it is unlikely that actual intakes would change in a fashion that increased energy from non-proteins while simultaneously holding proteins constant. Since Filipino children's diets included both intakes, on net any negative effects of other macronutrient sources would have been partly or fully offset by protein effects. For example, not including breastmilk, at age 6 months, 93% of children had some protein consumption and from ages 14 to 24 months, all did. Moreover, at age 6 months 75% of children are breastfed, which also provides protein intakes. In Section 4.5, we show that the model predicts that a dietary change (relatively rich in proteins but with some energy from other sources) indeed has positive effects on height and weight, despite negative coefficient estimates on non-proteins. There are several potential explanations for the finding that non-proteins are less robustly related to anthropometrics than proteins. First, it is possible that energy from macronutrients other than proteins do not affect height and weight, at least aggregating the other macronutrients as we do. Second, it may be that non-linearities are not captured. For instance, it could happen that carbohydrates and fat need some proteins to have an effect on anthropometrics—if protein intakes are zero or very low, other intakes would not affect height and weight. Third, dietary changes after children stop breastfeeding can result in poorer quality diets, especially poor quality of carbohydrates and low micronutrient density, weakening any potential link to anthropometrics. Fourth, the available instruments simply may not be powerful enough to detect effects of other macronutrients; protein and non-protein intakes are highly correlated (even before instrumentation), making it difficult econometrically to identify their distinct effects; in that sense, Guatemala greatly benefits from the experimental Atole intervention, which provides a clear and strong exogenous variation for protein, though it is less powerful for other macronutrients. Effects of other inputs and controls ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In addition to the different nutrition intakes, our analysis provides estimates of the coefficients on lagged height, lagged weight, breastfeeding, and diarrhea. The results clearly indicate some catch-up height and weight growth. The lagged height coefficient is consistently negative and mostly significant in the change-in-height equation, indicating that shorter children at the end of one period tend to grow more in the next period. Similarly, the lagged weight coefficient is consistently negative and mostly significant in the weight equation so that lighter children at the end of one period gain more weight in the following period. With the caveat that the estimates for breastfeeding and diarrhea are potentially biased due to endogeneity, our coefficient estimates for number of days without diarrhea are consistently positive and significant for weight in both samples, suggesting that diarrhea has detrimental effects on weight gain as generally found in the literature. The coefficient estimates for breastfeeding are positive and mostly significant for Guatemala. In the Philippines, the coefficient estimates generally show a positive association between breastfeeding and height while the associations between breastfeeding and weight show no consistent pattern, similar to findings from Adair and Popkin (1996).21 Counterfactual exercise: increasing nutritional intakes ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We next simulate the full effects of additional protein and non-protein intakes on child height and weight for the Philippines, complementing the simpler median predictions we used when interpreting individual coefficients. From the set of specifications with HJ P-values > 0.05, we select the specification with the highest CD. The simulation is based on adding one egg per week to a child's diet, assuming no other changes in diet and no change in diarrhea. Eggs are good for such simulations. They were widely available in the localities where these studies are situated and are easily consumed by infants. They not only contain highly bioavailable protein, but also contain energy from other macronutrients, similar to many other naturally protein-rich foods. A medium (44 g), whole raw egg contains on average 5.5 g of protein and 40.9 calories from non-protein.22,23 Based on our parameter estimates, a child who consumed an additional egg per week on top of existing diet, for 18 months – from 6 to 24 months of age – would gain an additional 0.72 cm in height and 265 grams in weight.","Arimond and Ruel (2004) described associations between children's dietary diversity and their height. We build on their insights, examining effects of diet and particularly diet composition on height and weight growth for children between ages 6 and 24 months, giving special attention to differences between diets rich and poor in proteins. We improve upon previous literature by making weaker identifying assumptions, considering two important anthropometric measures—height and weight, investigating the robustness of our results to the use of a number of different instruments, and separately investigating the effects of energy from proteins and from non-proteins while controlling for breastfeeding and diarrhea. We take advantage of two rich databases, one for Guatemala and the other for the Philippines, which have longitudinal information on height, weight, and protein and energy intakes with high frequencies of observations. IV estimation strategies are used to overcome endogeneity and measurement error problems, using food prices and, in the case of Guatemala, a randomized nutritional intervention, as instruments. Because there are many instruments and instrument combinations available, we present results that comprehensively summarize these combinations rather than selecting only a single set of instruments. Our findings indicate that increasing energy intake increases both height and weight in both countries. But the source of that energy, protein versus non-protein, matters. In these poor populations characterized by high levels of chronic undernutrition, increases in protein intake drive increases in child height and weight. These results provide evidence on an important puzzle in the literature while pointing to possible modifications to interventions designed to improve children's nutritional status. A systematic review by Manley et al. (2013) using meta-analysis techniques shows that while the average impact of income transfers from social protection programs on height-for-age is positive, effect sizes are small and not statistically significant. If households use these transfers largely to increase the quantity of calories consumed, if the increases in protein consumption is small in magnitude, or if these proteins are not allocated to children, then our results suggest that such transfers will have little impact on child height—precisely what Manley et al. (2013) find. Headey and Hoddinott (2015) examine impacts of Green Revolution-induced increases in rice productivity on children's anthropometric status. They find no impact of these on child height, results also consistent with what we observe here. Our findings, in conjunction with these other studies, suggest that interventions designed to increase household incomes may only improve children's nutritional status when they are linked to mechanisms that also improve the quality of children's diets. Such interventions, e.g., linking nutritional behavior change communication to social protection interventions or “nutrition-sensitive agriculture” await further study.","The authors thank Grand Challenges Canada (Grant 0072-03), Bill and Melinda Gates Foundation (Global Health Grant OPP1032713), and the Eunice Shriver Kennedy National Institute of Child Health and Development (Grant R01 HD070993) for Financial Support. The funders have no involvement in the analysis and interpretation of the data, writing of the paper, or the decision to submit the paper for publication.","There are no conflicts of interest."],["Obesity and overweight are spreading fast in developing countries, and have reached world record levels in some of them. Capturing the size, patterns and trends of the problem has, however, been severely hampered by the lack of comparable data in low and middle income countries. We seek to begin to fill this gap by testing several hypotheses on the determinants/correlates of overweight among women, related to the influence of economic and technological development. We undertake econometric analysis of nationally representative data on about 878,000 women aged 15-49 from 244 Demographic and Health Surveys (DHS) for 56 countries over the years 1991-2009. Our findings support most previously expressed hypotheses of what might explain obesity patterns in developing countries, but they also reject some prior notions and add considerable nuance to the emerging pattern. © 2014 The Authors. --------------------------------------------------------------------------------","Obesity has become a global phenomenon not solely confined to rich countries (James, 2008; Popkin, 2007), with the total number of overweight and obese people being estimated at about 1.5 billion in 2008 worldwide (Popkin et al., 2012). However, the accurate description of its size, trends and socioeconomic patterns has been compromised by the scarcity of globally comparable data. One recent study (Finucane et al., 2011) attempted to fill this void by estimating data on mean body mass index (BMI) among adults over the age of 20 living in 199 countries (analyzing 960 country years worth of data on 9.1 million participants). They found that between 1980 and 2008, mean BMI in all countries combined increased by 0.4 (0.5) kg/m2 per decade for men (women). In order to study the trends and determinants of obesity and overweight specifically in the developing countries, a small set of recent studies (Martorell et al., 2000; Mendez et al., 2005; Mendez and Popkin, 2004; Monteiro et al., 2004; Subramanian et al., 2011) has used data from the Demographic and Health Surveys (DHS), a particularly rich source of information that had hitherto primarily been used for the analysis of fertility and “traditional” developing country disease challenges in the area of maternal and child health. In this article we exploit to an even greater extent the DHS dataset, using data on almost one million women aged 15–49 from 244 Demographic and Health Surveys (DHS) for 56 countries over the years 1991–2009. We use this data to examine a broader range of questions than the previous studies that predominantly focused on how overweight, obesity or BMI in women varied between socioeconomic groups within a cross section of developing countries. By combining the DHS data with appropriate socioeconomic indicators or determinants from other sources, we are in a position to test selected hypotheses that are either derived from theoretical predictions or, for the most part, have been proposed in the literature (or expressed in the public debate) without yet having been submitted to closer empirical scrutiny. Taken together, our findings aim to produce a set of empirically confirmed stylized facts on patterns, trends and correlates of being above normal weight (defined as having the Body Mass Index (BMI) of 25, or greater), in the developing countries. As a number of potential explanations for increasing overweight prevalence have been suggested over the last years (e.g. see Popkin et al., 2012 for a recent overview), we grouped them into two main categories (necessarily omitting some well-known potential correlates related to early life events, e.g. famine (Popkin et al., 2012), as well as genetics- related explanations): (1) those that concern the relationship between proxies of economic development and individual overweight, and (2) those examining the association between proxies of urbanization and proxies of technological change on one hand and individual overweight status on the other hand. It is important to note that we apply a somewhat more general definition of “technological change” here than in the economic growth literature, where it is more narrowly understood as a change in total factor productivity. We broadly interpret technological change as the process of invention, innovation and diffusion of various technologies (e.g. labour saving and food production ones) that may facilitate changes in energy expenditure or consumption, and therefore may affect the likelihood of being overweight or obese. There is a significant public health and economic literature which considers the term “technological change” in this similar, broader sense (e.g. Finkelstein et al., 2005; Huffman and Rizov, 2007; Lakdawalla and Philipson, 2009; Philipson, 2001; Philipson and Posner, 2003a; Swinburn et al., 2011) Given the data constraints, we focus the empirical tests of the hypotheses on women only. We examine a third group of hypotheses, linking overweight prevalence to globalization-related determinants in a companion paper (Goryakin and Suhrcke, 2013). The role of economic development in overweight ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Overweight is a “disease of affluence”, but only up to a point: as countries grow out of extreme poverty, overweight among women will increase. However, as countries continue to grow richer, the increase should slow down at some level of per capita income. Why should overweight be positively related to national income? The concept of the nutrition transition (Popkin, 2001; Popkin et al., 2012) emphasizes the role played by increased affordability of processed foods, as well as foods rich in calories high in fat and sugar (both due to rising incomes, and advances in food technologies) in developing countries, especially since the 1970s. Similarly, economic theory (Philipson and Posner, 2003b) predicts that technological change (which drives and accompanies rising national income) entails lower costs of consuming calories and higher opportunity cost of expending them, which, taken together, increases weight. Within this theory, the nonlinearity can arise because greater income would also increase the demand for health (and thus a BMI closer to the medically ideal level), assuming that health is a normal good (Grossman, 1972). The theory predicts that the latter effect would at some point more than compensate the weight-enhancing effect resulting from the changes in calorie consumption and expenditure. An adverse economic shock (recession) will be associated with lower likelihood of being overweight among women. How might body weight respond to a sudden fall in income? It is tempting to infer from the hypothesized concave and positive relationship between per capita gross domestic product (GDP) and overweight prevalence from Hypothesis 1.1 that a recession should be associated with a reduction in overweight in countries below a certain threshold of per capita GDP and either no change or even an increase in overweight in middle income countries. However, the concave relationship at one point in time across countries (which was the focus of Hypothesis 1.1) is more likely to describe a long-term association. Relationships between short-term changes in per capita GDP and changes in overweight may well differ from the long-term relationship in levels. In a series of papers, Ruhm (2000, 2001, 2003, 2005, 2007) examined the impact of recession on health in a range of high income countries, finding evidence for improving (deteriorating) health in response to a recession (boom). Ruhm (2000) found that obesity declined in times of rising national unemployment rates in the US, likely due to (1) the opportunity cost of exercise falling as unemployment increases, (2) more time available for health enhancing activities, and (3) possibly due to less income being available for calorie consumption. In addition, a number of studies have shown that undernutrition prevalence increased rapidly during times of global economic crises in Cambodia, Bangladesh, Indonesia, Kenya and Mauritania (SCN, 2009). Similarly, Pongou et al. (2006) have found that child undernutrition increased considerably during times of economic crises and structural adjustment programmes in the 1990s in Cameroon. To the extent that undernutrition is the flipside of overweight, this should mean that overweight will decrease in low income countries experiencing a recession. In low income countries, women of higher socioeconomic status (SES) will be more likely to be overweight than those with lower SES, whereas in middle income countries, the burden of overweight will shift towards women of lower SES, resulting in an insignificant or mildly negative relationship between SES and the probability of being overweight. The previous hypotheses predicted the association between national income and overweight prevalence in all population groups. By contrast, Hypothesis 1.3 considers within-country differences between women of varying socioeconomic status (SES), proxied by their level of educational attainment. The hypothesis is inspired by a widely cited study that demonstrated the shifting burden of obesity from the rich to the poor as countries’ wealth grows (Monteiro et al., 2004), using largely DHS data from 1992 to 2000 for 37 developing countries. Their results suggested that the reversal in the obesity gradient occurred at a level of GNP per capita of about 2500 US dollars, which was broadly supported by a recent systematic review (Dinsa et al., 2012). Another recent paper (Subramanian et al., 2011), that used more data from DHS, an additional SES proxy (i.e. wealth) and a more detailed analysis, found that in low income countries, the overweight burden was mostly concentrated among higher SES people, and that the wealth gradient became less marked only at a level of per capita GDP of about 5500 USD. Moreover, Neuman et al. (2011) found that the SES gradient in overweight prevalence did not weaken over time in the low-to middle income countries studied. The theoretical justification for this hypothesis can be linked to the explanation given in the Hypothesis 1.2: before the threshold per capita income is reached, the higher SES people will have greater income, greater demand for calories and therefore greater weight. However, as economic development progresses, the demand for thinness and health (backed by better access to healthier lifestyle opportunities) among higher SES women may start to outweigh the demand for calories faster than among women belonging to lower SES groups (Jones-Smith et al., 2011), thus leading to the switchover at a certain income level. In addition, some empirical evidence (Popkin, 2001) suggests that the income elasticity of demand for certain foods high in fat is higher for people belonging to lower SES classes. If this is indeed the case, then rising income levels among the poor may lead to greater relative consumption of unhealthy diets among them, compared to people from upper SES classes. The role of urbanization and technological change in overweight ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our next set of hypotheses reflects the widespread notion that urbanization and technological change have been driving the rise in overweight and obesity. Previous studies have linked urbanization with various technologies potentially associated either with the reduction in energy expenditure over time (Monda et al., 2007; Popkin, 1999; Rivera et al., 2002; Swinburn et al., 2011), or with more abundant supply and consumption of cheaper, energy dense foods (Drewnowski and Popkin, 1997; Popkin, 1999; Popkin and Gordon-Larsen, 2004). Therefore, we have subsumed several hypotheses concerning the role of urbanization and different manifestations of technological change under the same group of hypotheses. Living in urban areas will be associated with higher probability of being overweight among women. The strong positive association between urbanization and obesity has been examined and confirmed quite extensively in previous research (e.g. Martorell et al., 2000; Mendez et al., 2005; Mendez and Popkin, 2004; Subramanian et al., 2011), although Neuman et al.’s recent paper (2013) suggests that the association between urban residence and excess BMI may be mediated by higher SES of urban residents. The novelty of our approach is in the use of a much larger set of individual-level data collected in 56 countries, and in the non-parametric approach used to examine the effect of living in urban areas on overweight, by taking into account the level of economic development across countries. In general, living in urban areas can increase the odds of being overweight for a number of reasons. For example, people living in cities may expend less energy than those living in rural areas, possibly because they have fewer opportunities to exercise, because they prefer to spend more time doing leisure activities (such as TV watching) at home, or because they work in less physically demanding occupations. In addition, those who live in the cities may transition faster to the so-called “Western” diets rich in calories from fats and refined carbohydrates, but relatively poor in calories from fruit, vegetables and cereals (Popkin, 2001; Popkin et al., 2012). This may occur for a variety of reasons, including greater exposure to food processing technologies, as well as marketing and distribution systems that promote fast and processed foods consumption. Next, we consider specific ways in which urbanization can affect the likelihood of being overweight: Possession of assets that facilitate sedentary behaviour will be associated with greater overweight prevalence among women Technological change has given rise to an environment that may make it harder (and thus more costly) to spend calories, with cars and TVs being but two very widespread manifestations thereof. We hypothesize that owning a car or a TV set will be related to greater risk of being overweight. The explanation is intuitive: having a car may predispose people to walk less frequently, and therefore spend less energy (Sassi, 2010). Previous research has demonstrated that owning a motorized vehicle was strongly correlated with the probability of being obese in China (Bell et al., 2002). Similarly, people who own a TV set are expected to be less likely to exercise and, hence, less likely to “spend” their calories. They may also consume more energy-dense food while watching TV (Roemling and Qaim, 2012). We will separately also consider if there is a relationship between frequency of TV viewing and overweight risk. TV viewing may expose people to mass advertising of certain foods that may encourage people to snack excessively and may thereby contribute to weight gain (Finkelstein et al., 2005). Some empirical work in this area (mostly confined to higher income countries) has largely focused on individual measures of exposure to TV (e.g. number of hours per day watching TV, measured by self-completed diaries, e.g. (Cutler et al., 2003), or possession of a TV (Roemling and Qaim, 2012)). Overall, this evidence suggests a rather limited role played by TV exposure, at least in the high income countries studied (Finkelstein et al., 2005), although Roemling and Qaim (2012) did find that owning a TV was related to a small but significant increase in BMI in Indonesia. Note that our goal here is to try and isolate the overweight-enhancing influence of owning a car and a TV set, rather than the effect of socioeconomic status and income, which can also be correlated with owning a TV. Our dataset will allow us to control for both per capita national income as well as for individual socioeconomic status. In addition, technological change and urbanization are often hypothesized to lead to increased supply and, hence, consumption of higher energy and fat density foods (Popkin, 1999, 2006, 2007). Our data will allow us to test the following related hypothesis: A greater per capita supply of calories will be associated with higher overweight probability among women. Some research has argued that an increase in the consumption of calories played a greater role in the rise in obesity, compared to a reduction in energy expenditure (Cutler et al., 2003; Finkelstein et al., 2005). Similarly, proponents of the nutrition transition theory emphasize the role played by the growing proportion of calories obtained from diets higher in fats and refined carbohydrates, and not so much the increased calories per se (Popkin, 2001). A recent prominent study (Swinburn et al., 2011) differentiated roughly between two groups of global determinants of obesity and overweight: those related to the changes in the built environment, and those related to changes in the food supply. The authors argued that the built environment has not changed simultaneously and universally to lead to observed changes in people's weight around the world; hence, the more important drivers are attributed to the food system domain (including greater supply of cheap, energy-dense foods, as well as more advanced distribution and marketing systems to increase people's access to these foods). Similarly, Cutler et al. (2003) argued that although per capita energy expenditure in the US decreased from 1.69 and 1.57 kcal/min/kg between 1965 and 1975, it remained fairly steady during the period of fast growing obesity prevalence. On the other hand, available information on trends in energy intakes in the US suggests that calorie consumption was significantly higher in 1990s compared to the early 1980s (Finkelstein et al., 2005). Although changes in food systems encompass much more than changes in per capita supply of calories (one reason being that the available calorie data is imperfect, and food supply does not accurately capture the calorie consumption by people living in these countries), testing if there is country-level association between per capita calorie availability and probability of being overweight, using data from 56 low and middle income countries, may still be very instructive, as the available evidence so far has been scarce, and largely confined to high income countries. On a macro level, limited empirical evidence exists that increased food energy supply was positively correlated with greater weight in USA and UK (Swinburn et al., 2011). On a micro level, a systematic review of several experimental as well as large prospective cohort studies concluded that there is strong evidence on the link between fast food consumption and increasing weight (Rosenheck, 2008). One study which examined the association between caloric intake and weight in a middle income country (Russia) concluded that the relationship was strong and positive (Huffman and Rizov, 2007). By contrast, comparing trends in obesity and diet patterns in the US population, other research found a counter- intuitive coincidence of decreasing calorie consumption and growing obesity, which was tentatively attributed to even greater reductions in energy expenditures resulting from physical inactivity (Heini and Weinsier, 1997). Technological change has also entailed a shift in the employment of people from agriculture and industry to the service sector, which in turn may have led to the reduction in energy expenditure: Working in the service sector is associated with a greater likelihood of being overweight among women. This may be the case because working in the service industry is – on average – less physically taxing than manual labour in agriculture. However, the existing empirical evidence is far from unambiguous. Thus, as Finkelstein et al. (2005) noted, it is true that in the US, for example, the proportion of workers employed in the manufacturing industries fell from 27% in 1980 to 19% in 2000, and that the proportion of obese men has almost doubled over this period. However, the proportion employed in manufacturing was considerably higher in 1960, i.e. 35%, and yet the proportion of men who were obese changed little between 1960 and 1980. Hence, a more fine-grained analysis, using individual-level data, is needed. We will test this hypothesis with individual-level indicators on the occupational status of women contained in the DHS dataset. Data ~~~~ The data in this study comes from several sources: the Demographic and Health Surveys (DHS), World Development Indicators (WDI),1 the Cross-National Time-Series Data Archive (CNTS)2 and FAOSTAT.3 The main source is the DHS: nationally representative, population based surveys, which collect extensive information in the areas of population, health and nutrition. DHS surveys are extensively described elsewhere (Corsi et al., 2012b; Subramanian et al., 2011). The WDI, CNTS and FAOSTAT datasets provide a range of relevant country-level economic indicators that we employed for our hypothesis-testing. The individual-level DHS data used in this paper covers 19 years (1991–2009) over 56 countries. In the regression analysis, all country-level variables were taken from WDI, CNTS and FAOSTAT datasets, and merged with individual-level DHS data for respective countries and years. More information on the data sources used is provided in Appendix. As the full data was only available on women aged 15–49, we necessarily had to restrict our analysis to this group. The main outcome variable of interest is being above normal weight (or overweight, for short), defined as BMI greater or equal to 25 kg/m2. In order to trim outliers, observations for women whose BMI exceeded 50, or were less than 13, or whose weight was either greater than or equal to 220 kg, or less than or equal to 20 kg, were dropped. In addition, observations for women whose height was recorded as either greater than or equal to 2.2 m, or less than or equal to 0.9 m, were also dropped. We restricted the sample to women who were non-pregnant at the time of the interview (about 7.9% of surveyed women aged 15–49 were pregnant). In all countries, information was collected on women of child-bearing age, and thus included those who had never had children. The proportion of women with no children (which included women with and without pregnancies at the time of the interview) varied from about 7–8% of the sample (e.g. in Egypt) to about 45% (e.g. in Morocco), and this difference was largely due to the age structure of the surveyed population. Thus, in Egypt, only about 8% of surveyed women were under the age of 20 in 1995, while in Morocco in 2003, about 28% of surveyed women were under 20. This issue may be important for the comparison of overweight prevalence rates across countries, but we take care of it by adjusting estimates by sampling weights provided with the DHS surveys. Overall, in the pooled sample, data on BMI was available for about 72% of the full sample of 1,225,808 non-pregnant women. Again, there was some variation across countries in terms of the proportion of non-pregnant women with available BMI data (from 99% in Uzbekistan, to only 26% in Brazil), that we account for when estimating country- specific overweight prevalence by employing sampling weights. In our analysis we used education as the main socioeconomic variable of interest. While socioeconomic status can also be proxied by other variables – e.g. wealth, as was done by Subramanian et al. (2011) – we preferred education, because the way it is defined in the DHS set (using six educational quartiles, ranging from incomplete primary school to higher education), it may be considered reasonably comparable across countries as a proxy for the absolute level of socioeconomic status. By contrast, wealth – as it is commonly constructed on the basis of DHS data – is a relative measure of SES specific for each country, and thus is harder to compare when all countries with different income levels are pooled together. Nevertheless, we also tested the robustness of our findings to the inclusion of an alternative wealth measure (i.e. car ownership) as education may not fully capture the income and wealth- related aspects of socioeconomic status. Education was defined using DHS dummies for six levels: people with no, incomplete primary, complete primary, incomplete secondary, complete secondary, and higher education. For the graphical analysis (e.g. Fig. 4), we divided education into four levels: (1) no education, (2) incomplete primary and complete primary, (3) incomplete secondary and (4) complete secondary and higher education. We then compared levels 1 and 4. Note that this gradation is different from the one used in some earlier papers (e.g. Monteiro et al., 2004), which defined top and bottom educational quartiles by the number of years of schooling in each country. We decided not to follow their approach in order to avoid dealing with countries for which too many observations were clustered around a certain value (e.g. zero years of education). In such cases, it may not be possible to clearly define four separate quartiles, and one would only be able to deal with three (or even two) groups (Monteiro et al., 2004). In our approach, the rule for selecting top and bottom “quartile” will apply equally to all countries. A country was defined as being in economic recession if the country's GDP per capita (in purchasing power parities [PPP], constant international $, 2005) decreased by 1% or more between adjacent years. We also checked the robustness of the results by defining a shock as a more severe economic contraction, i.e. a 5% reduction in GDP per capita year-on-year. Information on the working status for a woman was based on self-reports, and we captured service sector employment using an occupational status dummy. Specifically, occupations were first aggregated into three groups as follows: (1) unemployed; (2) services (professional and managerial; clerical; sales; household and domestic; services); (3) agriculture (agriculture employed and self-employed); (4) manual (skilled manual; unskilled manual). Furthermore, for the services group a dummy was assigned with a value of one, while for manual and agriculture occupations the assigned value was zero. The availability of cars and TV sets was measured by individual-level indicators on the ownership of these items. The hypothesis on the supply of calories was tested using country-level data on average food supply (kcal per capita per day) derived from FAO balance sheets. Table A1 in the Appendix describes all the variables used in this paper. In the regression analysis, we were interested in the association between several potential correlates and the composite category of “being above normal weight”. We therefore assigned a value of one to people whose BMI was greater than or equal to 25, and a value of zero to people whose BMI was less than 25. We also considered including the continuous variable BMI as the outcome variable, as was done in a recent paper by Subramanian et al. (2011). However, a change in BMI may have very different implications, depending on the initial value in its distribution. For example, the change in BMI from 18 to 19 (i.e. from being malnourished to having normal weight – a desirable outcome) has very different implication compared to an increase in BMI from 24 to 25 (i.e. from having normal weight to being overweight – an undesirable outcome). Measuring the association between covariates and BMI would not capture this difference, while measuring the effect of covariates on overweight (treated as a dummy variable) does. Most hypotheses require either splitting the sample, or introducing an interaction term to distinguish between low- and middle-income countries. In this analysis, we used the World Bank year-specific thresholds defining the split between low income and middle income countries. We then either split the sample between these two groups and ran the regression analyses separately, or we introduced an interaction term between the World Bank income level and the main independent variable of interest. Econometric specifications ~~~~~~~~~~~~~~~~~~~~~~~~~~ The estimates also allow controlling for time-invariant country-specific effects αc. For example, a country may have a certain endowment of human or natural resources that contributes to changes in the prevalence of being overweight, as well as to variation in some of the observed determinants included in the model. Not controlling for such factors may contribute to a bias in the estimated parameters. Note that this model will allow parameter estimation only when the main variables of interest are either individual-level ones or are country-level variables that vary over time. In addition, including time dummies Dt will allow us to control for other potential time dependence or for any world- wide factors (e.g. global economic crises) that could affect our associations of interest. As we are interested not only in the estimation of the main parameters of interest, but also in whether they are statistically significantly different in low and middle income countries, ideally we would have preferred to use specifications (2.1) or (2.2) only. However, as we have six separate educational categories, and our goal was to take advantage of the whole educational distribution, interacting each educational category separately with the medium income dummy would not produce very intuitively interpretable results. Therefore, we split the sample to estimate the association of education with overweight separately in low and medium income countries, as per specification (1). As we also had to control for urban residence and GDP per capita in such models anyway, we presented the results for these three hypotheses using the same specification (1).","In order to put the obesity “epidemic” in developing countries into perspective and evaluate its comparative severity, it may be informative to start by comparing what developing countries have recently been experiencing with what today's high income countries had experienced at earlier stages of the obesity epidemic. Although some recent studies have looked at the long-term trends in obesity and overweight prevalence in the US (e.g. Finkelstein et al., 2005; Flegal et al., 1998), there was so far no attempt to compare these trends with those observed in low and middle income countries from different regions. Fig. 1 compares the trends in the prevalence of female obesity between a prominent high-income country (USA) and the four regions that we have defined. The US data for women was taken from Flegal et al. (1998) and CDC (2010), while the prevalence data for the developing countries was estimated from DHS data. Please note that in this one instance, we consider the prevalence of being obese rather than above normal weight, to ensure comparability between the DHS data with that from the US. Specifically, for each multi-year period, we constructed a region-specific prevalence of the indicator of interest, averaging over countries belonging to the respective regions, and using each country's population as a weight. We describe the regional split in more detail in the Appendix, Table A2. Although the trends presented in Fig. 1 do not represent a formal econometric test and therefore are only indicative, they do paint an interesting picture: the Middle East appears as by far the most worrying region in terms of obesity trends. Back in 1960, the USA already had reached a considerably higher GDP per capita than the Middle East has in 2009, and yet it still lagged behind the Middle East for most of the observed period in terms of female obesity prevalence. Hence one may conclude that the obesity challenge faced by this region does substantially exceed what even the US has had to confront. Given that the Middle East still has a lower per capita income than the US did in the 1960s, further economic development is likely to pose considerable public health challenges in this region. The comparison to the other regions is more problematic in that they are at even lower per capita incomes than the Middle East countries on average and thus the US obesity data would have to go back even further in time to allow for a meaningful comparison. The role of economic development in overweight ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The results of our first set of more formally tested hypotheses – the role of economic development in overweight – are reported below: Overweight is a disease of affluence – up to a point? (cross-country perspective) A simple graphical illustration can be used as an initial exploration of Hypothesis 1.1. In Fig. 2, we fitted a lowess plot for the relationship between the country-level proportion of women above normal weight (i.e. having BMI greater or equal than 25), and GDP per capita, measured using PPP. This very simple analysis suggests that, indeed, the association between these two variables of interest is positive and stronger in low-income countries than in the middle-income ones. As Fig. 2 shows a very simple relationship between the level of economic development and overweight prevalence, without controlling for any socioeconomic or political factors (likely to vary widely between countries), this finding should not be over-interpreted. Nevertheless, the observed pattern confirms earlier findings derived from smaller samples of countries and individuals (Monteiro et al., 2004; Martorell et al., 2000; Ezzati et al., 2005; Egger et al., 2012). A few interesting cases are worth highlighting: in 2007 Egypt had a far greater female overweight prevalence than could be predicted from their per capita GDP level. On the other hand, another African country – Namibia – was in the reverse situation, being located below the regression line. The high prevalence of obesity in Egypt relative to its GDP level was discussed in other studies (e.g. Martorell et al., 2000). Future research should try to identify the factors that help account for these large differentials. In Fig. 3, we examine how the same relationship may have evolved over time. It emerges that the peak at which the prevalence of being overweight is reached is at approximately USD 5000–6000 per capita, PPP. Moreover, it appears that this peak level has been slightly shifting to the left over time, suggesting the existence of a learning process, which allows countries that reach higher income levels to switch to a descending trajectory for overweight in the following time periods. On the other hand, it is quite worrying that for a given level of economic development, the prevalence of overweight tends to increase over time. Thus, at USD 5000 GDP per capita, the prevalence of being above normal weight among women in 1991–1994 was approximately 35%. In 2003–2009, this proportion increased to almost 50%. A similar upward shift of the income-BMI curve over time was observed in another study, using US data (Ezzati et al., 2005). This patterning emphasizes once again the role played by determinants of overweight other than national income, the influence of which was apparently increasing over time. It is also worth noting that this reflects the reverse picture of Preston Curve (Preston, 1975) which indicates an improvement in life expectancy over the decades (attributed to technological change), while in the case of overweight we observe a deterioration. Given the trends mentioned in a previous paragraph, it appears that as time passes, countries may learn to switch to a downward trajectory faster (i.e. at lower level of GDP), but before doing so they tend to reach a greater overweight prevalence than previously. A more formal test of this hypothesis can be performed with the help of simple regression analysis as outlined in equation 1 above, using individual-level overweight status as an outcome variable. In Table 1, column 1, we see that the log of GDP per capita generally has a positive association with the probability of being overweight, suggesting that a 10% increase in GDP per capita is associated with about 1.2 percentage points (p.p.) higher probability of a woman living in that country being overweight. Put differently, a twofold increase in a country's GDP from, say, 500 to 1000 USD, is predicted to lead to about 12 p.p. higher probability of being overweight. This finding is robust to including country-level fixed effects as outlined in column 2 (although the magnitude of the effect becomes smaller). This is a relatively small effect, but note that it is averaged over all regions and years, and is derived after controlling for educational attainment and for the proportion of people living in urban areas. By contrast, results in columns 3 and 4 (as well as in 5 and 6) show that the estimate of the association between log GDP and the probability of being overweight among women is insignificant for the sample of low income countries, and positive and significant for middle income countries. This finding contrasts with the picture presented in Fig. 2. We will come back to possible explanation for this finding in the Additional Checks section below. An adverse economic shock (recession) will be associated with a lower likelihood of being overweight among women. The results for Hypothesis 1.2 are summarized in columns 1 and 2 of Table 2. The first column provides the OLS results. We see that economic decline, as measured by a dummy for countries that experienced a decline of 1% or more in their real per capita GDP compared to the previous year, is associated with an increase in the probability of being overweight in low income countries, and a decrease in the likelihood of being overweight in middle income countries. This finding is, however, not robust to including country-level fixed effects (see column 2): the fixed effects results show that with economic decline people in poor countries are less likely to be overweight – as Hypothesis 1.2 predicted – while those living in richer countries gain weight. Note that when we defined an economic shock differently, i.e. as a contraction by 5% or more in real GDP per capita, we obtained very similar results in the model without country fixed effects (CFE), and somewhat different results for the model with CFE. Specifically, while the interaction parameter was very similar, the economic shock parameter for low income countries was now positive and significant (b = 0.03), suggesting that the effect of a more significant economic shock may operate differently in the lower-income countries. Nevertheless, as only 3% of the DHS sample lived during years of such strong recessions, it is important to be cautious in interpreting this specific finding. We can also test more directly how work status of women is correlated with being overweight. In Table 2, columns 7 and 8 we see that in low income countries, working is surprisingly negatively correlated with overweight, while the reverse is true in middle income countries. Our findings thus partly support the previous evidence of a positive association between country-level recessions and weight reduction in the low income countries (SCN, 2009). As for the middle income countries, our findings contradict our hypothesis (and thus the work by Ruhm (2000, 2001, 2003, 2005, 2007)) but are in line with the small strand of literature arguing that recessions increase the propensity to be overweight (Morris et al., 1992; Ludwig and Pollack, 2009). One potential explanation for this finding is that in the more developed countries, economic recessions (and loss of income from the main breadwinner) may lead to greater employment of women to supplement family income as a coping mechanism (Lim, 2000). In turn, they may have less time for cooking and may thus be more likely to be overweight. On the other hand, the reverse may be true for women living in the less developed countries, who may have to cut their total consumption of calories altogether (thereby losing weight). Nevertheless, there is reason to doubt this explanation in that the scope of finding work during recessions may be quite limited for the previously unemployed women. Instead, it is possible that the observed increase in overweight among women in recession-hit middle income countries, and the reverse trend in the low income countries, is due to different patterns of dietary response. For example, women in poor countries may consume fewer calories during recessions due to more binding constraints on overall food-related spending, while women in middle-income countries may have an option to switch to less healthy, high calorie food instead. Unfortunately, it is not possible to explore this issue further given the available data. Having said that, the finding that individual working status is positively associated with overweight and obesity in the middle income countries, is in line with the literature that argues that being employed may be related to fewer opportunities for exercise and taking care of one's health, as well as to a line of reasoning that women who work may be less likely to spend time cooking nutritious food, thus contributing to the increase in overweight (Finkelstein et al., 2005). In low income countries, women of higher socioeconomic status (SES) will be more likely to be overweight than those with lower SES, whereas in middle income countries, the burden of overweight will shift towards women of lower SES, resulting in an insignificant or mildly negative relationship between SES and the probability of being overweight. The association between education and the probability of being overweight is generally positive for the overall sample (Table 1, columns 1 and 2). Thus, women who have the least education, are about 11p. p. less likely to be overweight than those with the most education. The country fixed effects specification gives very similar results. Again, this number conceals variation between low and middle income countries: columns 3 and 4, as well as 5 and 6 in Table 1, show that at all levels of education, the association between this measure of socioeconomic status and the probability of being overweight is weaker for the sample of middle income countries. Moreover, when country-level fixed effects are used, the relationship between education and probability of being overweight actually turns negative in middle-income countries, as expected. This finding was affected to only a small extent by the inclusion of the alternative wealth measure, i.e. car ownership (results not shown here but available from the authors upon request). Additional evidence on Hypothesis 1.3 is also presented in Fig. 4 below, showing that up until about 5000–6000 USD GDP per capita (PPP), those in the top education quartile are more likely to be overweight than those in the bottom quartile. Beyond this threshold, however, it is those with the least education who have consistently greater probability of being overweight. This association becomes even more pronounced – and the threshold declines to around 4000 USD – when the obesity prevalence is measured on the vertical axis (see Fig. A1 in the Appendix). Note that our finding is consistent with the results of other, smaller-scale studies. For example, a negative association between education and obesity probability was found for women living in China (Xuhong et al., 2008), Brazil (Marins et al., 2007), Eastern Europe (Pikhart et al., 2007), Iran (Hajian-Tilaki and Heidari, 2009), Argentina (Fleischer et al., 2008) – all countries at a level of around 4000 GDP per capita (PPP, international 2005 $) or higher. In contrast, many earlier studies focussing on less developed countries tended to find positive associations between education and obesity (Sobal and Stunkard, 1989). In a recent study on the relationship between educational attainment and overweight among Chinese women (Jones-Smith et al., 2011), no relationship was found using 1989 data, but a strongly negative relationship (OR = 0.22 (CI 0.11–0.42)) was found for 2006. In another recent study that used World Health Survey data (Moore et al., 2010), the general finding was a positive relationship between education and overweight. Note, however, that their results were not split by country income classification, so it is hard to directly compare the results to ours. Monteiro et al. (2004) also found that as national income increased the burden of obesity was shifting from the top to the bottom socioeconomic status. The fact that our finding is similar, using considerably more data, and using a different definition for the top and bottom educational attainment, is encouraging. Similarly, Martorell et al. (2000) found that more educated women living in very poor countries were more likely to be obese than the least educated ones, while this gradient disappeared as countries got richer. On the other hand, Subramanian et al. (2011) used wealth as an additional proxy for socioeconomic status. Their finding was different in that there was no shift in the burden of obesity from the richest to the poorest people as national income was increasing (although there was a shift when the richest quartile was compared to the third and second quartile). That said, those results are not directly comparable to ours, since they used predicted BMI level, rather than prevalence of being above normal weight as an outcome variable. They also found that the association between wealth and BMI was considerably stronger than the association between education and BMI. However, since including both education and wealth in the same equation may mean that the influence of one variable is captured in another, it is hard to conclude from this with certainty that wealth is more important than education for nutritional health. The role of urbanization and technological change in overweight ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In what follows we summarize the results from our four hypotheses that could be subsumed under the factors capturing indicators of technological change. Living in urban areas will be associated with higher probability of being overweight among women. The tests of Hypothesis 2.1 are also contained in Table 1. It is obvious that this hypothesis is supported by the results from all specifications, suggesting that, in general, living in urban areas is related to an about 7–12 p.p. greater probability of being overweight, even after controlling for education. Moreover, the effect of urban residence on overweight is weaker in countries with a higher level of income – consistent with the finding from another study (Popkin et al., 2012) which used a different method, and different data – again suggesting a potential learning curve that occurs in countries as they undergo economic transformation. Also, Fig. 5 shows that for all levels of GDP per capita, the probability of being overweight is higher in urban areas, and thus unlike education, there appears to be no switchover in obesity occurring from urban to rural areas as national income grows larger. This confirms previous findings by Mendez et al. (2005), which however where based on less comprehensive data and used a smoothed plot, rather than lowess regression. Similarly, a recent paper using data from 42 low and middle income countries (Popkin et al., 2012) also found that on average, urban women had higher baseline prevalence of being overweight, compared to rural women, and this was true for all countries they studied. Moreover, they found that in countries with higher GDP per capita, there was relatively little association between residence and overweight prevalence, and that this association was much stronger in lower income countries. Given the worldwide trend towards urbanization that is set to continue in developing countries (United Nations, 2011), the consistently overweight-enhancing effect of urbanization does raise concerns that this will remain a factor driving a further spread of obesity in the developing world. Possession of assets that facilitate sedentary behaviour will be associated with greater overweight prevalence among women. Hypothesis 2.2 was confirmed both in OLS specification and CFE specifications (see Table 2). Thus, having one car is associated with an about 12–14 p.p. higher probability of being overweight. The effect was also considerably stronger in the low-income countries compared to the middle income ones, suggesting that changing transportation habits there may entail a particularly high risk of adverse weight outcomes. Similarly, TV ownership was significantly related to female overweight in both OLS and fixed effects specification (see Table 2). Interestingly, the magnitude of association is very similar to what we found for car ownership variable, suggesting that having a TV is related to an about 12 p.p. greater probability of being overweight among women in the low income countries. The association becomes slightly smaller in size in the middle income countries when country fixed effects are included. We also found a quite strong dose-response relationship between frequency of watching TV and overweight risk (using the country fixed effects specification – detailed results available upon request). Thus, compared to the reference group of never watching TV, those who watched it less than once a week; at least once a week; and almost every day, had an 4.5 p.p., 7.2 p.p. and 11 p.p. greater risk of being overweight, respectively. A greater per capita supply of calories will be associated with higher overweight probability among women. The food supply hypothesis was tested by running specifications reported in columns 9–10 of Table 2. As expected, there is a significant positive association between average, country-level per capita supply of calories (for convenience of interpretation, expressed in 1000 kcal per capita) and the probability of female overweight in the simple OLS model. The effect is also stronger in the middle income countries. To put this into perspective, the average total supply of calories, across all countries, is about 2400 kcal/person/day. Increasing this availability by 1000 kcal, or by about 41%, would lead to about 11.9 p.p. greater probability of women being overweight in the middle income countries. The effect remains significant for the middle income countries in both models. It should be emphasized that this is a test of the relationship between average per capita supply of calories and the probability of being overweight, and, as noted above, the former is a crude measure of the actual consumption of calories in the country. This measure does not take into account the potential changes in the patterns of the diet either, e.g. the potential shift to higher consumption of calories from refined carbohydrates and fats (Popkin et al., 2012). Nevertheless, the size of the association that we found for the middle income countries supports the hypothesis that greater supply of food plays a role in increasing overweight risk, at least in the middle income countries. Working in the service sector is associated with a greater likelihood of being overweight among women. This hypothesis was also confirmed in both the OLS and fixed effects specification (see columns 3 and 4 in Table 2). Thus, working in the service sector was associated with an about 8 p.p. higher probability of being overweight among women, and the effect was somewhat weaker in the middle income countries (although it still remains significant).","One potential concern with our findings is that traditionally used BMI cut-off points ultimately are arbitrary and at least in some countries or regions inappropriate to capture the prevalence of weight problems (James, 2004, 2008). To tackle this problem, we ran the same regressions, but with differently defined cut-offs for being overweight. The results are presented in Fig. 6, focussing – for illustrative purposes – on the role of education in overweight and its sensitivity to changes in the cut-off. All underlying regression estimates included the same controls as in Table 1. In the low-income countries, the effect of education is very sensitive to the cut-offs used. Thus, when being overweight is defined as having a BMI greater or equal to 23, the people with some or complete university education have about 15 p.p. greater probability of being overweight than the reference group. However, this differential falls by more than 50, to about 9% p.p., when the cut-off is at about 26. The effect of university education on being overweight and obese is at its smallest when the BMI cut-off is 30. Comparing this picture to the one emerging for middle income countries (Fig. 6), we can see that there is much less sensitivity in the effect of education by different BMI cut-offs. We also checked how sensitive the results were to using different thresholds defining the split between low and middle income countries. Table 3 shows that the OLS results were particularly sensitive for the log GDP per capita variable (please note that OLS version was chosen over country fixed effects, as there was an insufficient number of countries to estimate parameters for log GDP per capita variables in columns 2 and 4). It is important to bear in mind that in the results presented in Table 1, the effect of GDP on the probability of being overweight was surprisingly insignificant in low income countries and positive in middle income ones. However, now we see that this is only true when comparatively low income levels are chosen for the threshold separating low from middle income countries. When the middle income level was defined as having a GDP per capita greater than 3000 USD or higher (PPP, 2005 international dollars), our earlier unexpected finding is reversed, and is now in line with the original hypothesis. Note that this finding is also in line with what we found in Fig. 2. Education turned out as another category that is quite sensitive to the cut-offs used: the relationship in the predicted direction was again at its strongest when a higher cut-off was used. There was little change in the association between urban living and being overweight regardless of the threshold level chosen. Table 4 provides the same information for the remaining hypotheses (for convenience, only the main parameter estimates are presented). We see that in this case the choice of the cut-off had little effect on almost all variables. Hence we conclude that the choice of a cut-off for the income level made a substantial difference only for Hypotheses 1.1 and 1.2. Note that in this case again the OLS version was chosen over country fixed effects, as there was an insufficient number of countries to estimate several parameters in columns 4 and 5. We also examine whether the effect estimates of the main determinants vary by region. In Table 5 we present OLS results for the equation reported in the first two columns of Table 1, but this time split by 4 regions, and for two different outcomes: being above normal weight (i.e. when BMI is 25 or greater), and being obese (i.e. when BMI is 30 or greater). Here the OLS version was again chosen instead of country fixed effects estimates, as there was an insufficient number of countries to estimate parameters on the log GDP per capita variable for columns 4 and 8. Several interesting findings emerge. First, in the Americas and the Middle East, education has a different association with the probability of being overweight among women, than in Africa and Asia (qualitatively, the same results are obtained when using obesity as the outcome variable.) While in the Americas and the Middle East, women with the most education have a lower probability of being overweight than almost all other educational categories (except the lowest one), this is not the case in other regions, where less education is generally associated with a lower probability of being overweight. A possible explanation for this pattern is that most countries in the Americas and Middle East are already past the GDP threshold at which the overweight burden tends to shift to lower socioeconomic groups. In addition, log GDP per capita has a considerably stronger association with overweight in the Middle East: in this region a 10% greater GDP is associated with an about 12 p.p. higher probability of being overweight. It appears that in the Middle East, economic development may fuel the rise in overweight prevalence much faster than in the rest of the world. Also, the association between age and obesity is much stronger in the Middle East and in the Americas than in other regions (not shown here). This implies that as the population ages in those regions, the obesity burden there may become particularly acute.","While the growing obesity challenge in the developing world may have been fairly widely acknowledged by now, precisely capturing the size, patterns, trends and determinants of overweight and obesity in developing countries has thus far been severely hampered by the lack of comparable micro data in low and middle income countries. To date, most of the broadly related studies focussed either on calculating country-specific prevalence of obesity, or on examining how socioeconomic status of individuals and their place of residence is related to obesity burden. By taking advantage of a considerably bigger sample than used in previous studies, we have re-examined previous findings and have tested a number of additional hypotheses derived from existing theoretical literature, or from widely held notions about alleged determinants of being overweight, that had been taken for granted without having been empirically tested. Wherever possible, we tested for several hypotheses simultaneously, by including relevant variables of interest at the same time. On the whole, our results confirmed the set of hypotheses about the association between economic development and overweight: Hypothesis 1.1: The relationship between national per capita income and obesity is positive and concave; Hypothesis 1.2: In an economic recession, people in poor countries lose weight and, hence, are less likely to be overweight (although this was not the case in the middle income countries); Hypothesis 1.3: The relationship between education (as a proxy for socioeconomic status) and the probability of being overweight is positive in the low income countries and negative in medium-income countries. The latter finding in particular suggests an important policy implication: in order to minimize negative consequences of economic development in the form of greater overweight, there may be reason to target attention at preventing overweight and obesity among higher socioeconomic groups in low income countries, and at lower socioeconomic groups – in middle income countries. Moreover, it appears that the greatest returns to better nutritional status from more education are in the Americas and the Middle East regions (see Table 5). Given that the Middle East is the most problematic region in terms of obesity prevalence, this finding is especially promising, as additional investment into female education in this region is likely to bring the greatest returns in terms of reducing overweight. The picture emerging from the next set of hypotheses suggests that urbanization and related technological change do play at least some role in the growing overweight prevalence: Hypothesis 2.1: Our finding that living in urban areas is related to significantly higher likelihood of being overweight for countries at all income levels suggests that the continuous urbanization process taking place in the developing countries poses a serious public health challenge. Hypothesis 2.2: Both car and TV ownership are robustly related to overweight outcomes, with the magnitude of the association being smaller (yet still positive) in the middle income countries compared to the low income countries. Hypothesis 2.3: Greater per capita calorie intake is positively related to overweight, but only in the middle income countries. Hypothesis 2.4: Shifting patterns of employment from agriculture into services, traditionally associated with urbanization and technological change are also significantly related to a greater probability of being overweight in both low and middle income countries. In addition, we have dealt with the previously mentioned cut-off problem (e.g. James, 2004, 2008) in a systematic way. For example, we have demonstrated that the results can be quite sensitive to choosing different thresholds for defining overweight. Moreover, we checked how the estimation results varied depending on the level of income threshold chosen to differentiate between low and middle income countries. Inevitably, our study suffers from several limitations. For example, the sample was necessarily restricted to women only, mostly of child-bearing age. Therefore, generalizing our findings to women of all age groups, let alone to both genders, is impossible. Nevertheless, since the age group 15–49 represents the most economically active group of women, who also typically have a number of people depending on them (including children and the elderly), focussing our attention on this demographic segment seems to cover a sizeable and important share of the population. Also, by the very nature of the sampling undertaken for the DHS surveys, most sampled women were mothers with at least one child under 5 years of age (Monteiro et al., 2004). This can be problematic in that such women may be more likely to be overweight, and therefore the estimated prevalence of obesity may be slightly overestimated. Nevertheless, since our goal was not to estimate country-level obesity prevalence, but rather to test the statistical association between overweight status and its correlates, this should not be a significant concern in our study. The sample composition was also changing depending on the availability of certain variables. For example, information on the occupational status of respondents was available for about 50% of the sample with BMI information, while for most other variables the data was much more available. In addition, since the data we used is of a cross-sectional nature, we have to caution against deriving too strong causal inferences about the identified relationships. Rather, our goal was a more modest test for any potential partial correlation between a number of relevant socioeconomic variables and overweight. Finally, we did not explore how variation in overweight status was partitioned between national and community levels. A recent paper by Corsi et al. (2012a), using similar data, did focus on this aspect. Despite the limitations, this study presents an important step towards quantitatively examining the role that different proxies of economic development, urbanization and technological change play in explaining overweight in developing countries, using a unique global dataset. While the findings support several previously expressed hypotheses of what might explain obesity patterns in developing countries, they also reject some prior notions, and add considerable nuance to the emerging pattern. More research is needed to assess the extent to which the findings reflect a true causal relationship."],["Whether money is exogenous or endogenous is the subject of one of the most important and intriguing debates in monetary economics. The aim of this article is to contribute to this longstanding debate through detailed examination of different notions of endogeneity and exogeneity of money. I argue that the debate has been too simplified. In reality, money can be either endogenous or exogenous, depending on several factors. Not only has the debate prompted some economists to draw too far-reaching conclusions, but it also misses the point. --------------------------------------------------------------------------------","Whether money is exogenous or endogenous is the subject of one of the most important and intriguing debates in the monetary economics. The discussion of whether the money supply is a cause or an effect of economic activity goes back to the dispute between the Currency School and the Banking School, if not before (Artesis and Howells, 2001). The neoclassical synthesis adopted an exogenous-money position, which was challenged, among others, by Kaldor (1970), Chick (1973), and Moore (1998). Although it is now a consensus position that central banks set interest rates and allow the quantity of reserves to float, as the post-Keynesian economists argued, the debate is alive and of key importance for monetary theory and policy. At the most general level, exogeneity means the supply of money is independent of demand. Endogeneity means the opposite. On the one hand, it should be clear that money is endogenous. After all, the majority of money is supplied by commercial banks in the form of bank deposits. It would be strange to assume that commercial banks do not respond to the public's demand for loans or preferences regarding the form of money it wants to hold. On the other hand, it is similarly obvious that money is exogenous. Do not central banks have a monopoly on cash issuance? It would be rather awkward to assume that a monopoly producer does not control the supply of its product, as this is the very definition of a monopolist. In this article, I cut this Gordian knot. To do so, I start by discussing different notions of monetary endogeneity/exogeneity, as the term may be understood differently. For example, Palley (2002) lists several different types of money endogeneity/exogeneity.1 As these notions are often conflated, I analyze them to clarify the debate. The meanings I distinguish among are shown in Table 1. In the remainder of the article, I analyze the notions and prove that economists often wrongly color conclusions about one concept of exogeneity with conclusions about another concept. As one can see in the table, I differentiate between exogeneity/endogeneity in the following senses: origin and evolution (Section 2), quantity (Section 3), bank money (Section 4), statistics and policy (Section 5), and control (Section 6). In Section 7, I discuss the debate between the horizontalists and the structuralists. Section 8 concludes.","The first notion of exogeneity/endogeneity concerns the origins of money. According to Knapp (1924) and his followers, on the one hand, money was introduced by the authorities. Money is thus an exogenous creation of law and the state. On the other hand, Menger (1871, 1883, 1892) developed the concept of spontaneous or “organic” genesis of money, which eventually became the mainstream theory of the origin of money. According to him, money was not the top-down result of an act of will, but the unplanned product of market mechanisms. Money evolved spontaneously in the market through the self-interested actions of individuals who wanted to improve their position. Menger (1871) noted that economizing individuals are inclined to exchange their products not only for those goods that they desire to consume directly, but also for more marketable goods, as the possession of goods with high marketability improves their chances to acquire what they really want. As more and more people start to accept these goods in their exchanges, their marketability rises even further. This gradual, self-reinforcing process eventually leads to the selection of a common medium of exchange—that is, money—although nobody intended it. Hence certain commodities were endogenously selected as monies on the basis of their ability to reduce transaction costs and facilitate exchange. The second notion of exogeneity/endogeneity concerns the evolution of money after its origin—in particular, the switch from commodity money to fiat money. Mainstream economists believe that money endogenously evolved from commodity money to fiat money, as individuals economized on production and transaction costs (for example, Thornton 2000). As it required many more resources to produce commodity money, society replaced it with cheaper, fiat money. Meanwhile, Austrian economists argue that fiat money did not endogenously (read: spontaneously) emerge on the free market, but resulted from exogenous government intervention in the monetary sphere (for example, Hoppe, 1994; Hülsmann, 2008).","The origins and evolution and nature of money are one thing, and the quantity of money is another. The debate about exogeneity/endogeneity of money generally concerns the determination of the latter. Indeed, for Fontana (2003, 291–92), “the essence of endogenous money theory is that the stock of money in a country is determined by the demand for bank credit, and the latter is causally dependent upon the economic variables that affect the level of output.” Similarly, for Rochon and Rossi (2013, 214), the essence is that “the supply of money expands and contracts with the needs of production, in response to expectations of aggregate demand, through the banking system.” However, the debate over whether the quantity of money is endogenous or exogenous is too general. Our position depends not only on our understanding of exogeneity/endogeneity, but also on the monetary system and whether we are analyzing outside or inside money. For example, it is widely believed that under the gold standard—understood here as a monetary system without a central bank and government intervention in the monetary sphere, with gold as the outside money—the money supply is exogenous, as the total stock of gold cannot be increased at will. Commercial banks have no control over the money supply. In a full- reserve system, they cannot create money at all, while in a fractional-reserve system, they need gold reserves, the supply of which cannot be increased at will—that is, the reserves constitute an exogenous constraint on credit creation. This is the core of the main argument for and against the gold standard: neither the government nor the banking system is able to arbitrarily increase the money supply. For supporters of this system, it protects people against inflation, while for opponents it makes the monetary system not elastic enough to stimulate the economy during recessions. However, it is not the case that the supply of gold is a totally independent factor under the gold standard. After all, the supply of money under the gold standard can only be what is consistent with the underlying market fundamentals. Indeed, the allocation of gold between monetary and nonmonetary uses, gold mines’ decisions about the production level, and gold flows between countries are endogenously determined by the price mechanism operating within the economic system and the resulting changes in the balance of payments. Hence, if we use “endogenous” as a synonym for “within the market system,” then money is endogenous under the gold standard (although the total stock of gold is given exogenously, the supply of gold flowing into market is, by definition, endogenous, as it depends on unfettered decisions of market participants). In contrast, under a fiat standard, the supply of money is exogenous, when the world is understood as being outside the market system, as it is determined, at least partially, by the central bank, a nonmarket institution. In principle, under a fiat standard, the central bank, a monopolistic producer, can choose the supply of outside money to be whatever it wants, but in practice, it can decide to respond more or less to commercial banks’ demand for reserves, rendering it dependent on the state of the economy and thus endogenous. Hence, whether money is endogenous or exogenous depends on the monetary system and its institutional details.","The endogeneity/exogeneity debate often boils down to the debate about the proper theory of banking, or, to be more specific, the right theory of money creation and bank lending. According to the exogenous view, banks need reserves in order to grant loans on top of them. Meanwhile the endogenous view claims that loans create deposits. The confessions of central banks’ practitioners and the failure of quantitative-easing programs to revive credit activity and economic growth provide evidence of endogeneity in that sense. Banks do not lend simply because they have more reserves. Rather than paying attention to the relationship between commercial and central banks, or between banks’ loans and reserves, some researchers focus on the relationship between commercial banks and their clients, or the key role of both in the payment system. The supply of money is considered endogenous in this view as it is determined by firms’ need to pay for the costs of production. The production decisions of companies generate the demand for loans (Moore, 1988). Commercial banks set the interest rate on loans (the policy rate plus a markup) and accommodate the demand for loans, so money is endogenous. However, commercial banks do not merely respond to the demand for loans. They can ration credit and alter the borrowing rate and other credit conditions to affect the demand for loans. This means that although commercial banks cannot determine the loan supply at will (as they have to find creditworthy borrowers), they do not merely accommodate the demand for loans. Hence money is not fully endogenous in this sense. Certain economists believe that money is endogenous by nature. As Rochon and Rossi (2013, 224) put it, “Money is and has always been an endogenous phenomenon, owing to its being essentially tied to the nature of debt and the need for a means of final payment that has to be provided by a third party on the agents’ demand.” The endogeneity of money does not depend on any institutional factors or the adopted policy by the central bank; rather, it is a natural occurrence in a capitalist economy, even a logical necessity.2 This is because money is used in every payment in the economy, and the payments are elicited by the “needs of trade.” If “endogenous” means here money that is not just a veil, as the neoclassical economists assume, I agree with this notion. After all, because of the Cantillon effect, changes in the money supply are not neutral, even in the long term (Sieroń, 2019). And I concur that money is an inherent feature of capitalism, although for different reasons (not because it is allegedly elicited by the need of trade, but because it enables rational economic calculation). However, referring to its nature is a rather unusual way of framing the debate about the endogeneity/exogeneity of money. Money, of course, plays a key role in the market economy, but it does not follow that institutional features of the monetary system do not matter. Given that the debate about the endogeneity/exogeneity of money often boils down to whether central banks control the money supply, 3 one cannot abstract from central banks’ policies.","Moore (1988) distinguishes between exogeneity in the statistical and control senses (I analyze the latter in Section 6). In econometrics, or statistics, exogeneity means a variable is independent of other variables, while endogeneity implies a variable is jointly determined with other variables in the system. Of course, in reality all variables are interrelated.4 Hence, in a statistical sense, whether a variable is exogenous or endogenous can only be determined in the context of a particular model. We thus say that a variable is exogenous when its determinants do not include variables determined within the model, while endogenous variables are those whose values are determined by the model (Chick, 1973). Now, it should be clear that the debate is not about exogeneity in the statistical sense, as “it takes little technical expertise to identify exogenous variables in an explicit model” (Chick, 1973, 85). We all know, for example, that in the quantity theory of money, the money supply is supposed to be exogenous. It is similarly uncontroversial that the money-multiplier model assumes the monetary base is exogenous. These assumptions are what transform tautologies into theories. But what really matters is the adequacy of the model, not the nature of the variables per se. Hence the debate is not really about the exogeneity of money, but the determinants of the money supply in the real world, where the term “exogeneity” does not apply. The confusion stems from the fact that what really matters is causality, not exogeneity, but only exogenous variables can be independent sources of change within a particular model. So, if we investigate the impact of changes in the money supply on prices, it sounds reasonable to assume that the money supply is exogenous in a particular model, although in reality the money supply can be affected by other variables. The bewilderment is compounded by the fact that some people use the term “exogenous” to describe policy changes. The government is seen as being outside the private sector, so this approach seems to be justified. As a reminder, exogenous variables are those whose values are determined outside the system under consideration. Therefore, when economists model the private sector, including its response to some policy changes, it makes sense to assume that money is exogenous in the policy sense—that is, that policy parameters are exogenous. However, there are some problems related to this approach. Although it is true that the government is not subject to the profit-and-loss mechanism, it cannot fully abstract from developments in the private sector. Actually, public choice theory shows that the private sector demands certain policy actions (Tullock, 1993). It is also true that at least some components of the money supply are provided by the public sector. But we should not on this basis draw the conclusion that the money supply is exogenous, as in doing so we would be assuming something that needs to be proven. After all, the majority of the money supply in the modern monetary system is created by commercial banks in the form of bank deposits (McLeay et al., 2014a). Even if one agrees that the monetary base is determined exogenously by the central bank, a nonmarket institution (and not all economists agree with this), it does not follow that the whole money supply is exogenous. Exogenous variables should not be conflated with policy-determined variables. After all, monetary policy is set generally in response to past, current, or expected economic developments (Goodhart, 2001).","As I have already shown, the endogeneity/exogeneity of money can be understood in a few distinct ways. However, most authors focus on the control sense. In that sense, the exogeneity of the money supply implies that “its nominal size is or can be controlled by the monetary authorities and does not automatically change as people make payments or try to build up or run down their money holdings” (Yeager, 1997, 131). Endogeneity in the control sense means the opposite. Graphically, the debate may be expressed using the IS/LM framework. A vertical LM curve expresses the notion of full exogeneity, as the supply of money is fixed and unresponsive to money demand. The second extreme case is the horizontal LM curve, which implies a perfectly elastic supply of money and money endogeneity. The central bank sets the interest rates at which it supplies money on demand. To resolve this debate, one must first clarify the concepts of “money supply” and “control.” Obviously, the central bank does not directly control the total supply of money. Nobody argues otherwise. However, monetarists claim that the monetary authorities control the monetary base and that there is a strong link between it and the broader monetary aggregates; hence the money supply as a whole can be taken as exogenous. The assumption of such a link is, of course, necessary, because the money supply in the contemporary monetary system consists mainly of the liabilities of the banking system. Therefore monetary policy relies on the banking system: its actions have to be transmitted by the banking system to exert a desired effect on the total money supply. The theory that links the behavior of central banks and commercial banks (or the monetary base and the total money supply) is the money- multiplier model. The money multiplier itself is an arithmetical tautology that shows how large the total money supply can be with a given monetary base and reserve ratio. But the monetarists make two assumptions that transform the mere tautology into theory: first, the monetary base is exogenously determined; and, second, there is a strict link between it and the total money supply. The nonmonetarist camp presents many counterarguments. Some researchers focus on the first assumption, arguing that the central bank does not control the monetary base. For example, Goodhart (1994) notes that in the modern banking system, commercial banks convert currency into deposits (and vice versa) at par. This implies that fluctuations in the public's demand for cash must be accommodated. Post-Keynesian economists offer an even more radical argument, as they reverse the monetarists’ causation between the monetary base and bank deposits. They argue that banks first grant loans based on demand factors and perceived risk and profitability. Banks look for reserves only later, and the central bank accommodates that demand. The fact that the central bank supplies reserves on demand implies that the monetary base is not exogenous (in the sense of being the source of change in the economy) but rather a function of investments and demand for loans, or the needs of trade. So it responds endogenously to changes in the demand for money and other developments in the economy (Chick, 1973; King, 1994; McLeay et al., 2014b).5 Others pay more attention to the second monetarist assumption. The moderate critics point out the instability of the money multiplier, as commercial banks can choose to hold excess reserves while the private nonbank sector may change its preference for cash and deposits. So the critics notice some breaks in the supposedly tight link between control of the monetary base and control of the broad money supply. In other words, the money multiplier is believed to be unstable—that is, the relation between base money and total money supply changes over time in an unpredictable manner. For example, the M2 money multiplier in the United States contracted by about 50% during the Great Depression (Von Hagen, 2009) and also during the Great Recession, as one can see in Fig. 1. Meanwhile, as I have already discussed, the more radical opponents of the money-multiplier model reverse the causal link between bank deposits and the monetary base. There are two key elements of this multiplier à rebours. First, commercial banks are no longer mere intermediaries of monetary policy that mechanically multiply reserves given by the central bank. The post- Keynesian model gives banks some degree of control over their own business, which is a more plausible assumption. In consequence, there is slippage between the monetary base and the broad money supply. Although this model weakens the case for the money-multiplier model, it does not fully demolish it, as the slippage might be less quantitatively important than changes in the monetary base. The second element seems more significant, as it refers to the central bank, which is said to passively accommodate the demand for money. Such behavior is related to being a lender of last resort. It is believed that to adequately exercise that function, the central bank should provide liquidity on demand.6 This second component also refers to the central bank's policy-response function. The argument goes as follows: the central bank reacts to economic developments; hence the monetary base cannot be exogenous. Exogenous variables are those whose values are determined outside the system under consideration. But central banks alter their policy targets in response to developments in the economic system, implying the endogeneity of money (Moore, 1998).7 One problem is that it is not clear what the word “control” means in the context of the money supply. Does it mean determining the amount of money supply at will, or just exerting a dominant influence (however that may be understood)? Another problem is that controllability is often confused with exogeneity. Logically, the money supply may be endogenous and either controllable, through the central bank's reaction function, or not, because of money supply disturbances generated from within the economy and outside the central bank's control (Palley, 2002). Similarly, the money supply may be exogenous and either controllable, through the monetarist multiplier theory operating under the fiat-money regime, or not, as under the gold standard. In other words, as Fischer (1993, 248) notes, “the claim that money is a controllable policy variable does not necessarily imply that money is exogenous in the statistical sense. The monetary authority may not have perfect control over the monetary aggregate it targets because the money multiplier is endogenous. Alternatively the central bank often watches several economic variables while implementing its monetary policy, hence the quantity of money depends on the central bank's reaction function. Moreover, it is extremely difficult to counteract unexpected changes in aggregate demand.”","The debate about the endogeneity/exogeneity of money is also complicated by the fact that there are two main strands of the endogenous-money approach. The first is horizontalism, the most extreme version of money endogeneity. According to this view, the central bank exogenously sets the interest rates at which it fully accommodates demand for reserves that results from banks’ decisions to meet all creditworthy borrowers’ demand for loans. Thus, under this approach, the money supply curve is horizontal—that is, perfectly interest elastic, as, although the central bank sets the short-term interest rate, it has to always accommodate bank demand for reserves in order to preserve the stability of the financial system (Moore, 1998). The second strand is the structuralist view. It also adheres to the core proposition of the endogenous-money approach that bank lending drives the money supply. However, the structuralists additionally take account of the role of portfolio preferences, uncertainty, balance sheet positions, profit-seeking behavior, microeconomic financial constraints, financial innovations, and expectations in influencing the money supply (Wray, 1990).8 As the structuralists accept the upward- sloping money supply curve, they are between the verticalists and the horizontalists. The proponents of this view argue that the central bank retains some control over the supply of reserves and only partially accommodates the demand for them. This implies that the central bank can decide which instrument to target and whether to accommodate demand for reserves. That is, it does not automatically provide whatever liquidity is demanded, as it is an active player that additionally operates under a set of constraints that affect its ability and willingness to conduct a fully accommodative policy (Palley, 2013). Moreover, in contrast to the horizontalist view, the structuralist approach underlines the importance of commercial banks’ asset and liability management, which allows them to (partly) overcome reserve constraints. Thus, according to Fontana (2003), the latter view is broader and more sophisticated than the former. For the structuralists, endogenous money is more than a horizontal money supply curve or the central bank's reaction function. They also take the liquidity preference of households and banks into account, maintaining that commercial banks are not merely quantity takers in the credit market that only passively accommodate the demand for loans. Hence, although both groups share the fundamental insight that bank lending drives the money supply, the horizontalists focus on the control sense—that is, whether the central bank accommodates the commercial banks’ demand for reserves. In this approach, the endogeneity of money depends exclusively on the monetary authority and its willingness to meet the demand for reserves. Meanwhile, the structuralists underline in addition the banks’ asset and liability management activities. In that approach, the endogeneity of money depends on both the central bank's behavior and the actions of commercial banks, which are not fully dependent on the central bank (Palley, 1994). In other words, the horizontalists claim that money is endogenous, because central banks willingly and deliberately supply the reserves on demand. The money supply curve is thus horizontal, as the suppliers of money always fully accommodate the demand for money at a given interest rate. Meanwhile, the structuralists argue that money is endogenous, because—thanks to innovative techniques of management of assets and liabilities—commercial banks can lend largely free of any central bank constraint (Howells, 2005). Hence, although monetary policy is effective in the short run, it is “outwitted and subverted by the banking system in the longer term” (Chick and Dow, 2002, 588). The money supply curve is upward sloping but moves rightward in the long run.9","The conclusion is clear. The debate about endogeneity versus exogeneity of money has been too simplified. The issue is more complex than many economists would like to admit (Chick, 1973). What complicates the matter is the fact that many participants in the debate use the terms without proper defining them. In particular, the concepts of controllability and exogeneity (and policy-determined and exogenous variables) are often conflated. The problem exists because central banks have deliberately chosen to set the interest rate and supply reserves on demand. Hence they have retained their control over the money supply at some ultimate level but changed the way they exercise it.10. After all, central banks can set short-term interest rates because they have a monopoly right to create money.11 Economists also confuse the worlds “control” and “influence.” This is partially understandable as to control is sometimes defined as “to maintain influence.”12. But it is clear that the debate is not really, or should not be, about control of the money supply, as absolute control is impossible in our monetary system, but about initiative and influence. Central banks might not have full control of the money supply, but as long as they maintain a non-negligible influence on markets, it does not really matter. There is no simple answer to the question of whether money is endogenous or exogenous. Money can be either endogenous or exogenous, depending on several factors. For one thing, it depends on the components of the money supply. The question is too general, as money is too vague a concept and money supply is too large an aggregate. In the modern economy, different monies coexist. Commercial banks create deposits while central banks create cash and bank reserves. Hence the debate is poorly framed and assumes a false dichotomy: the money supply must be either determined exogenously (by the central bank) or endogenously (by the commercial banks and the public). In reality, the money stock is jointly determined by the interaction of the central bank, commercial banks, and the public (Brunner and Meltzer, 1981; Chick and Dow, 2002).13 The answer thus depends on its components. Commercial banks do not control the aggregate amount of reserves (they cannot create or destroy them; they can only transfer reserves among themselves), while central banks do not control the banks’ deposits (banks create them when they grant loans or buy assets). Further, the answer depends on the sense of the concept of money. For example, money is endogenous in the origin sense, as it evolved spontaneously in the marketplace, but it is exogenous in the evolutionary sense, as fiat currencies were imposed by governments rather than spontaneously emerging within markets. Endogeneity/exogeneity also depends on the model (in a statistical sense, money is exogenous in the quantity theory of money but endogenous in other theoretical frameworks) and the monetary system (for example, money is exogenous under the gold standard, in the sense that the money supply is outside the control of authorities, but endogenous in the fiat system, in the sense of being elastic—that is, being subject to central banks’ responses to the needs of the economy14). It depends on the institutional setting. For example, Niggle (1990) distinguishes five stages of the development of modern monetary and banking institutions. He argues that money was exogenous in stage one (an economy with a commodity money, such as gold, or a convertible paper currency with strict gold-backing rules) and three (an economy with a central bank that effectively imposes reserve requirements on the banks). Stage three is an approximation of the US monetary system in the decades immediately after the establishment of the Federal Reserve System. However, later innovations in the form of liability- and asset-management techniques allowed commercial banks to relax the Fed's constraints on lending and money creation. Moreover, the Fed assumed the role of lender of last resort for the system as a whole, which makes it unable to constrain bank-reserve growth. Hence, “under these circumstances, even bank reserves are determined endogenously, as the central bank acts to stabilize financial markets by accommodating the banks’ need for reserves adequate to cover their loan-created deposits” (Niggle 1990, 448).15 Finally, the answer depends on the central bank's behavior. The central bank might implement both active and passive policies.16 It can sometimes accommodate the private sector's demand and react passively to changes in the economy and banks’ behavior, but it can also initiate changes in the marketplace. Implementing large-scale asset purchases might be the best example of central banks’ actions on their own initiative, as they go beyond the passive reactions that merely accommodate the banks’ demand. Hence it is practically impossible to provide any meaningful contribution to the debate without reference to a particular context. Money has both endogenous and exogenous aspects. It is not completely endogenous, as this would exclude hyperinflations, which occur from time to time when monetary authorities increase the money supply excessively.17 Such instances clearly show that central banks are not merely passive institutions that only accommodate the demand for money. However, money is not fully exogenous either, as that would imply that central banks are powerful while commercial banks are powerless. But we know that monetary policy, at least the traditional one, operates through the banking system. When it broke down in the aftermath of the global financial crisis, central banks implemented unconventional monetary policies. Fully exogenous money would also imply that financial crises are caused by a decline in the monetary base. But the monetary base did not drop in any meaningful way before the Great Recession. As Fig. 2 shows, the monetary base in the US (and other advanced countries) increased enormously in the aftermath of the global financial crisis, as central banks tried to counteract, effectively or not, the endogenous contraction in total credit to the private nonfinancial sector caused by the financial crisis. Another problem with the debate about the endogeneity/exogeneity of money is that the economists involved in it, even if they are right in some respects, often make too far-reaching conclusions. In particular, the post-Keynesians are right when they criticize the money-multiplier model. It is true that the impact of monetary policy is less mechanical and more subtle than the monetarist textbooks seem to describe. However, the fact the model is erroneous does not imply that the money supply curve is horizontal or that central banks are powerless to affect the money supply.18 Central banks do not affect the loan supply directly through the multiplier, but they do affect it indirectly through the demand for loans, creditworthiness of borrowers, value of collaterals, expected bank profitability, the risk-taking channel of the monetary transmission mechanism, and more. So they do not directly determine the money supply by controlling the reserve-requirement ratio and the size of the monetary base, but only indirectly by setting interest rates. Similarly, from the fact that the amount of loans depends on the demand for credit it does not follow that commercial banks merely react passively to borrowers’ demands for loans. Actually, they are active agents that set credit conditions, buy securities, or make strategic decisions. They create money, so macroeconomists should pay more attention to their behavior. However, it does not follow that commercial banks can create money without any restrictions or that the supply of money is purely credit-demand driven. Central banks try to reach certain macroeconomic goals, and they set the policy-rate targets accordingly. The resulting monetary policy actions somehow constrain the money-creation process. In other words, it is true that central banks can accommodate the growth in bank lending, in line with the endogeneity of bank money, according to which commercial banks grant loans first and look for reserves only later. This is certainly possible, and there are good reasons to believe that this is how central banks very often indeed behave. But it is not an inevitable result; rather, it is a matter of central bank policy. So Wray (1990) might be right in writing that deposits make reserves as long as the central bank decides to deliver the latter. Hence, the money supply is endogenous in the control sense, but only because the central bank allows it to adjust.19 Analogously, the fact that the central banks affect the money supply is not incompatible with the role ascribed to commercial banks in endogenous-money-creation theory, according to which deposits are created whenever banks grant loans (Chick and Dow, 2002). Therefore, in a sense, the whole debate about endogeneity/exogeneity of money misses the point. The dichotomy between central banks’ affecting the economy through the money supply or through interest rates is false. Monetary policy “generally consists of a combination of quantity effects and price effects, none of which is deterministic” (Chick and Dow, 2002, 605). What really matters is not what the central bank's target is and whether money is exogenous or endogenous in a control sense, but whether money is created out of thin air. In our monetary system, when the bank advances the entrepreneur a loan, it is not the transfer of existing purchasing power, but “the creation of new purchasing power out of nothing … which is added to the existing circulation” (Schumpeter, 1934, 106). This is of great importance because money is created with the simultaneous creation of an asset and a liability by the banking sector. The implication is that—contrary to mainstream economics, in which credit remains invisible and irrelevant, as commercial banks are considered merely intermediaries—changes in the aggregate level of debt are of great macroeconomic importance and credit expansion endangers financial stability. In other words, what really matters is not whether central banks target the supply of money or interest rates, but whether they affect the money supply at all—whether the quantity of money is determined solely by market forces. Contrary to the real-bills doctrine and post-Keynesian monetary theory, the fact that money is endogenous does not mean that credit expansion, and the resulting malinvestments and misallocation of resources, is endogenous to the market economy. Actually, the opposite is true. As Hayek (1937) notes, cyclical activity is a standard feature of an economy with an elastic currency—that is, an economy in which the supply of money either wholly or partially responds to changes in the demand for money or the demand for credit. This is an important argument against loose monetary policy and credit expansion. No matter whether the increases in the money supply are endogenous—whether they merely respond to the needs of trade—they still distort the economy.20"],["Integrated agricultural-nutrition programs are often implemented under the premise that program effects are durable and spillover. This paper estimates one year post-program effects, three-year aggregate program effects and spillover effects using treated and untreated household cohorts. Two treatment interventions implemented agricultural interventions with behavior change communication strategies varying implementers using either village health committees or older female leaders. In the post-program period, program effects deteriorated relative to program period impacts documented in Olney et al. (2015), but the three-year agricultural, nutrition knowledge, health care practices and severe anemia impacts remained statistically significant. Despite the non-rival nature of nutrition education and promoted production techniques, there is little evidence of agricultural technology or health knowledge spillovers to non-treated households within treatment communities. Spillover effects measured for appropriate treatment of diarrhea (10 pp increase in giving rehydration salts rather than traditional medicine), wasting (20 pp lower probability of wasting) and children's anemia status (7 pp reduction in severe anemia) significantly improve in later cohorts. The aggregate program effects and spillovers are generally robust to multiple hypothesis testing. --------------------------------------------------------------------------------","Post-program and spillover effects are often intended consequences of agriculture and nutrition programs designed to increase their potential effect over time (World Bank, 2007). Programs may have spillover effects if they generate non-rival goods which diffuse in communities or, to the contrary, lead to disadoption after the program implementation period because monitoring, knowledge decay or the return on the productive or nutritional investment is no longer positive. The technology adoption literature provides many examples of the diffusion of new technologies (Munshi 2004, Bandiera and Rasul 2006, Beaman et al., 2015), the diffusion of knowledge about new technologies (Beaman and Dillon, 2018) and the disadoption after the initial period take-up (Moser and Barrett, 2006, Duflo et al., 2011, Gilligan et al., 2014). The behavior change communication (BCC) literature in nutrition also documents increased knowledge, yet low adoption of improved infant and young child feeding (IYCF) practices, in response to nutrition education campaigns (World Health Organization, 2010). Estimation of post-program and spillover effects due to nutrition interventions is limited in the developing country context (Benjamin-Chung et al., 2017), though in developing countries health spillovers between siblings (Ho, 2017) and from health education programs (Mora et al., 2015) have been documented. This paper builds on these two literatures by, first, testing whether program impacts endure in the first year after program implementation ends among treated households, and second to test for spillover effects by estimating whether implementation period agricultural and nutritional program impacts estimated in Olney et al. (2015) are also measured in non-treated households in treatment villages. We disentangle the mechanisms of nutritional post-program effects and spillovers by estimating treatment effects on nutritional ‘inputs’ such as mothers’ health and nutrition knowledge and agricultural production mechanisms independently. To estimate spillovers of agricultural production techniques, health knowledge, and children’s nutritional status, we build on a cluster-randomized control trial conducted in partnership with Helen Keller International (HKI) in Burkina Faso. From 2010–2012, HKI implemented an Enhanced Homestead Food Production (EHFP) program in Burkina Faso with the specific objectives of improving women’s agricultural production of nutrient-rich foods and children’s nutritional status.1 Two treatment groups received a standardized production intervention and a nutrition behavior change communication (BCC) intervention with the only difference between the two treatment groups being the BCC implementers: either older women leaders (OWL) or health committee (HC) members. The production interventions were standardized across each treatment group focused on knowledge of the nutritional content of specific foods that were locally grown to encourage adoption, access to inputs, including seed and tools, to establish village gardens, and materials to encourage the establishment of a homestead garden. It was expected that the program would have the biggest impact on improving the nutritional status of children during the first 1,000 days of life, therefore households with children 3–12 months of age were targeted by the program. Mothers with children 3–12 months of age living in the treatment villages were invited to participate in the EHFP program, along with their husbands and children. Our previous analysis of the program impacts (Olney et al., 2015) was based on the analysis of a baseline study that was conducted between February and May 2010 and an endline study that was conducted between February and June 2012. In 2012, during the endline study, an additional cohort of younger children 3–12 months old were sampled in treatment and control villages. A post-treatment follow-up survey was also administered between March and June 2013 to track this younger cohort of children, their mothers and households, in addition to the original households interviewed in treatment and control villages at baseline. The EHFP program was successful in improving women’s agricultural production during the program implementation period, specifically women’s production of vitamin A-rich foods. No additional agricultural treatment effects on treated households or non-treated households were found after the program implementation period. Women’s use of manure and fertilizer decreased in the post- program period, though manure use and number of plots cultivated by women were still statistically significantly higher over the three-year period of the study in treatment as compared to control villages. Increases in women’s production of vitamin A-rich crops found during the program implementation, however, were not sustained in the post-program period. These results suggest that program extension advice, monitoring and input provision were critical to the small production gains found during the program implementation period. No spillover effects on new cohort households were found in terms of utilization of inputs or production outcomes. We also found positive program effects on mothers’ knowledge of optimal IYCF and hygiene and healthcare practices as compared to non-treated mothers. The analysis of the post-program period impacts suggests that mothers’ health and nutrition knowledge deteriorated after the end of the program, though knowledge gains were statistically significantly higher among treated as compared to non- treated mothers for the over the three-year period of the study. Spillovers on health care practice knowledge and exclusive breastfeeding practices for children less than six months of age were estimated between treatment village cohorts. We estimate a 23 percentage point increase in later OWL treatment village cohorts on exclusive breastfeeding of children less than 6 months old. In the HC and OWL later treatment cohorts, there was a 10–11 percentage point decrease in treating diarrhea with traditional medicine and in the HC treatment cohorts, a corresponding 10 percentage point increase in giving rehydration salts for diarrhea treatment. With respect to children’s nutritional outcomes, the primary findings during the program implementation period were that the EHFP program improved hemoglobin concentration and decreased the prevalence of anemia, wasting and diarrhea among children living in HC villages relative to those living in control villages (Olney et al., 2015). In the post-implementation period, no further increases in hemoglobin concentration were found. However, the relative increase in hemoglobin concentration that was found among children 3–5.9 months of age at baseline were sustained one-year after the program ended, although the impact estimate for the three-year study period was only marginally statistically significant. No spillover effects were measured on the new cohort of children 3–12 months of age at endline in the post-implementation period in terms of their anemia status or hemoglobin level. During the program (2010–2012), the prevalence of wasting among children 3–12 months of age at baseline living in the HC villages relative to those living in the control villages decreased significantly by 8.4 percentage points. In the three-year estimates, the prevalence of wasting in HC villages had declined by 5.4 percentage points more as compared to control villages, but this difference was not statistically significant. In addition, there is some evidence that there were significant protective effects of the program on preventing wasting among children born into treated households during the program period. In control villages, the prevalence of wasting increased from 2012 to 2013 among children 3–12 months of age. In 2012, there was no change in wasting in OWL villages, and in HC villages there was a decline, resulting in 19.5 and 20.4 percentage point differences between each of the treatment groups and the control group in the prevalence of wasting from 2012 to 2013. These differences were statistically significant and indicate a potential role of the EHFP program in protecting children’s nutritional status during this vulnerable age. Because there are multiple hypotheses for agricultural, knowledge, practices, or nutritional set of measures, we correct all estimates using multiple hypothesis testing following Anderson (2008) to assess the robustness of our results. Among the statistically significant effects estimated for the post-program period, three year overall program effects or spillover effects, we report corrected p-values that account for multiple hypothesis testing. The overall program effects are robust to multiple hypothesis testing (9 out of 11 statistically significant results), while only half the post-program effects are statistically significant after multiple hypothesis corrections (3 out of 6 results). The spillover effects are generally robust to multiple hypothesis corrections (4 out of 6 results). In the next section of the paper, we present the experimental design, while the third section details the post-program and spillover effects identification strategy. The fourth section presents the results, while the last section concludes. Program description ~~~~~~~~~~~~~~~~~~~ HKI’s EHFP program implemented in the province of Gourma in Burkina Faso consisted of an agricultural production component and a nutrition and health BCC strategy. The agricultural production activities included input distribution (e.g. seeds, saplings, chicks and small gardening tools) and agricultural training provided by female village farm leaders at village model farms. Program beneficiaries started their own homestead food production activities after the receipt of inputs and training. Production activities primarily included the promotion of micronutrient-rich fruits and vegetables, eggs and poultry production. The BCC strategy was designed using the essential nutrition actions framework that focuses on seven practices: women’s nutrition, anemia prevention and control (e.g. intake of iron-rich foods and use of bednets to prevent malaria), iodine intake, prevention of vitamin A deficiency, breastfeeding practices, complementary feeding practices, and nutritional care for sick and severely malnourished children (Guyon et al., 2009). Beneficiary women received bimonthly home visits from either an older woman leader (OWL) or a health committee (HC) member who implemented through these visits a behavior change communication curriculum during which beneficiaries were taught about optimal health and nutrition practices, and discussed successes and challenges related to the adoption of these practices. Study design and participants ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The EHFP program was evaluated using a cluster-randomized controlled trial2 to estimate the effects of the program on agricultural production, knowledge, practices and ultimately nutritional outcomes of young children3 . Villages in four departments in the Gourma province (Diapangou, Diabo, Tibga, Yamba) were selected for inclusion if they had access to water in the dry season to enable participation in the agricultural intervention. Fifty-five villages were stratified by department and village size and randomized into three groups: 1) control group which received no interventions (25 control villages); 2) OWL group that received the agricultural production intervention with the BCC strategy implemented by older women leaders (OWL) (15 OWL villages); and 3) HC group that received the same agricultural production intervention with the BCC strategy implemented by health committee (HC) members (15 HC villages). The two types of behavior change implementers were selected because the effectiveness in improving knowledge and eliciting behavior change may vary by the type of implementer delivering the messages. In rural areas, HC members delivered health and nutrition interventions and could facilitate direct links with health services, while OWLs were the main providers of ante- and post-natal counseling and delivery care and thus could be more influential in changing infant and young child feeding and care practices. Within the selected villages (n = 55) all women with children 3–12 months of age were invited to participate in the study.4 Trained fieldworkers explained the study to eligible households and informed consent was obtained from either the household head or the mother of the selected child. We used a longitudinal design and followed the same households, mothers and children over the two-year program implementation period. The baseline study was conducted between February and May of 2010 (when children were 3–12 months of age), and the endline survey between February and June of 2012 (when children were 24–40 months of age). In 2012, during the endline study, an additional cohort of younger children 3 to 12 months of age was sampled in treatment and control villages from both treated households, who participated in the EHFP program or impact evaluation, and from a new cohort of non-treated households in both treatment and control villages. A post-treatment follow-up survey was administered between March and June 2013 to track this younger cohort of non-treated children, and their households, in addition to the original treated households, mothers and children interviewed in treatment and control villages at baseline and endline. Thus, we have a cohort of treated households, mothers and children 3–12 months of age at baseline interviewed at baseline in 2010, endline in 2012 and follow-up in 2013. In addition, we have a cohort of younger non- treated children 3–12 months of age at endline born into the treated and non-treated households in treatment and control villages. This younger cohort and their households were surveyed in 2012 and again in 2013. Half of these younger cohort children are siblings with the same mother, and the remaining are siblings with a different mother (co- wife) and cousins due to the frequency of polygamous household structures. All age eligible children were interviewed in both the original and follow-up cohorts.","Three treatment effects are specified to measure different types of program effects after the program implementation period (2010–2012): a post-program effect (2012–2013), a three- year effect of the program (2010–2013) and spillover effects to non-beneficiary households. The specifications are estimated with the following regressions for the 2010 cohort restricting the sample to treated households in OWL and HC villages and control households from the 2010 cohort: The identification of spillover effects in our study builds on the recent literature (Baird et al., 2018; Angelucci and Giacomo De, 2009; Miguel and Kremer, 2004)) which uses the exogenous variation in program assignment and eligibility status. As households were screened for treatment eligibility if households had a child between 3–12 months at baseline and the new cohort selected in 2012 used the same survey selection criterion, the group on whom spillovers would occur is well defined in our data. A strength of the above econometric strategy is the direct testing of production, knowledge, practice and nutrition outcomes which the program was designed to impact. One statistical concern is that because there are multiple primary and secondary effects of an integrated agriculture-nutrition intervention, our results could over report false positive effects without corrections for multiple hypothesis testing. To address this concern, we first present all treatment effect results. We follow Anderson (2008) to correct treatment effect p-values within agricultural, knowledge, practices and nutritional outcome groupings among related outcome measures that are statistically significant to assess their robustness to multiple-inference corrections. The paper’s main results and their robustness with Anderson’s q-values are presented in Table 5. Balancing tests and descriptive statistics ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To test whether the sample has statistical balance in observable characteristics between the treatment and control groups, balancing tests on household- and child-level characteristics are estimated for the 2010 and 2012 cohorts in Tables A1 and A2 in Supplementary material, respectively. We report household characteristics by treatment group, HC or OWL, and control, as well as the p-value of the test of mean equality between the three groups for the non-attrited sample. Household characteristics such as household size, value of men’s and women’s assets, and household head’s and mother’s educational status were all balanced across the three groups at initial randomization in 2010. One housing characteristic, having a dirt floor, was not balanced between the treatment and the control groups, with higher prevalence of dirt flooring in HC and OWL villages, indicating a lower level of housing quality. Children’s characteristics, including the gender and age, were balanced at baseline, though children’s nutritional indicators were not. Household attrition does not seem to be a significant source of observable characteristic imbalance, though there are differences in attrition rates between the treatment and control groups. In Table A2 in Supplementary material, we test the mean equality of the same household and child characteristics at selection for the 2012 cohort7 . Though villages were not re-randomized in 2012 as the objective of collecting information was to estimate spillover effects, we compare mean characteristics across the treatment and control groups to assess whether treatment status potentially affected household composition and welfare characteristics, apart from the nutritional, production or knowledge indicators on which we could potentially observe a spillover. Table A2 in Supplementary material reports that household characteristics, value of assets and livestock owned, as well as children’s characteristics were well balanced across the treatment and control groups. Children’s nutritional indicators were not all balanced across groups, particularly children’s hemoglobin status. We test for balance between the old and new cohorts and report those results in the Table A3 in Supplementary material. Due to variation in agricultural production and disease environment between years, it is not necessarily expected that cohort characteristics will be balanced between years. The balancing tests demonstrate that household and child nutritional indicators are not balanced between the cohorts while characteristics such as child age, child gender, household size and women’s assets are balanced between cohorts. Differences between cohorts motivate the inclusion of time fixed effects in the between cohort spillover specification described above. Post-program, three-year program, and spillover effects ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ HKI’s EHFP program in Burkina Faso was designed to improve household agricultural production of nutrient-rich foods and the nutritional status of the targeted children who were 3–12 months of age at the time of the baseline study through an integrated set of production and nutrition interventions. We estimate treatment effects of the program on the primary outcomes – household agricultural production of nutrient-rich foods, women’s health and nutrition-related knowledge, IYCF practices, and children’s nutritional status including anemia and anthropometric measures. As expected, HKI’s EHFP program was successful in increasing women’s agricultural production during the program implementation period, particularly women’s production of vitamin A-rich foods (Table A4 in Supplementary material). Although there was no overall impact on household or on men’s production, impacts on women’s agricultural production were fairly consistent, albeit small, across the different types of crops. These results are consistent with other studies that have shown small impacts of small-scale agriculture interventions on household food production (World Bank, 2007; Masset et al., 2012; Olney et al., 2009; Dillon et al., 2019). In the post-program period, women’s use of manure and fertilizer decreased in the post- implementation period, though manure use and the number of plots cultivated by women were still statistically higher over the three-year period of the study. Increases in women’s production of vitamin A-rich foods found during the program implementation period were not sustained in the post-implementation period (Table 1-Panel A). While the number of plots cultivated by women and manure use were statistically higher in the treatment groups relative to the control over the 2010–2013 period, one-year post-program implementation agricultural production effects were not sustained (Table 1-Panel B). These results suggest that program extension advice, monitoring and input provision were critical to the small production gains found during the program implementation period. No spillover effects on non-beneficiary households were measured (Table 1-Panel C), suggesting that if agricultural programs are expected to improve children’s nutritional status through increased food availability, the scale of the program and post-program extension advice are likely important program design features. In terms of health and nutrition-related outcomes, we found a significant impact of the EHFP program on improving mothers’ health and nutrition-related knowledge in both treatment groups compared to the control group during the program implementation period (Table A5 in Supplementary material). Treatment effects for the post-implementation period suggests that mothers’ knowledge deteriorated after the program implementation, though overall knowledge gains were significant and substantial for the three-year period (Table 2-Panel A and B). The only statistically significant declines were with respect to knowledge of vitamin A- rich foods like dark, green leafy vegetables in the HC group and eggs in both the HC and OWL groups. Overall program effects including one-year post-implementation on mother’s health and nutrition- related knowledge remained significant and ranged from 13 to 24 percentage point change in the OWL group and 12–21 percentage point change in the HC group (Table 2-Panel B). Following on the improvements in mothers’ knowledge, our analysis of birth cohorts at endline indicated that mothers living in OWL villages were more likely to have exclusively breastfed children less than 6 months old as compared to mothers living in control villages (Table 3- Panel A). However, these estimates are suggestive as they do not track the same cohort of birth mothers over time, so we cannot rule out that the composition of the endline group may have also influenced some of the measured differences in IYCF practices. A significant between cohort spillover effect was also measured among non- treated mothers in the new cohort OWL villages. We estimate an 23 pp increase in exclusive breastfeeding children less than 6 months of age and a 2.4 pp reduction in feeding a child using a bottle as compared to mothers in control villages (Table 3- Panel B). In any integrated agriculture and nutrition intervention targeted to children, the primary outcomes of interest are children’s nutritional status indicators. The primary findings during the program implementation period indicated improved hemoglobin concentration, and reductions in the prevalence of anemia and, wasting among children living in HC villages relative to those living in control villages (Table A5 in Supplementary material). In the post-implementation period, no further increases in hemoglobin concentrations were measured or reductions in the prevalence of anemia or severe anemia (Table 4-Panel A). In Table 4-Panel B, most of these nutrition impacts were sustained one year after the end of program implementation. Children 3–12 months of age at baseline living in HC villages were 5 percentage points less likely to be severely anemic as compared to those living in control villages. Among children 3–5.9 months of age at baseline living in HC villages, hemoglobin concentration was higher, and the prevalence of severe anemia was 9 percentage points lower compared to children 3–5.9 months of age living in control villages. To measure spillovers within households, the effect of the intervention on younger siblings within treated households is estimated in Table 4-Panel C. Among the younger siblings of the target children (children 3–12 months of age in beneficiary households in 2012), there was a protective effect of the EHFP program on the normal increase in the prevalence of wasting that is seen in this population over the first two years of life. Whereas the prevalence of wasting increased from 2012 to 2013 among children that were 3–12 months of age in 2012 in control villages, in OWL villages there was no change and in HC villages there was a decline, resulting in a 20 percentage point difference in each of the treatment groups and the control group. These differences were statistically significant and indicate a potential role of the EHFP program in protecting children’s nutritional status during this vulnerable age. We observe a rise in anemia in siblings of HC treated children (11 pp increase) and among siblings of HC treated children 3–5.9 months (18 pp increase), though the full sample treatment effect is not statistically different than zero (p-value = 0.11). No other spillover effects were measured for the new cohort of children with respect to their nutritional status estimating the spillover using a between village specification (Table 4-Panel D). Estimating within village cohort spillover effects, non-treated children living in HC villages were 7 percentage points less likely to be severely anemic compared to treated children assessed at baseline (Table 4-Panel E). Robustness of main results to multiple hypothesis testing corrections ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As we estimate different treatment effects over varying time periods and multiple spillover effects, Table 5 presents a summary of the main statistically significant results by time period, noting the effect size and pvalue of the null hypothesis that the coefficient estimate is statistically different than zero for statistically significant treatment effects. Table 5 also presents the corrected q-values for multiple hypothesis testing following Anderson (2008). The overall program effects are robust to multiple hypothesis testing (15 out of 21 statistically significant results where p-values rejected the null hypothesis that the treatment effect was equal to zero). The sustained impacts on anemia and hand washing were not robust to multiple hypothesis testing, while the results on women’s agricultural land holdings, inputs, breastfeeding knowledge, and nutrient rich food knowledge were robust to multiple hypothesis testing. Three out of the seven post- program effects are statistically significant after multiple hypothesis corrections. The results which were not robust to multiple hypothesis testing were the results on mother’s knowledge of vitamin A rich foods (dark leafy vegetables and eggs) and treating diarrhea using traditional medicine. Results on agricultural inputs and breastfeeding knowledge were robust to multiple hypothesis testing among the post-program effects. Five out of the eight statistically significant spillover effects estimated are also robust to multiple hypothesis correction. The results that were not robust to multiple hypothesis testing included the within and between cohort results on breastfeeding and the anomalous result on increased anemia prevalence among younger siblings. The spillover effects on the treatment of diarrhea, anemia and wasting remained robust to multiple hypothesis testing.","This paper demonstrates the importance of post-program effects in evaluating the impact of an integrated agriculture and nutrition program where disadoption dynamics, knowledge deterioration or spillovers may affect overall program impact assessment. Despite deterioration of some of the program implementation period effects and limited production and knowledge spillover effects, we found a significant impact of an agriculture and nutrition program on improving children’s hemoglobin concentration during the program period and sustained impacts one-year after the program ended. We also demonstrate a potential protective effect of the program on younger children born into beneficiary households during the program period on preventing wasting during the first two years of life. While no impacts across treatment cohorts were estimated on children’s anthropometric outcomes, within HC villages, younger children were statistically less likely to be severely anemic. Production or nutrition BCC mechanisms provided some explanation for these sustained and spillover effects. We found small effects from increased agricultural production over the program implementation period and limited spillovers to non-treated women from production interventions in the post-program implementation period. Knowledge related to some optimal health and nutrition practices deteriorated during the post-implementation period, though it remained significantly higher than knowledge among mothers in control villages over the three-year study period. Overall, program spillovers are limited with respect to health and nutrition-related knowledge, except for an important decline in the belief that traditional medicine should be used to treat children’s diarrhea. Finally, we found a statistically significant child feeding spillover effect on exclusively breastfeeding children under 6 months of age, though this result is not robust to multiple hypothesis testing corrections. Despite significant positive effects during the program implementation period on hemoglobin concentrations, several factors may explain the lack of impact on other nutritional outcomes with respect to the between village spillovers. First, agricultural production impacts one year after program implementation were not detected in our sample. Given these disadoption dynamics, availability of nutrient-rich foods did not significantly change for beneficiary households or non-beneficiary households in treatment villages. Without program support, the agricultural component implemented over the two-year project was not able to sufficiently resolve the production constraints including access to water during the dry season, faced by rural Burkinabe producers. Second, agricultural production requires labor that women with young children may not have been able to sustain after the initial program period. Without program monitoring, the incentives to provide labor may have decreased, resulting in lower production. With respect to the role of nutrition education in sustaining program impacts or promoting spillovers, the impact estimates clearly show that the modality of providing nutrition education either through health committees or drawing on the experience of older women leaders was critical to sustaining impacts and spillovers within beneficiary households and across cohorts. In Olney et al. (2015), initial program period effects are likely explained by the mixed gender health committee’s effectiveness in communicating nutrition messages, but also potentially by having more technical health knowledge to connect mothers and children to the formal health system. If this is the case, sustained effects and spillovers are potentially explained through the program’s mobilization and demonstration of new roles for the health committee and OWLs in monitoring and providing information to new mothers. These nutritional resource people including both village health committee leaders and older women leaders remained in treatment communities which could explain the presence of and differences in within village spillovers. Further research on why different types of leaders led to differences in spillovers is important to understanding the role of BCC on nutritional outcomes. The potential integrated connectivity of households within treatment villages either to other households or to those implementing the nutrition education components of the program will be an area of future research which we will explore to further understand the mechanisms driving the anemia and wasting spillover effects we measure in this paper."],["Focusing on groups of high-income countries (OECD, EU, and EMU), this study shows that the finding of a non-linear, hump-shaped impact of financing on economic growth is robust to controlling for financing composition in terms of the sources (bank credit, debt securities, stock market) and the recipients of finances (households, non-financial and financial corporations), or both. In particular, we obtain the following results. First, the non-linear impact of total bank credit is more pronounced than that of either household credit alone, or the sum of bank credit, debt securities, and stock market financing. Second, credit to non-financial corporations tends to have a positive, while credit to households a negative impact on growth, even after allowing for non-linearities. Third, debt-securities and stock market-based financing have a different impact on growth. Finally, the estimated turning point of the non-linear relationship is close to that found by Cournède and Denk (2015) for the OECD countries, and lower than that established by Arcand, Berkes, and Panizza (2015) for a broad set of countries. --------------------------------------------------------------------------------","The relation between financial development and economic growth is much debated. As was hypothesized by Schumpeter (1934) and supported by King and Levine (1993) with numerous papers thereafter, differences in the level of the development of financial systems affect economic growth differentials among countries. The impact channels vary from additional financial funds, available to finance investment projects due to larger volumes of savings, to more efficient reallocation of funds, thus reaching the right entrepreneurs and leading to higher productivity (see e.g. Beck et al., 2000; Aghion et al., 2005; Levine, 2005). The early empirical literature (see overviews ibidem or Panizza, 2014) suggested a positive association between financial development and economic growth, the former measured by e.g. the amount of domestic private credit or stock market capitalization relative to gross domestic product (GDP). The dominant positive attitude towards financial expansion encouraged a sharp increase in financial penetration, and the median level of private bank credit (in higher income countries with data reported by the Bank for International Settlements, BIS hereafter) constituted around 90 percent of GDP in 2014. In a number of countries, it has reached levels much greater than their GDP (see Fig. A1 in Appendix A for some selected countries). Such high levels of financial penetration, together with recent and contemporary financial crises started casting doubt on the benefits of such a degree of financial deepening (see e.g. Beck, 2012). The corresponding more recent empirical work provides evidence of either a vanishing positive impact (as e.g. in Rousseau and Wachtel, 2011), or a potentially non-linear (often an inverted U-shape) relationship as documented in numerous contemporary studies.1 Although this relationship can be complex and may vary, among others, with a country's level of economic development and quality of financial institutions (Cecchetti and Kharroubi, 2012; Demirgüç-Kunt et al., 2013; Masten et al., 2008; Rioja and Valev, 2014), the particular functions performed by the financial sector (Beck et al., 2014), the speed of expansion of financial sector (Cecchetti and Kharroubi, 2012; Ductor and Grechyna, 2015), the ‘normality’ of the period under investigation (Balta and Nikolov, 2013; Breitenlechner et al., 2015; Gambacorta et al., 2014), or the strength of patent protection (Chu et al., 2016), the high current levels of financial penetration and the recent findings of a non- linear impact of financial development on economic growth point to a potential of ‘too much finance’ in many countries, thus questioning the desirability of large financial sectors. Such findings are typically established using total finances (mostly: credit), and the apparent non-linear impact of totals can stem from a substantial structural change in the composition of finances, that has been taking place during the recent decades. Though there are some studies going beyond total finances, they usually look at the impact of certain financing components separately or using ratios, which may bias the estimation and lead to incorrect conclusions. The main aim of this paper is to advance on this front, and consider both the composition of finances and its potentially nonlinear impact. We incorporate both the sources (bank financing, debt securities financing, and stock market financing) and the recipients of finances (households, non-financial corporations, and financial corporations). We believe that it is important to include them in this analysis for a number of reasons. First, there is already evidence that the type of financing may matter. Different sources of finances (bank-based versus market-based financing) can have an uneven impact (see e.g. Beck and Levine, 2004; Cournède and Denk, 2015; Demirgüç-Kunt et al., 2013; Gambacorta et al., 2014; Langfield and Pagano, 2016; Mishra and Narayan, 2015). At the same time, fund recipients (users of finance) might also matter nontrivially for the outcome. For instance, Chu and Cozzi (2014) develop a model in which the nature of liquidity constraints (whether they affect consumers or firms more) matters for the performance and optimality of macroeconomic policies. Beck et al. (2012) stress that a substantial household credit expansion might be hurting economic growth. In parallel, Bezemer et al. (2014) point out that the share of credit to nonfinancial business credit decreased sharply since 1990, while it had a significantly positive effect on growth. Along these lines, although warning for a small sample size, Arcand et al. (2015) indeed find that the non-linearity of household credit is more significant than that of firm credit. Second, the analysis of the importance of financial structure is currently quite limited. The impact of different components of financing is mostly analyzed individually or looking only at a few of them (see e.g. Cournède and Denk, 2015), thus creating potentially an omitted variable bias. Even when the analysis is performed including several subcomponents together (see e.g. Gambacorta et al., 2014), the difference between their individual and joint impact (e.g. that of total financing) is not investigated. Besides, though the dependence of economic growth rates on bank credit financing and stock market financing is often analyzed, the influence of debt securities is rarely considered. Moreover, when it is, like in Langfield and Pagano (2016), the stock market and debt securities financing is often merged, which might impose an incorrect restriction and lead to biased inference. Finally, to our knowledge there is no study that jointly and not individually investigates the impact of both the sources (bank financing, debt securities financing, and stock market financing) and the recipients of finance (households, non- financial corporations, and financial corporations), not to mention also the non- linearity. Last, but not the least, the changing structure of financing can lie behind the vanishing or non-linear impact of finance on economic growth;2 therefore, it is crucial to investigate if the impact remains non-linear after controlling for the detailed structure of finance that accounts for potential changes.3 Third, understanding of the composition- driven versus the total-finance-driven impact on growth can shed some light on the desirability of certain economic policies. If ‘too much of’ total finances created problems, the curbing of the whole financial sector could be considered. If the bank credit caused the troubles, the desirability of taxation systems promoting the debt-biased finance should be questioned. If non-linearity and reduction of economic growth rates at some point were connected with some structural-shift in the composition of finance, some structural and/or competition policies could be invoked to promote the favorable composition, and so on. Another main feature of our analysis is our focus on a homogenous sets of (developed) countries, the OECD, EU and EMU in particular. This is important because the empirically identified non-linearity can be an artefact of mixing different groups of countries. For instance, Karagiannis and Kvedaras (2016) show,4 using the original Arcand et al. (2015) data set, that their non-linearity finding vanishes when considering more homogeneous sets of countries (such as that of the OECD, or the European Union).5 Nevertheless, some other recent research (see e.g. Cournède and Denk, 2015; Cournède et al., 2015; and Samargandi et al., 2015) have also concentrated on smaller sets of more homogeneous countries like the OECD or middle-income developing countries, and found significant non-linearity. It is of further interest therefore to investigate whether similar results hold for the EU countries and/or the founding member states of the European Monetary Union (EMU1999). These groups are interesting also because they are quite homogeneous in general as well as in terms of financing structure in particular, namely, they have strongly bank-biased financing (Langfield and Pagano, 2016). The usage of a smaller number of more homogeneous countries and the need of detailed financial series limit the number of observations, and influence the choice of the econometric methodology that can be properly employed in our case. However, in order to be more confident in the obtained empirical results, we do not restrict ourselves only to the EU and EMU 1999 member-states samples, but also provide the results for a broader set of countries; namely, the OECD countries where the required data are available. This does not only enable us to compare our findings obtained using a different methodology with the already available ones (namely, Cournède and Denk, 2015; Cournède et al., 2015), but also allows us to be more confident in the results obtained for the EU and EMU member states, given that the established patterns are fairly robust across all investigated groups of countries. Our econometric research strategy is to start from a simple log-linear specification with only few financial variables, and then to introduce richer specifications with more detailed structure and/or non-linearity. Namely, we first consider the impact of different types of finances (distinguished by their source and/or recipient). This includes evaluating whether it is sufficient to use various ratios (like bank credit to stock market, or bank credit to the sum of stock market and debt securities, as e.g. in Demirgüç-Kunt et al., 2013, or Langfield and Pagano, 2016), or additional disaggregation is required due to the non-homogeneity of the impact (for such evidence see e.g. Kaserer and Rapp, 2014). Then we add a non-linear term to these specifications. These results can be read in two main ways. On the one hand, they assess whether the findings about the different impacts of different types of finances are robust to the inclusion of non-linearities. On the other hand, they also explore whether a non- linear impact remains valid after controlling for the types of finances. Finally, we investigate the source of the non-linear effect, whether it is due to total finances, bank credit or household credit. This last question is motivated by the findings of Beck et al. (2012), among others, and was investigated by Arcand et al. (2015). We obtain the following results, which prove to be quite stable in our extensive robustness analysis. First, the non-linear impact of total bank credit is more pronounced than that of either household credit alone, or the sum of bank credit, debt securities, and stock market financing. Second, credit to non-financial corporations tends to have a positive, while credit to households a negative impact on growth, even after allowing for non-linearities. Third, debt-securities and stock market-based financing have a different impact on growth. Finally, the estimated turning point of the non-linear relationship is close to that found by Cournède and Denk (2015) for the OECD countries, and lower than that established by Arcand et al. (2015) for a broad set of countries. The paper is structured as follows. Section 2 discusses data sources and variables. Section 3 characterizes the econometric modelling approach. Section 4 presents and discusses the main empirical findings and Section 5 concludes. Finally, some further details and robustness checks are delegated to the Appendix.","In order to evaluate the effects of the composition of domestic private finance on economic growth and their potential role in the non-linear impact of finance on growth, we need disaggregated data on the split of financing by the source (bank, debt securities, and stock market financing) as well as the recipient (households, non-financial firms, and financial corporations). For this, our most important source is the Bank for International Settlements (BIS) database of private non-financial sector credit and debt securities, as it provides a fairly detailed split of these series by the sources and users of finance. Appendix A contains a detailed description of the sources of all the variables that we use (Table A1). All the employed financial variables are expressed in relative terms to GDP and used after the logarithmic transformation (Table 1 describes the actual transformations of variables). This is first of all prompted by a better fit we obtained, and also suggested by the marginal impact of credit on growth rates estimated and presented by Cournède and Denk (2015) in their Figure 5 − using the logarithmic transformation we obtain the same shape of the marginal impact (see Fig. 1 in Section 4.1). Whenever the original BIS data is quarterly, we use the last quarter to align the frequency with the annual periodicity of other data. The BIS credit database contains directly the ratio of credit to nominal GDP series (with a split by credit to households and credit to non-financial corporations). For the outstanding debt securities (with a split into issued by non-financial corporations and financial corporations), we calculate these ratios to GDP using the BIS debt securities data and the GDP data from the World Bank's (WB) World Development Indicators (WDI) database. It should be pointed out that private bank credit data at the aggregate level (without splitting into household and firm credit) are also available from the WB Global Financial Development Database (GFDD). However, the GFDD credit series have a number of structural breaks, whereas the BIS credit data are adjusted for breaks. Fig. A1 in Appendix A presents several comparisons between data from the two sources, and those from the GFDD contain obvious structural breaks. This motivated us to use the BIS data in the econometric analysis. To represent the stock market financing of listed domestic companies, we use the market capitalization (in percentage of GDP) indicator from the WDI database. It should be pointed out that the usage of turnover ratio of domestic shares from the same database yields qualitatively similar results, but loses the significance, which is consistent with the analogous finding by Mishra and Narayan (2015). Another reason for preferring the market capitalization series is that its ratio to GDP is more natural and therefore aligns better with the other employed series that are also ratios to GDP. All the mentioned databases were downloaded in June 2016, and the respective extract of series is available upon request from the authors. The data period and number of observations to be used in further estimations varies depending on the particular question/specification at hand and the availability of data. The typical estimation period is from 1990 to 2014, whereas the number of actually available countries varies from 9 to 23, depending on the particular group of countries under investigation (OECD, EU, EMU1999) and data availability. The number of countries is always displayed in the tables containing the results. In addition to the discussed financial series, a set of usual control variables is included, comprising GDP per capita, enrolment in secondary education, government final consumption expenditure to GDP, trade openness to GDP, and inflation of consumer prices. These indicators come from the WB WDI database, and are also annual. The additional transformations of these original data are described in Table 1, and the specific choices ensure comparability with Arcand et al. (2015). Modelling strategy, employed model, and parameter estimation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our econometric research strategy is to start from simple log-linear specifications with only few financial variables, and then to introduce richer specifications with more detailed structure and/or non-linearity. Namely, we first consider the impact of bank credit, debt securities and stock market on growth, i.e., the impact of different sources of financing. Afterwards, we further decompose finances not only by sources, but also by fund users. Finally, we merge both specifications discussed above with non-linear components. While presenting the whole picture, this gradual approach thus reveals also the sensitivity of different specifications, without falling into potential problems connected with relatively low degrees of freedom and possible overfitting if only the richest specification were reported. The vector of explanatory variables xi,t can contain various linear and non-linear terms (logarithms, their squares, interactions, etc.) of economic series. The two main groups comprise the control variables and financial series that were summarized in Table 1. Let us turn to the parameter estimation. When the number of periods T grows to infinity, θh in Eq. (1) can be consistently estimated by e.g. the fixed effects estimator. However, when T is fixed, due to the problem of incidental parameters, consistent estimation of θh cannot be directly obtained from Eq. (1) and the instrumental variable-based estimators of Anderson and Hsiao (1981, AH hereafter) or the generalized method of moments (GMM) estimator of Arellano and Bond (1991) or Arellano and Bover (1995) and Blundell and Bond (1998) are usually applied. In larger samples, the GMM estimator is known to be more efficient when T is small and N is large, but it has large biases when T is relatively large. On the other hand, the AH estimator is consistent under both N and T asymptotics (see e.g. Phillips and Han, 2014). This last property is very convenient in our case, because we want to estimate the impact of financial deepening on economic growth in the sample of EMU countries, which has a very limited number of countries, thus forcing us to rely more on the increase in T rather than N. Because of this, and in order to increase the number of observations, we do not aggregate the initial data into e.g. 5 or 10 years periods (as in the baseline estimations of Arcand et al., 2015). That would not only substantially reduce the number of effective periods to a few, but also might induce pre-aggregation bias; while the removal of business cycle effects by such a simple aggregation is also questionable, because the length of business cycles might vary both in time and among different countries. Consequently, the AH instrumental variable estimator will be used hereafter. As a robustness check, we also run the Arellano and Bond (1991) difference GMM estimator. In all the cases, the robust inference is based on standard errors adjusted for clustering by countries. Caveats ~~~~~~~ The presented results should be considered with some caution due to several reasons. First, given our focus on a homogenous set of developed countries (most importantly: the EU and EMU1999), the sample size is quite limited, whereas the number of parameters is large due to the consideration of a detailed structure of financing. To tackle this, we use yearly data and not multi-year averages, as that would further shrink the number of observations. In addition, to increase the number of observations we consider also a larger group of countries (the OECD countries) and, given consistent results among various country groups, we are more confident in the findings established for the EU and the EMU1999. Note that a larger group can also cover potentially less homogenous countries where the impact of financial deepening and/or its structure therefore might also differ. Second, estimations that rely on the employed period (typically 1990–2014 or part of it) are informative about processes that took place during these years, but might be less indicative for other periods (either past or future). It is particularly true if there were substantial changes in the conditions, for example if there were important alterations of the financial structure or the inter-dependence between the structural components. In order to account for this, we try to control as much as possible for all relevant aspects and include all components of interest, which however limits the degrees of freedom. Consequently, there is a tradeoff between weak inferences versus potential biases due to omitted variables. Third, in order to avoid endogeneity stemming from simultaneous relationships, we use lagged explanatory variables in Eq. (1), i.e., it is always the future growth rates that are under prediction. However, this does not completely eliminate endogeneity, as expectations about future growth conditions can affect the choice of current levels of financial penetration, which may lead to a correlation between the financial series and the error term. It is however difficult to find the necessary (large number of) proper instruments needed in our case, due to the detailed analysis of the structure. Therefore, we present our results without taking into account this aspect. Fourth, the consideration of totals together with various levels of subcomponents (even though in a non-linear model) might lead to multicollinearity and thus weaken the statistical inference. Therefore, it is possible that some estimates would turn significant when adding more data, once they become available in the future. Fifth, the complete disaggregation of finances is not available: for example, credit to households or financial corporations are reported from all sectors and not from banks only, data coverage on private domestic or total outstanding debt securities varies across countries. Financing composition and non-linearity in bank credit ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 2 presents estimation results for the impact of composition with and without the non-linear term for bank credit. In general, there are always consecutive triplets of columns, using the same specification but for the different country groups (OECD, EU and EMU1999). In particular, columns (1)−(3) present a basic specification with financing split only by its source (bank credit, debt securities, and stock market). Columns (4)−(6) refine the analysis of columns (1)−(3) by further splitting bank and security based financing by its user. Columns (7)−(9) check how much the results in columns (1)−(3) change if one adds the non-linear component of bank credit. Finally, columns (10)−(12) augment the more detailed split of finances of columns (4)−(6) with the non-linear component of bank credit. Columns (1)−(3) of Table 2 reveal that even when using the log- linear approximation of the impact of finance on growth, the impact varies substantially (even in terms of its sign) for different types of financing: bank credit and debt security have a significantly negative impact on growth, whereas stock market financing tends to have a significantly positive influence. In terms of bank and stock market financing, we find that the latter is more beneficial for growth, at least in high-income economies. This is consistent with the evidence found in many previous papers (see e.g. an overview by Valickova et al., 2015). In short, it is not all types of financing that affect growth negatively. Our estimates are also economically meaningful. An increase of the bank credit to GDP ratio from 90% to 100% (a log difference of 0.105) implies a 0.17–0.12 percentage point reduction of average growth rates.7 The impact of debt securities and equities is proportionally smaller. The results also reveal that the impact of the different types of sources is not homogenous. In particular, the absolute values of the coefficients of bank credit and stock market capitalization are significantly different; therefore, the data does not support the use of their ratio. Next, the finding that outstanding debt securities have a negative, while stock market capitalization has a positive effect (see e.g. Kaserer and Rapp, 2014, for a similar finding for the EU countries) reveals that merging/pooling all sources of market-based financing (as e.g. in Langfield and Pagano, 2016) is not supported. Consequently, the equal promotion of different types of market-based financing can be suboptimal from an economic policy point of view. Turning to the impact of an even more refined financing structure (both by sources and users of finance) presented in columns (4)−(6), we confirm earlier findings that bank credit to households is a drag on economic growth, whereas bank credit to firms tends to promote economic growth rates significantly. A similar though somewhat weaker conclusion can be drawn about the importance of the structure of outstanding debt securities. Namely, the coefficient of debt securities issued by financial corporations tends to be significantly negative, whereas that of debt securities issued by non- financial corporations is insignificant. Hence, economic growth would have been higher during the analyzed period if outstanding debt securities were issued more by non- financial corporations than by financial corporations. Nevertheless, the coefficient of debt securities of non-financial corporations is still negative. Although it is insignificant, this negative sign contrasts sharply with the positive coefficient of stock market capitalization, which also tends to be significant. As columns (7)−(9) show, the same conclusions are robust to the introduction of the non-linear impact of bank credit (CREDIT2). The only difference is that the linear term is positive for bank credit, while the quadratic term is negative. Thus, the non-linear impact of bank credit remains significant (at least at the 10% level) after taking into account the split by the source of financing. The finding that the linear term is positive while the quadratic term is negative implies that there is a turning point in the impact of bank credit on growth (see the end of this subsection for a detailed analysis of this). It should be pointed out that CREDIT and CREDIT2 are highly correlated by construction, which is partly responsible for the moderate significance of CREDIT and CREDIT2 observed in the OECD and the EU. Looking the other way round, i.e. at the stability of results about the role of financial structure to the inclusion of the non-linear term, a few changes emerge. First, the findings about the relative benefits of promoting stock markets become even stronger as the coefficients of stock market capitalization become larger and more significant. Next, the differentiation between the influences of different types of debt securities becomes more blurred. Similarly, the positive impact of bank credit to non-financial corporations becomes significant only in the EMU1999 case (although there it becomes more significant than without the non-linear term). Nevertheless, the relative inferiority of credit to households remains strongly valid. Summarizing the main findings of this analysis, we find that the impact of finance on economic growth differs substantially among the different types, and these findings are robust to the presence or absence of the non-linear bank credit term. During the analyzed period, bank credit was on average a drag on economic growth rates, but the bulk of this stems from the negative impact of household credit. Nevertheless, the non-linear impact of bank credit is robust to controlling for the main structural composition of financing, both in terms of its source and its user. Therefore, a part of reduced growth can also come from the non-linear impact of ‘too much credit’, given that most countries in our sample have already reached credit levels higher than the turning point (peak of maximum contribution of credit to growth, to be characterized shortly). Higher stock market capitalization seems to be robustly connected with higher economic growth, whereas larger outstanding debt securities to GDP have a negative impact (and significantly so for financial corporations, when the non-linear credit term is absent). Although these conclusions might be specific to the period under investigation, they are quite robust despite substantial changes in model specifications. Finally, let us discuss the estimated turning points of the non-linear impact of bank credit on growth rates. Fig. 1 plots the marginal impact of bank credit on growth, with the turning point estimate identified where the marginal impact equals zero. First, it can be seen that the estimated turning point is smaller when finance is split only in terms of sources. In this case, it is below 50% of GDP and varies from 37% to 46% depending on a group of countries. Furthermore, considering the 95% confidence bounds, the marginal impact of financing here is never found to be significantly positive. On the other hand, the positive contribution becomes significant when a more detailed split of financing is employed (also by the user of finance). In this case, the turning point also increases and ranges from 61% to 72% in the different country groups. It is interesting to note that these point estimates (in particular, 62% of GDP for the OECD) compare well with that obtained by Cournède and Denk (2015) for the OECD countries, using a longer intermediate credit series (their estimated turning point is about 60% of GDP). However, these point estimates are in general lower than those established by Arcand et al. (2015), using their global sample of countries. Nevertheless, the mentioned difference is less evident once looking at the confidence bands: for some specifications provided in Arcand et al. (2015), the difference is statistically significant, whereas for others it is not. Financing structure and other non-linearity questions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In this subsection, we explore whether the non-linearity of the effect of finance on growth is sufficiently captured by the non-linear term of bank credit alone. Maybe the total amount of financing from all the different sources is more relevant than bank credit alone in generating the non-linearity, conditionally either only on the sources of financing, or the sources and users of financing. Alternatively, maybe household credit is solely responsible for the non-linear impact of bank credit,8 thus, after taking it into account, the non-linearity of total bank credit vanishes. In order to answer these questions, we investigate the statistical significance of the respective non-linear terms. Table 3 presents the corresponding empirical findings. Columns (1)−(3) include both the non-linear term of bank credit and that of the total financing, conditioning on the sources of financing. Columns (4)−(6) also condition on the users of finance. Finally, columns (7)−(9) compare the relative significance of the non-linear terms of total bank credit and of household credit only. Comparing the significance of the linear and non- linear terms of bank credit (CREDIT, CREDIT2) and total financing (TOTAL, TOTAL2) in columns (1)−(6) of Table 3, one can see that the impact of bank credit is consistently more significant than that of the total financing. Although the difference is moderate in columns (1)−(3), where we control only for the sources of finance, there is little doubt about the substantial difference in significance when a detailed financing structure is taken into account (columns (4)−(6)). Therefore, we can infer that bank credit seems to dominate in the hump-shaped finance-growth relationship. One can draw similar conclusions from columns (7)−(9), regarding the relative significance of the non-linearity of household credit and (total) bank credit. Bank credit retains uniformly not only the sign of both its linear and non-linear terms, but also the significance, whereas the non- linearity connected with household credit does not only change signs irregularly, but also becomes insignificant in the OECD and EU samples. In the EMU1999 case, the terms of household credit are significant, but it is more likely to occur due to the small number of observations, potentially coupled with multicollinearity of bank credit and household credit terms (and their squares). We therefore can infer that, even after controlling for the structure of financing in a detailed manner, the hump-shaped, non-linear impact of finance on growth seems to be most strongly connected with (total) bank credit. Robustness checks and extensions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In this subsection, we summarize the implications of some robustness checks and further extensions. We look at the impact of varying the length of future horizons (h), excluding outlier observations, including dummy-interaction variables for the latest after-crisis period, reducing the number of variables (dropping period effects, dropping controls, leaving only the most significant principal component of controls), using ratios to represent the composition of financing instead of an unconstrained estimation, additional modeling of dynamics (by including the changes of explanatory variables or including autoregressive terms of the dependent variable), and using a different estimator (the Arellano and Bond, 1991, difference GMM estimator). Finally, we explore the inclusion of some additional indicators: an index of patenting rights, and a dummy for accelerating real housing prices. Appendix B describes the implementation details. In order to save space, we mostly concentrate on the sensitivity analysis and extensions of the main results provided in Table 2: either the whole table whenever possible, or a part of it, namely, the specification which has the most detailed split of financing composition (columns 10–12). Due to the same reason, all tables associated with the empirical estimation results are delegated to Appendix B. The results of the performed robustness analysis can be summarized as follows. In general, the previously discussed main findings are quite robust to the considered deviations from the baseline specifications considered in Table 2. The least robust one is about the impact of the composition of outstanding debt securities: although the negative sign of debt securities issued by both the financial and non-financial corporations is dominant, the ranking of its subcomponents becomes less obvious in many of the performed investigations. Some additional interesting aspects are worth singling out. First, the negative impact of household financing seems to emerge more over longer periods, and is much smaller in shorter horizons, as revealed both by Tables B1 and B8. Second, it should be pointed out that these robustness checks reveal that our findings are not driven by the recent financial crisis period: they remain present while taking into account the potential crisis period effects as in Table B3. The results are also not driven by a few outlying countries or observations: in addition to that they are quite homogenous among the OECD, EU, and EMU groups of countries, the results are robust to cutting outliers and removing up to nearly a third of observations (see Table B2). Our two extensions also yield important insights. First, in line with Chu et al. (2016), Table B11 shows that patenting is connected with the channel through which an increase in household financing (relative to firm financing) influences economic growth rates. Second, Table B13 illustrates that the positive impact of stock market financing seems to be mostly observed during periods of accelerating real housing prices, after which economic growth is significantly lower, but less so in countries that relied more on capital markets during the associated housing market spur. The analogous impact of debt securities was not observed and even had a negative sign, which can be connected also with the bank strategies to finance housing loans by issuing debt securities.","This paper contributed to the analysis of the impact of finances on economic growth by incorporating the structure of financing and allowing for the non-linearity of the impact of finances, in homogeneous groups of high-income countries. Our results reveal that different types of finances (both by their sources and their recipients) have very different effects on economic growth, and the significance of the non-linear impact of bank credit is robust to controlling for a fairly detailed composition of private finances, and also for an inverted U-shape relationship between patents and innovation (and consequently, economic growth rates). Furthermore, results are very similar in all the three high-income groups of countries considered (member states from the OECD, EU, and EMU1999). Besides its robustness, we find the following additional features of this non- linearity. The non-linear impact of total bank credit is more pronounced than that of either only household credit or the joint sum of bank credit, debt securities, and stock market financing. The estimated turning point/threshold of the identified non-linear relationship is smaller than that established e.g. in Arcand et al. (2015) using a global panel, while it is in line with that estimated for the OECD countries by Cournède and Denk (2015). Therefore, a large bank credit penetration relative to GDP (especially with heavy financing of households) might be more harmful to economic growth in high-income countries than thought previously. At the same time, due to the dominance of bank-biased financing in the EU, even a simple reduction of bank credit relative to GDP, e.g., by the means of removing the debt-favoring pattern in the taxation systems, could result in improved economic growth rates in a number of EU countries, especially if stock markets, having a positive impact on economic growth, were promoted instead. It should be noted that the positive impact of stock markets on growth is sometimes perceived as being driven by the differences in the Anglo-Saxon and continental models of economic and financial systems, but our findings reveal that this feature persists while considering only the EU or EMU countries. We also find and/or confirm many important aspects of the role of financing composition, even after controlling for the non-linearity discussed above. First, the impact of bank credit to households and non-financial corporations qualitatively differ: in our sample, the former had a strongly negative, whereas the latter tended to have a positive impact on economic growth. Consequently, if a reduction of bank credit were beneficial for a particular economy in general, the strongest promotion to growth could be achieved by shrinking household credit. This established empirical finding seems to support the hypothesis that, in the long run, household credit diverts funds of limited supply from firms that could generate longer-lasting positive development. One potential mechanism is the non-trivial interaction of patents with finance, as put forward by Chu et al. (2016). Namely, whenever funds are diverted from firms towards households, financially constrained firms fail to develop their R&D projects further, slowing down the speed of innovation and eventually economic growth. We find such interactions to be empirically significant. The diversion of funds from firms to households can become especially acute during housing market booms, periods that facilitate the expansion of credit to households by creating larger values of collateral acceptable to banks and larger returns in this market. We indeed find that, during periods of significantly positive real housing inflation, growth was further reduced besides what has already been captured by the amounts of credit to households directly. Thus, either housing credit has a further negative impact on long-term growth relative to total household credit (e.g., it may create a drag on households’ willingness to work productively), or the actually realized amounts of household credit do not reveal its whole negative influence (e.g., banks shrank firm financing more by foreseeing the need of additional household borrowing in the future). At the same time, the larger share of stock markets during the housing booms was curbing the negative effect on growth, which might simply indicate that the bank potential to reallocate the means towards housing was curbed, but at the same time it could point to the importance of the development of a balanced financial system with sufficient competition between different modes of finance. Next, the growth impact of stock market and debt security financing are qualitatively different: stock market financing has a positive, whereas debt securities tend to have a negative influence on growth. Looking from both the methodological and policy perspectives, this would suggest that the use of financing aggregates and the equal promotion of all types of market-based modes of financing might be just as misleading as cutting all types of bank credit. Although statistically less clear-cut, we have found some evidence that shifting currently outstanding debt securities from financial corporations towards the non-financial ones could be beneficial for growth. This can be due to several factors at play. First, a substantial part of debt securities issued by financial institutions is connected to the financing of housing, which we find to have a negative impact on growth. Furthermore, international financial markets are highly integrated, and financial institutions issuing debt securities can outsource domestic savings from high-income economies to other countries easily, thus reducing the local funding of investments. On the other hand, given the increased total globalization of corporate activities, it can be a potential explanation also for the negative sign (though smaller absolute value) of the impact of non-financial corporations. Finally, from the policy perspective, our results point to several alternatives connected with the financial deepness and its structure that would promote economic growth. Regarding the banking sector, growth could be enhanced by directing more credit towards non-financial corporations and by reducing the bank credit to GDP levels in a number of European countries (especially, from the EMU). The reduction of household credit, which simultaneously diminishes the total amount of credit and favorably changes its composition, can have the largest economic impact. However, the effect of a reduction of the total amount of bank credit also depends nontrivially on the initial conditions of a particular economy (namely, the actual distance from the peak impact of credit, the level of penetration of all modes of finance, etc.). Therefore, for economies that are close to the turning point of the non-linear impact, a balanced compositional shift towards firm financing without affecting the total amount of credit might be best suited. The further development of market-based financing seems to be mostly beneficial through the fostering of stock markets. The above findings are in line with the European Union's ongoing strategy to generate a greater choice of funding through a diversified financial system, while making the financial system more resilient (through the capital markets union)."],["This paper estimates the magnitudes of government spending and tax multipliers within a regime-switching framework for the U.S. economy during the period 1949:1-2006:4. Our results show that the magnitudes of spending multipliers are larger during periods of low economic activity, while the magnitudes of tax multipliers are larger during periods of high economic activity. We also show that the magnitudes of fiscal multipliers got smaller for episodes of low growth, while they got larger for episodes of high growth in the post 1980 period. Analyzing the effects of government spending and taxes on consumption and investment spending indicates that the magnitude of the effects of fiscal shocks on consumption and investment is very small. --------------------------------------------------------------------------------","The role of fiscal policy in stabilizing business cycles came under scrutiny by researchers and policymakers about three decades ago. As argued by Beetsma and Guilidori (2011), expansionary fiscal policies implemented in response to oil price shocks did not provide the desired results and therefore raised concerns regarding the efficiency of fiscal policy during business cycles. Moreover, fiscal consolidations in Europe during the 1980s, contrary to Keynesian wisdom, led to an increase in output in the short-run and in the long-run and therefore led economists and policymakers to question the established theories regarding fiscal policy. Recent studies, including Alesina et al. (2002) explain this puzzling result with the fact that, certain fiscal shocks, namely shocks to government wages and salaries, can have non-Keynesian effects. They show that negative shocks to government wages and salaries result in an increase in economic activity both in the short-run and in the long-run by decreasing labor demand and wages, and therefore increasing business profits and investment. With the global financial crisis of 2008 turning into a global recession, there has been a revival of interest in the effects of fiscal policy on major macroeconomic variables. Especially in the U.S., with President Obama’s fiscal stimulus package, there is a heated discussion regarding the effectiveness of fiscal policy as an economic stimulus tool. In Europe, similarly, fiscal consolidations in many countries pushed economies deeper into recession, quite differently from what we observed in the 1980s. This particular observation certainly suggests that the magnitude- and even the sign- of fiscal multipliers might change during the business cycle. There are different approaches in estimating tax and spending multipliers, but the two most common approaches employ structural macroeconometric models or vector autoregressive (VAR) models. Among these two approaches, the VAR models occupy a more prominent role in the recent literature. The studies using VARs identify fiscal shocks either by employing the structural VAR approach or the narrative approach. The structural VAR approach uses either economic theory or institutional information to identify the variance/covariance matrix, and therefore the fiscal innovations (Blanchard and Perotti, 2002; Perotti, 2002). The multipliers estimated with this approach are close to (in most cases less than) unity. Perotti (2002) also argues that the tax multipliers tend to be negative but small, despite some evidence on positive tax multipliers. Finally, he argues that the U.S. is an outlier in many dimensions, so the responses to fiscal shocks estimated on U.S. data are often not representative of the average OECD country. Most VAR studies reach the conclusion that the post-1980 fiscal multipliers are smaller (Perotti, 2002; Favero and Giavazzi, 2009). This particular result is generally interpreted as fiscal policy becoming more ineffective over the years – most probably due to increased labor and capital mobility. The narrative approach, identifies exogenous fiscal shocks by a narrative based dummy (Ramey and Shapiro, 1998) or the defense news measure (Ramey, 2011) or the exogenous tax measure (Romer and Romer, 2010a). While Ramey and Shapiro, 1998) use large exogenous increases in defense spending, like the Vietnam War, the Korean War and the Carter-Reagan military build-up to identify shocks to fiscal policy and Ramey (2011) constructs a new defense news variable which measures the present discounted value of expected change in military spending, Romer and Romer (2010b) use information from the official U.S. budget documents to classify exogenous tax changes. Ramey (2011) estimates the spending multipliers to be between 0.6 and 1.2, while Romer and Romer (2010a) find that an exogenous tax increase of 1% of GDP lowers real GDP by almost 3%. Among the more recent studies that do not use VARs, Barro and Redlick (2011) estimate defense spending multipliers with two-stage least squares, using annual data for different samples where the estimated multipliers lie between 0.6 and 0.7. The approaches mentioned above, with the exception of Auerbach and Gorodnichenko (2012), employ linear models in estimating the tax and spending multipliers. A common characteristic of these studies is that the magnitude of the multipliers does not vary over the business cycle. Auerbach and Gorodnichenko (2012) employ a regime switching VAR where transitions across recessions and expansions are smooth. By imposing the restriction that the U.S. economy is in recession 20 % of the time, they estimate that the total spending multiplier is 0.57 during expansions and 2.45 during recessions, while the defense spending multiplier is 0.8 during expansions and 3.56 during recessions. In this paper, we investigate empirically whether fiscal multipliers are quantitatively different in magnitude during “good times” and “bad times”. To do so, we use a multiple regime framework first suggested by Hamilton (1989). We contribute to the literature by estimating a non-linear model within a Markov-switching framework and obtaining government spending and tax multipliers during periods of low and high levels of economic activity. Our paper differs from Auerbach and Gorodnichenko (2012) in two respects. First, the Markov switching model that we employ has different properties from the STVAR model used by Auerbach and Gorodnichenko (2012). The model employed in this paper provides additional information as it estimates the transition probabilities (the probability of staying in each of the two regimes, low economic activity and high economic activity). Second, government spending multipliers are identified from variations in the defense news variable constructed by Ramey (2011) and the tax multipliers are identified from the exogenous tax variable constructed by Romer and Romer (2010a). Ramey (2011) shows that defense spending accounts for almost all of the volatility of government spending, but also argues that shocks to government spending or defense spending can be anticipated ahead of actual spending. This has important implications because anticipated future changes in government spending can affect current economic activity. She shows that the standard VAR shocks do not reflect news about defense spending accurately and that the Ramey–Shapiro war dates Granger-cause the VAR shocks. Ramey (2011) also acknowledges that the simple dummy variable approach does not exploit the potential quantitative information available regarding the news about military spending and for this purpose constructs a new measure of defense news variable, which reports the anticipated changes in defense spending. We use this measure to identify shocks to government spending and to calculate the spending multipliers. One major obstacle in calculating tax multipliers is endogeneity. As GDP increases, we observe an increase in tax revenues and vice versa. This makes the calculation of tax multipliers very difficult. Romer and Romer (2010b) argue that most changes in revenues are endogenous responses to non-policy developments. They analyze federal tax actions from 1945 to 2007 and identify four categories. Of these four categories, spending-driven and countercyclical tax changes are defined as endogenous tax changes, while deficit-driven long-run tax changes are categorized as exogenous tax changes. We use the exogenous tax changes in estimating the tax multipliers. The non- linear model employed in this paper separates periods of high and low states of the world for the endogenous variable (the change in real GDP per capita scaled by the real GDP per capita of the previous period, which can also be interpreted as per capita growth), and therefore allows us to estimate separate fiscal multipliers for periods of low growth, and periods of high growth. We find that the spending multiplier is 2.91 for periods of low growth and 0.13 for periods of high growth, while the tax multiplier is −0.19 for periods of low growth and −0.66 for periods of high growth. Our results show that the magnitudes of the spending multipliers are larger during episodes of low growth, while the magnitudes of tax multipliers are larger during episodes of high growth – a result that emphasizes the importance of non-linearities for fiscal multipliers. Moreover, the non-linear framework used in this study provides larger multipliers for periods of high growth and smaller multipliers for periods of low growth during the post-1980 era when compared to the whole sample period, which indicates that previous findings about the post-1980 multipliers might be biased since they do not differentiate between different states of the economy. As a further analysis, we investigate how changes in government spending and taxes affect investment and consumption spending. Our results indicate that consumption spending rises during episodes of low economic activity and falls during times of high economic activity, but only the rise in consumption during times of low economic activity is statistically significant at the 5% level. Investment spending falls in response to an increase in government spending during episodes of both low and high economic activity. However, only the decline in investment during periods of low economic activity is marginally significant. We find that both consumption and investment spending fall in response to an increase in exogenous taxes during both episodes of low and high economic activity, but none of the parameter estimates is statistically significant. The remainder of the paper is organized as follows: Section 2 discusses the methodology employed in the paper. Section 3 presents the empirical analysis and results. Section 4 concludes. Data ~~~~ We added squared government bond spread as an indicator of monetary credit conditions, following Barro and Redlick (2011). Since the spread is endogenous with respect to GDP growth, its lagged value is used rather than its current value. As indicated by Barro and Redlick (2011), if the yield spread is thought as analogous to a distorting tax rate, then the square of the spread approximates the deadweight loss from the distortion. Nominal GDP and the GDP deflator are from the Bureau of Economic Analysis and interest rates are from the Board of Governors of the Federal Reserve System. Nominal present discounted value of expected change in defense spending and total population, including armed forces overseas are from Ramey (2011) and the change in nominal exogenous tax liabilities are from Romer and Romer (2010b). Summary statistics for the variables considered are reported in Table 1. Our dataset covers 1949:1–2006:4 (quarterly) for a total of 232 observations. Results ~~~~~~~ The null hypothesis of linearity against the alternative of Markov regime switching cannot be tested directly using the standard likelihood ratio (LR) test. We properly test for multiple equilibria (more than one regime) against linearity using the Hansen’s standardized likelihood ratio test (1992, 1996). The value of the standardized likelihood ratio statistics and related P-values (Table 1) under the null hypothesis (see Hansen, 1992, 1996) for details) provides strong evidence in favor of a two – state Markov mean–variance regime-switching specification.2 Maximum likelihood (ML) estimates of the model described above are reported in Table 2. The model appears to be well identified, parameters are significant, and the standardized residuals exhibit no signs of linear or nonlinear dependence (Ljung–Box statistics for dependency in the first moment and for heteroskedasticity). The periods of high and low economic growth seem to be accurately identified by the filter probabilities, which clearly separates the two regimes. Our findings are in sharp contrast with the empirical findings of the previous VAR studies that use linear VAR models. These studies do not make a distinction between periods of low and high economic activity and estimate government spending multipliers mostly around 1. An exception to these VAR studies is Auerbach and Gorodnichenko (2012), who employ a regime switching VAR and find a government spending multiplier of 0.57 during expansions and 2.45 during recessions. Even though there are some methodological differences between our single-equation model employing a Markov switching framework and the regime switching VAR model of Auerbach and Gorodnichenko (2012) and the way we identify fiscal shocks, both approaches have similar findings that are in sharp contrast with the previous literature. It should be noted that the spending multipliers that we have estimated are actually the multipliers associated with anticipated changes in military spending based on news. As argued by Barro and Redlick (2011), multipliers associated with non-defense purchases would be more relevant to evaluate fiscal stimulus packages, but they acknowledge that these multipliers are hard to estimate since there is a great deal of the endogeneity between the movements in non-defense spending and real GDP and therefore they estimate multipliers for defense spending. They argue that the defense spending multiplier would provide an upper bound for the non-defense multiplier. Our multipliers should be interpreted similarly. They represent an upper bound for the non-defense multiplier during periods of high and low growth. Extensions ~~~~~~~~~~ Another issue that we investigate in this paper is related to how consumption and investment spending are affected by changes in defense news measure of Ramey (2011) and the exogenous tax measure of Romer and Romer (2010b) during periods of low and high growth. The investigation of this issue is important in evaluating the predictions of the real business cycle (RBC) theory and the Keynesian analysis. According to the RBC theory, increase in government spending financed by future lump-sum taxes, creates a negative wealth effect, which decreases consumption and increases employment. Increase in employment raises the return to capital and increases investment. According to the Keynesian analysis, an increase in government spending financed by future lump-sum taxes increases disposable income, which leads to an increase in consumption. On the other hand, increase in government spending also raises the real interest rate, which leads to a fall in consumption and investment spending. The empirical evidence on this issue is mixed. Fatas and Mihov (2001), Blanchard and Perotti (2002), Perotti (2002), Caldara and Kamps (2008), and Mountford and Uhlig (2009) find that consumption increases in response to government spending, while Ramey and Shapiro (1998), Edelberg et al. (1999), Burnside et al. (2004), Cavallo (2005) and Ramey (2011) find that government spending lowers consumption. The common thread among these studies is that they do not take into account the possibility that the response of consumption to government spending may differ during expansions and recessions. An exception to this is Tagkalakis (2008) who investigates the effects of fiscal policy changes on private consumption in recessions and expansions in the presence of binding liquidity constraints on households. Using an unbalanced yearly panel data set (1970–2002) of nineteen OECD countries, Tagkalakis (2008) finds that fiscal policy is more effective in boosting private consumption in recessions than in expansions. Our methodology differs from Tagkalakis (2008) in two respects. First, we identify “bad times” and “good times” from a regime switching model that allows for shifts in the mean whereas Tagkalakis (2008) identifies “good” and “bad” times by extracting the cyclical component of real GDP by applying the Hodrick–Prescott (HP) filter. Second, we use the instruments used and suggested by Ramey (2011) and Romer and Romer (2010a) as measures of fiscal policy to avoid problems related to endogeneity. We find that both consumption and investment spending fall in response to an increase in exogenous taxes during both episodes of low and high economic activity, but none of the parameter estimates is statistically significant at the 5% or 10% level. We have some similarities and differences with Tagkalakis (2008) in relations to empirical findings on consumption. We both find that consumption increases in response to an increase in government spending during “bad times”, which is statistically significant. However, our results indicate that the impact of government spending on consumption is very small in magnitude. We find that consumption decreases during expansions in response to an increase in government spending, though this is not statistically significant. Tagkalakis (2008) finds a statistically significant increase in consumption during expansions, though not as big as in recessions. We find a decrease in consumption in response to an increase in taxes, while our findings are not statistically significant. Tagkalakis (2008) has similar, but statistically significant results. We interpret Tagkalakis’s findings as strong evidence and ours as weak evidence for Keynesian models. The differences could be due to different methodologies and different data sets employed.","By identifying fiscal policy shocks, using the narrative approach, we estimate the magnitude of fiscal multipliers within a non-linear framework. The empirical results show that the magnitudes of the spending multipliers are larger during times of low growth, while the magnitudes of tax multipliers are larger during times of high growth. Our results imply that there is a role for fiscal policy as a stabilization tool by using the “right instrument” at the “right time”. Another contribution of our paper is, contrary to the previous literature, to show that multipliers during periods of low growth get smaller, and multipliers during times of high growth get larger in the post-1980 era relative to the whole sample period. This particular result implies that the comparison of fiscal multipliers between periods in linear VAR studies might be biased. Analyzing the effects of government spending and taxes on consumption and investment spending indicates that the effects are very small. Further avenues for research may include further disaggregation of fiscal shocks to find out exactly which budget items can be used to stabilize the economy during recessions (or expansions). As Perotti (2002) contends that the U.S. is an outlier in terms of response to fiscal policy actions, there is also some benefit applying our framework to other countries, given data availability."],["We utilize longitudinal data on nearly 1800 children in Vietnam to study the predictive power of alternative measures of early childhood undernutrition for outcomes at age eight years: weight-for-age (WAZ8), height-for-age (HAZ8), and education (reading, math and receptive vocabulary). We apply two-stage procedures to derive unpredicted weight gain and height growth in the first year of life. Our estimates show that a standard deviation (SD) increase in birth weight is associated with an increase of 0.14 (standard error [SE]: 0.03) in WAZ8 and 0.12 (SE: 0.02) in HAZ8. These are significantly lower than the corresponding figures for a SD increase in unpredicted weight gain: 0.51 (SE: 0.02) and 0.33 (SE: 0.02). The heterogeneity of the predictive power of early childhood nutrition indicators for mid-childhood outcomes reflects both life-cycle considerations (prenatal versus postnatal) and the choice of anthropometric measure (height versus weight). Even though all the nutritional indicators that involve postnatal nutritional status are important predictors for all the mid-childhood outcomes, there are some important differences between the indicators on weight and height. The magnitude of associations with the outcomes is one aspect of the heterogeneity. More importantly there is a component of height-for-age z-score (at age 12 months) that adds predictive power for all the mid-childhood outcomes beyond that of birth weight and weight gain in the first year of life. --------------------------------------------------------------------------------","Studies on the importance of early-life anthropometry for later human capital development recently have become prominent. Behrman and Rosenzweig (2004), Black et al. (2007), Victora et al. (2008), Rosenzweig and Zhang (2013), and Figlio et al. (2014) find birth weight to have significant associations with long-run adult health, education and earnings. Gupta et al. (2011) and Krishna et al. (2016) are the most relevant previous studies for the birth weight-related contents in our study.1 Gupta et al. (2011) use data from the Danish Longitudinal Survey of Children (DALSC), which followed children born in 1995 with surveys in 1996, 1999, 2003 and 2007. For mid-childhood outcomes, their findings imply that the associations of low birth weight (2.5 kg or lower) with weight and height are statistically significant, but not the associations with behavioral outcomes. The data used in Krishna et al. (2016) are from Young Lives, which is described in the next section. They find that prenatal conditions, reflected in birth weight, are more strongly associated with height trajectories than postnatal factors; they do not consider weight and educational outcomes. A number of child health and nutrition researchers have focused on the concept of a critical window for investing in early childhood nutrition during the first 1000 days after conception (Martorell et al., 1994; Victora et al., 2008, 2010; Prentice et al., 2013; Lundeen et al., 2014). Stunting at age 2–3 years, which is indicated by deficits of two standard deviations or more below the median height-for-age for a well-nourished reference population, has become an important policy concern (Engle et al., 2007, 2011; Grantham-McGregor et al., 2007; UNICEF, 2013; Richter et al., 2016). Inadequate prenatal nutrition results in low birth weight and inadequate postnatal nutrition results in low weight and height gains in early childhood. There have been few studies that focus on the impacts of infant weight gain and height growth. Even fewer studies investigate the relative importance for predicting later development of weight gain versus height growth. Huang et al. (2013) reviews the limited literature on the predictive power of infant weight gain and height growth and reports substantial variety in the findings. Li et al. (2004) find that height growth during the postnatal period (birth to age 2 years) is the only variable predictive of Guatemalan women’s educational achievement and weight gain is not statistically significant. Corbett et al. (2007) find in UK data that “postnatal weight gain in the first 2 years of life is at most weakly related to cognitive and education attainment at age 10. In contrast, birth weight is clearly associated with cognitive and educational attainment at age 10…” (2007: 62). There is, however, an issue that may affect Corbett et al’s results: factors such as birth order, maternal education and family environment are not included in their analysis. More recently, Huang et al. (2013) evaluated the relative associations of birth weight and postnatal growth (weight gain, height growth, or head circumference growth) with cognition and behavioral development in over 8000 Chinese children. They find that, for full-term children, both birth weight and postnatal growth are associated with child’s IQ at age 4–7 years, but the sizes of the associations are small. While Huang et al. (2013) used up to seven years for their postnatal period, we consider the postnatal period of the first year, which is within the usually-emphasized critical window. We study the nexus of the early childhood undernutrition with anthropometric and educational outcomes in mid childhood. There are many studies that use height-for-age z-scores to predict longer-run outcomes and many that use birth weight (e.g., Crookston et al., 2013; Victora et al., 2008). While height-for-age is often interpreted to measure the nutritional status over the whole period from conception to measurement, weight-for-age is expected to have an advantage in capturing the effect of recent health shocks before the survey, such as diarrhea in the few months before the survey. In the same vein, weight gain after birth might be relevant in addition to birth weight and weight-for-age. The timing of the impact is important because of the debate on the “critical period” during which brain development is most sensitive to poor nutrition. Dobbing (1976), for example, argues that the period from birth to six months is the most critical. For other authors, such as Doyle et al. (2009), the prenatal period is more important. To separate the pre- versus postnatal influences of nutrition status, we apply two-stage procedures in which we derive indicators of unpredicted nutritional changes. We define the unpredicted weight gain (height growth) in the first year of life as the component of weight-for-age (height-for- age) z-score at age 12 months that is unpredicted based on what is known at birth, including birth weight. Further, we define the conditionally unpredicted height-for-age as the component of the unpredicted height growth that is uncorrelated with weight-for-age at age one year. The conditionally unpredicted height-for-age is useful for comparison of predictive powers of the indicators on height growth versus that of weight gain in the first year of life. The sample of children we work with in this study differs from those in the aforementioned studies on birth weight in a number of dimensions. Our data contain no twins, so the distribution of birth weight differs from that in the papers that use data on twins since the distribution of birth weights for twins is to the left of the distribution of birth weights for singletons.2 In part because our data do not have twins, less than five percent of children in our sample have birth weight under 2.5 kg (the standard cutoff for low birth weight), even though we use a semi-purposeful pro-poor sample from Vietnam. Also the age patterns of undernourishment are very different than reported for other contexts in previous studies. In the first decade of the 21st century, for Vietnamese under-5 years old moderate and severe stunting rates were as high as those for West and Central Africa, while the Vietnamese percentage of low birth weights was the same as for high-income countries (UNICEF, 2009). We examine different combinations of variables for birth weight, weight-for-age, height-for-age, unpredicted weight gain (height growth) and conditionally unpredicted weight-for-age (height-for-age), together with a set of controls, to estimate their associations with mid-childhood outcomes. We find that a standard deviation higher unpredicted weight gain (height growth) in the first year of life generally is more associated with positive outcomes in mid-childhood than are the adverse outcomes associated with a standard deviation lower birth weight. This suggests that for most of the children, the adverse outcomes related to being born low weight may be partially or totally avoided by weight gain (height growth) in the first postnatal year. In addition to the question of prenatal versus postnatal timing, we also investigate another aspect of heterogeneity of the relationships of the early childhood nutrition indicators and mid-childhood outcomes: the difference between height and weight. We find that height-for-age z-score at age one year contains a component that adds explanatory power for the variation of all the mid-childhood outcomes, beyond that of birth weight and weight gain in the first year of life. Young lives data for Vietnam ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The sample is part of Young Lives, which is an international comparative study of child poverty. Since 2002, Young Lives has been following ∼12,000 children in Ethiopia, India, Peru and Vietnam. The Young Lives sample consists of two cohorts: ∼8000 children born in 2001–2 and ∼4000 children born in 1994–5. The Vietnamese Younger Cohort sample, which we use in this study, consists of 20 commune-based clusters in Northern Uplands, Red River Delta, Central Coast, and Mekong Delta. The subsamples for each of these regions contain four rural clusters, except that for the Central Coast, which consists of four urban clusters in the city of Da Nang and four rural ones in the province of Phu Yen. The children in this study were on average 11.6 months old in Round 1, with the youngest being 5 months old and the oldest being 18 months old. Consistent with the design of the Young Lives study, we limit consideration to children in the age range from 6 to 17.9 months and therefore exclude 21 observations.3 Furthermore, following the same method as in Lundeen et al. (2014) and Schott et al. (2013), we exclude 45 observations with anthropometric measures in Rounds 1–3 either too low or too high or with unusually large changes (in height or weight) from one survey round to another.4 We also drop a further 65 observations because of missing data on either weight or height in Rounds 1–3. There are missing values on other variables discussed in the following section. Variables ~~~~~~~~~ As presented in Table 1, the characteristics of children include: sex, birth order, birth weight, year of birth, weight at one year, and length at one year. To allow nonlinearity in the effect of birth weight at the high end, we use a dummy variable that equals 1 for birth weight >3.5 kg.5 The children under study are divided into two groups by birth year. A dummy variable is defined for the children born in 2002. By the Education Law of Vietnam, the children in this group started school one year after the ones born in 2001, regardless of birth month. The Law was not strictly followed, however.6 Data on birth weight were recorded in birth documents by health clinic staff at the time of birth. Other anthropometric data were collected in the Young Lives surveys. Birth weight has missing observations for 11% of the children. We impute values for these missing observations on the basis of an OLS regression for the 89% of the sample that has birth weight data and the main caregiver’s recollection (at age 1 year) of the child’s size at birth and other relevant factors (Appendix A). We normalize birth weight in standard deviation (SD) units for our estimates to facilitate comparisons of coefficient magnitudes across the nutritional indicators by having them all in SD units. Supine length (Round 1) and standing height (Rounds 2&3) were measured with length/height boards using standardized WHO methodology (WHO, 2008) and measurements precise to 1 mm. Weights were measured precise to a tenth of a kilogram (kg). The height (weight) measurements were converted into height-for-age z scores (HAZ) and weight-for-age z scores (WAZ) using WHO standards (WHO, 2006a,b; de Onis et al., 2007) based on well-nourished populations. As mentioned earlier, child ages varied between 6 and 17.9 months in Round 1 in our sample. In the analysis that follows, it is desirable to convert these Round 1 anthropometric measures to the same age in months because on average in this and in most undernourished populations HAZ declines significantly over this age range (e.g., Victora et al., 2010). Following the method used in previous studies (e.g., Stein et al., 2010; Crookston et al., 2013; Schott et al., 2013) we project the value of HAZ for the children either older or younger than 12 months as follows. The average HAZ for age i (in months) is defined as the mean of HAZ for all the children age i − 1, i and i + 1 months.7 The projected value of HAZ at age 12 months is the HAZ in Round 1 for each child minus the age-group average HAZ for the actual age of the child in Round 1 plus the average HAZ for children 12 months of age in Round 1. That is, each child is presumed to be the same number of HAZ units above or below the average HAZ curve for all children at age 12 months as the child was at the age actually measured.8 The projected value of HAZ at age 12 months is denoted by HAZ1. We also apply this method, using WAZ in Round 1 to calculate the projected WAZ at age 12 months, denoted by WAZ1. All the outcomes are assessed at age eight years. The anthropometric outcomes are height-for-age (HAZ8) and weight-for-age (WAZ8). The educational outcomes are the test scores in Early Grade Reading Assessment (EGRA), math and Peabody Picture Vocabulary Test (PPVT) the children took at age eight years. The PPVT assesses the receptive vocabulary outcome. Cueto and Leon (2012) analyze the psychometric characteristics of the math test and the PPVT. Household characteristics include: mother’s height, mother’s weight, mother’s ethnicity, mother’s schooling attainment, father’s schooling attainment, and a wealth index in Round 1. The wealth index is the simple mean of three components: (1) housing quality, which is the scaled (0–1) mean rooms per person as a continuous variable and dummies on floor, roof and walls9; (2) the value of consumer durables, which is the scaled (0–1) sum of nine dummies for basic consumer durables10; and (3) the value of health-related infrastructural services, which is the simple mean of drinking water, electricity, sanitation facilities, and fuel, all of which are 0–1 variables.11 Regional characteristics are represented by dummy variables for Urban, Red River Delta, Mekong Delta, and Mountainous. We apply the official list by Committee for Ethnic Minorities of Vietnam to define the Mountainous category.12 Mountainous areas in Vietnam are considered less developed than other areas. Among the mountainous clusters, however, we find one with average birth weight above that of whole sample. Looking into the details, we find high concentration of Tay and Nung ethnic groups in this cluster. Together with the ethnic majority of Kinh, these ethnic groups count for 84% of observations in this cluster. These ethnic groups do relatively well in economic development, and in the decade under research, ethnic minorities groups Tay and Nung were more prosperous than the other ethnic minority groups in the Young Lives sample. For these reasons, we add a dummy variable for this cluster. Similarly, another dummy variable is included for the coastal part of the Mekong region because these sites are distinguished considerably from the other part of the Mekong region. The gaps between the coastal and the inlands of Mekong are most significant in parental schooling and services in general.13 Our analysis involves many variables from multiple rounds of longitudinal data. So potentially there might be a problem of missing values beyond that discussed above with regard to birth weight. We find, however, that there is no concentration of missing values for any particular variable or for any group of children. The final panel consists of 1758 observations, or 89% of the children aged 6–17.9 months from the initial sample. Model specifications ~~~~~~~~~~~~~~~~~~~~ Even though the unpredicted terms y1 and y2 correlate strongly with the corresponding z-scores, they are not perfectly correlated so there is some leeway for having differential predictive power. The correlation between WAZ1 and y1 is 0.90, and that for HAZ1 and y2 is 0.86. The correlation between y1 and y2 is weaker than that between WAZ1 and HAZ1 (0.69 against 0.74). The residual y3 is orthogonal to birth weight, height-for- age and all the regressors in X, and is therefore, orthogonal to the unpredicted height growth y2 as well. Symmetrically, the residual y4 is orthogonal to birth weight, the unpredicted weight gain y1 and all the controls. We consider two groups of models corresponding to the following groupings of variables y0, z1 and z2. In both groups there is a set of models (Models 1–2 in Group I and Models 3–4 in Group II) that include one relation with a weight-related indicator as the dependent variables for Eqs. (1a) and (2a) (indicated by W) and one relation with a height-related indicator as the dependent variables for Eqs. (1b) and (2b) (indicated by H). Group I consists of models that include zero or one of the early childhood nutritional variables: ▓▓Model C: controls only, no variable on child nutritional status.15 ▓▓Model 0: birth weight y0 is the only nutritional variable, e.g. β1 = β2 = 0. ▓▓Model 1W: WAZ1 represents z1, β0 = β2 = 0; ▓▓Model 1H: HAZ1 represents z2, β0 = β1 = 0; ▓▓Model 2W: y1 (Eq. (1a)) represents z1, β0 = β2 = 0; ▓▓Model 2H: y2 (Eq. (1b)) represents z2, β0 = β1 = 0. In Group II, all models contain birth weight and one or two additional nutritional variables: ▓▓Model 3W: birth weight y0 and y1 represents z1, β2 = 0; ▓▓Model 3H: birth weight y0 and y2 represents z2, β1 = 0; ▓▓Model 4HW: birth weight y0, y2 presents z1 and y3 (Eq. (2a)) represents z2. ▓▓Model 4WH: birth weight y0, y1 presents z1, and y4 (Eq. (2b)) represents z2. By estimation of Model C, we will find the explanatory power of the controls, and therefore can assess how much the nutritional variables added in the other models increase the predictive power in the estimation of Eq. (3). Models in Group I are used for estimation of the associations for each individual early-life nutritional variables with the outcomes. Models in Group II contain two or three variables, of which at least one captures postnatal nutrition. These specifications are to demonstrate the relative importance of postnatal nutritional indicators. The purpose of Model 4HW, for instance, is to investigate if weight-for-age, which contains y3 as a component, adds to the predictive power beyond that of the combination of birth weight and the unpredicted height growth. On the other hand, Model 4WH permits investigation of whether the height-for-age, which contains y4, has predictive power beyond that of birth weight and the unpredicted weight gain in the first year of life.","Ordinary Least Squares (OLS) regressions are used for all the estimations. We relax the requirement that the observations be independent, allowing for the possibility that community factors cause correlations within communes that results in heteroskedasticity. To increase the likelihood that the standard errors are “robust” to heteroscedasticity, we apply an option aiming at robust estimators. For the estimates in Tables 2–5 and A2 (in Appendix B) , the Stata option used in the OLS estimations is vce(cluster community), with community being the commune id. Estimates for controls ~~~~~~~~~~~~~~~~~~~~~~ As described above, there is no nutritional indicators among the controls for Model C. The estimates for Model C are presented in Table 2. It can be seen that the gender factor (being a boy) is negatively significantly associated with height-for-age at eight years of age. This association means that the gap in HAZ8 between eight-year boys in Vietnam and the boys of the same age in the WHO sample is less favorable than the gap between eight- year girls in Vietnam and the same-aged girls in the WHO sample. The estimate for gender also implies that boys perform poorer than girls in EGRA. The gender gaps are not statistically significant for the other outcomes. Being first-born is positively significantly associated with anthropometric outcomes and moderately significantly associated with EGRA. Birth year is an important predictor for the educational outcomes, but not for anthropometric outcomes. The omitted category for the year of birth consists of children born in 2001, so the results imply that older children perform better than younger ones. Schooling might have been a part of reason. In fact, in the academic year 2009–10, 92% of the children born in 2001 were in grades 3–4, while only 30% of children born in 2002 were in grade 3 or higher. Mother’s anthropometrics are significant for child anthropometric outcomes at age eight years. With respect to the educational outcomes, there is slight difference between the two anthropometric indicators for mothers. Mother’s height is moderately associated with the EGRA and math tests. The children of ethnic minority mothers perform significantly less well than those of ethnic majority mothers for all the outcomes, except weight. Socioeconomic status (parental schooling and wealth index) is significant for all the educational outcomes. For the child anthropometric outcomes, the associations of mother’s schooling are statistically significant, but those of father’s schooling are not. Finally, residing in the urban sector has statistically significant associations with anthropometric outcomes and reading. Basic estimates for early childhood nutrition indicators ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As discussed in Section 2, 11% of the children have missing data on birth weight, and for these observations, the imputed values are used in the same way as the actual data for the regressions in Table 2. For the estimates in Tables 3 and 4, in addition to the controls used in Model C, a dichotomous variable for “Birth weight imputed (1/0)” is included, as well as the dummy variable on “Birth weight greater than 3.5 kg”. To ease any concern about possible bias related to imputation of data on birth weight, we also run regressions on the reduced sample with only observations that include birth weight data, excluding the imputed values. The estimates are presented in Table A2 in the Appendix B. For the remaining models, we present R2, rather than the differences, which are significant at least for the anthropometric outcomes, unlike that for Model 0. The indicator of weight- for-age at age one year WAZ1 in Model 1W is strongly associated with all the outcomes, except PPVT. The goodness of fit increased from that for Model 0, and the change is most significant for the WAZ8 outcome (0.17, or more than 50%, above that for Model C). The estimates for the nutritional indicator in Model 1W are very similar to the corresponding ones in Model 2W. Weight-for-age z-score at age one year and the unpredicted weight gain in the first year of life are statistically significantly associated with reading and math outcomes, unlike birth weight. The estimates imply that for the anthropometric outcomes, the explanatory power of WAZ1, which includes prenatal nutrition status plus postnatal weight gain, is slightly greater than that of y1. For the non-anthropometric outcomes, however, R-squares in Model 1W are almost identical to the corresponding R-squares in Model 2W, which is further evidence that variation in birth weight has little predicative power for the educational outcomes. For Model 1H, in contrast to Models 0 and 1W, there is a statistically significant association between the nutritional indicator, in this case HAZ1, and PPVT. The estimates for Model 1H are very similar to those for Model 2H. The estimate for HAZ8 in Model 1H is consistent with previous literature, such as Krishna et al. (2016).16 The estimates for Models 3W and 3H in Table 4 are broadly consistent with the parallel estimates in Models 2W and 2H. However, the inclusion of birth weight together with the unpredicted weight gain (height growth) allows comparisons of the predictive power of these variables. The estimates for Model 3W suggest that a one SD increase in the unpredicted weight gain in the first year is associated with variation in the anthropometric outcomes more than double of those for a one SD increase in birth weight. For the EGRA and math scores, the estimated associations of a SD change for the unpredicted weight gain are clearly more significant than the corresponding estimates for a SD change of birth weight. Neither y0 nor y1 is statistically significantly associated with PPVT. As for Model 2H, unpredicted height growth [in Model 3H] is associated significantly with PPVT – that is not the case for the unpredicted weight gain in Models 2W and 3W. Estimates for Model 3H imply that a child born with one SD of birth weight under the mean but with one SD of unpredicted height growth in the first year of life is expected to have better outcomes in mid-childhood than a child born with average birth weight who grew normally (as predicted) in the first year of life. The estimates for Model 4HW suggest that with the presence of the unpredicted height growth (y2) among the regressors, the predictive power of the conditional unpredicted weight-for-age y3 is not statistically significant for the height and PPVT outcomes, and for the math outcome it is not statistically significant at five percent. Related to these estimates, the corresponding results in Table A2 are not statistically significant for all the outcomes in height, math and PPVT. On the other hand, the conditional unpredicted height-for-age y4 in Model 4WH is statistically significantly associated with all the outcomes. In fact, we can see that y4 contributes to prediction of variations in anthropometric outcomes, math and PPVT, all at the level of five percent significance. At the ten percent significance level, one can reject the hypothesis that y4 has no association with EGRA. In Table A2, however, the estimate for y4 is statistically significant at five percent for all the outcomes. By definition, the conditionally unpredicted height-for-age y4 is orthogonal to birth weight y0, the unpredicted weight gain y1 and all the control factors. That means the variations of the outcomes that are predicted by y4 are beyond those predicted by y0, y1 and all the control factors. Finally, it can be seen in Tables 3, 4 and A2 that there are slight differences between the two sets of estimates concerning the magnitudes but not the statistical significance, except for estimates for y3 and for y4 in Model 4WH and 4HW, as discussed in the preceding paragraph. There are no such inconsistencies with respect to the associations of the indicators with birth weight or any of the z-scores, or unpredicted terms y1 and y2. Overall, the results in Table A2 are consistent with the basic findings in Tables 3 and 4.","We start with considering three standard measures of nutritional status in early childhood. The first measure, birth weight, reflects prenatal nutritional intakes. The second and the third are respectively the weight-for-age and height-for-age z-scores at age 12 months, both of which incorporate both prenatal and postnatal nutritional developments. In contrast to some previous studies we do not find strong associations of birth weight with educational outcomes. Probably this is because in the 21 st century Vietnamese context, low birth weight is much less prevalent than in many of the data sets used for previous studies because of their use of twins or because they were for 20th century low- and middle-income contexts with higher prevalence of low birth weight. The estimates for birth weight in the current study are consistent with those in Gupta et al. (2013). In both studies birth weight predicts anthropometric, but not the other outcomes. Our results further are consistent with the studies related to the Dutch famine in 1944–1945. This famine led to lower birth weights of infants whose mothers experienced severe nutritional deprivation during their pregnancy. Stein et al. (1972), however, find that at the age of 19 years there were no detectable adverse effect on cognitive ability of the children born around the time and location of the event. Back to the current study, even for the outcomes for which birth weight has the most significant predictive power – anthropometries at age eight years – it predicts a smaller part of the variance than predicted by any of the weight or height z-scores. Weight-for-age and height-for-age z-scores at age 12 months are strongly associated with all the anthropometric outcomes and educational outcomes at age eight years except the former does not predict the receptive vocabulary outcome. Together with the controls, height-for-age at age 12 months predicts half of the variation in height-for-age at age eight years (more than that predicted by weight-for-age at 12 months for weight-for-age at eight years of age). Correspondingly but not equivalently, weight-for-age at 12 months together with the controls predicts less than half of the variation in mid-childhood weight-for-age. That is evidence for some difference between height and weight in their predictive power – more evidence is summarized below. Closely related to the aforementioned weight and height z-scores are their postnatal components. We define the concepts of unpredicted weight gain (height growth) in the first year of life as the parts of weight-for-age (height-for-age) z-scores that are uncorrelated with birth weight and all the control factors, including mother’s anthropometrics and household socioeconomic status. These unpredicted components of the weight and height z-scores are substantial. In fact, the standard deviations of the variables on birth weight, the unpredicted weight gain, and the unpredicted height growth are 0.43, 0.92 and 0.94 respectively. The sizes of anthropometric outcomes at age eight years associated with the standard deviation of unpredicted weight gain (height growth) are more than double those of a standard deviation in birth weight. Unlike birth weight, the unpredicted weight gain in the first year of life is strongly associated with two of the educational outcomes. Unpredicted height growth in the same period is strongly associated as well with the receptive vocabulary outcome at age eight years. Our study has limitations, including inadequate information on birth outcomes other than birth weight and no information on anthropometrics at age 24 months, which often has been emphasized in the nutritional literature as particularly important (e.g., Victora et al., 2008, 2010). With regard to the latter, however Schott et al. (2013) report that “in analysis using the Institute for Nutrition in Central America and Panama (Guatemala) nutritional supplementation study data (Stein et al., 2008:286), in which there were multiple measures over the first 7 years of age, we found that HAZ measures at ages 6–17.9 months predict well HAZ at age 24 months (for ages 6, 12, and 18 months, correlations with HAZ at age 24 months were r = 0.74, 0.83, 0.91, respectively). If similar correlations across ages also hold for populations of Young Lives countries [in our case, particularly in Vietnam], the cross-sectional patterns in HAZ at ages 6–18 months that were observed in the Young Lives data represent fairly well cross-sectional patterns in HAZ that likely held for these same children at 24 months, even if the overall distribution in HAZ may have declined fairly substantially from 12 to 24 months.” Despite these limitations, our investigation sheds considerable light on how, in the context of Vietnam in the early 21st century, early-life nutritional indicators, together with family characteristics, link to latter outcomes. We find the associations of the socioeconomic status with all the mid-childhood outcomes consistently strong. Our estimates imply a standard deviation higher unpredicted weight gain (or unpredicted height growth) in the first year of life has larger associations with height and weight outcomes at age eight years than a standard deviation higher birth weight. Parents might expect significant adverse outcomes in height and weight at age eight years for children born with lower birth weights. However, expectations about mid- childhood outcomes can change substantially at age one year thanks to developments that were unforeseen at birth because the unpredicted weight gain (height growth) over the first postnatal year are by construction independent of birth weight, socioeconomic status and maternal anthropometry. In addition to the aforementioned difference between the prenatal and postnatal nutrition with respect to their magnitudes of associations with the outcomes, we find difference between the postnatal indicators on height versus weight. That is, there is a component in height-for-age (at age one year) that adds explanatory power for all the outcomes beyond that due to the linear combination of birth weight, weight-for-age z-score (at age one year) and all the controls."],["Research from richer countries finds that dairy consumption has strong positive associations with linear growth in children, but surprisingly little evidence exists for developing countries where diets are far less diversified. One exception is a recent economics literature using the notion of incomplete markets to estimate the impacts of cattle ownership on children's milk consumption and growth outcomes in Eastern Africa. In addition to external validity concerns, an obvious internal validity concern is that dairy producers may systematically differ from non-dairy households, particularly in terms of latent wealth or nutritional knowledge. We re-examine these concerns by applying a novel double difference model to data from rural Bangladesh, a country with relatively low levels of milk consumption and high rates of stunting. We exploit the fact that a cow's lactation cycles provide an exogenous source of variation in household milk supply, which allows us to distinguish between a control group of households that do not own cows, a treatment group that own cows that have produced milk, and a placebo group of cow-owning households that have not produced milk in the past 12 months. We find that household dairy production increases height-for-age Z scores by 0.52 standard deviations in the critical 6–23 month growth window, though in the first year of life we find that household dairy supply is associated with a 21.7 point decline in the rate of breastfeeding. The results therefore suggest that increasing access to dairy products can be extremely beneficial to children's nutrition, but may need to be accompanied by efforts to improve nutritional knowledge and appropriate breastfeeding practices. --------------------------------------------------------------------------------","Worldwide, child undernutrition is increasingly recognized as a significant global health problem and a major constraint to economic development. Child undernutrition is associated with almost 3.1 million child deaths (Black et al., 2013), impaired cognitive development in early childhood (Walker et al., 2011; Grantham-McGregor et al., 2007), reduced school attainment in childhood, and lower labour productivity and wages in adulthood (Shekar et al., 2006; Victora et al., 2008; Hoddinott et al., 2008). Nutritionists, moreover, have increasingly emphasized that it is good nutrition in early childhood – in utero and the first 24 months after birth – that is truly critical for ensuring healthy growth (Shrimpton et al., 2001; Victora et al., 2010). A particularly striking nutritional feature of developing country populations is that growth faltering appears to be particularly pronounced from roughly 6 months of age to 20 months of age, a period that coincides with the introduction of complementary foods that are often low in high quality protein and micronutrients, such as rice, wheat, maize or starchy roots and tubers. Previous research has found that calorie intake alone is not always a strong predictor of child growth in developing countries settings (Griffen, 2016), perhaps because calorie requirements for infants are relatively modest. Instead, many researchers point to low consumption of animal-sourced foods (ASFs) as a critical constraint (Allen, 2003; Brown, 2003; Demment et al., 2003; Headey and Hoddinott, 2016; Neumann et al., 2002; Puentes et al., 2016; Randolph et al., 2007). Indeed, in the absence of fortified foods, young children cannot meet their micronutrient needs without daily intake of ASFs (PAHO/WHO, 2003). Dairy constitutes a particularly important complementary ASF for young children because of the familiarity of its taste to exclusively breastfed children, and because of its nutritional profile. Dairy is high in all three macronutrients (energy, fat and protein), as well as important micronutrients such as vitamin A, vitamin B12, and calcium (Murphy and Allen, 2003). Moreover, like other ASFs, dairy contains several essential fatty acids that are hypothesized to be critical for processes of cellular growth and bone formation (Semba et al., 2016). Dairy has a protein digestibility corrected amino acid score (PDCAAS) of about 120%. Many studies also suggest dairy intake affects child growth through a stimulating effect on plasma insulin-like growth factor 1 (IGF-1). Milk also contains minerals such as potassium, magnesium, and phosphorus, that could also be a factor stimulating growth, as well as lactose. Consistent with this biological evidence, a range of research has linked linear growth to childhood dairy consumption, albeit mostly in developed country samples (Iannotti et al., 2013; de Beer, 2012; Dror and Allen, 2011; Wiley, 2005, 2009; Sadler & Catley 2009). In developing countries there have been remarkably few efficacy trials of dairy supplementation on growth in infants or young children, though several dairy consumption programs have demonstrated some impact on linear growth at older ages (Iannotti et al., 2013).1 Because of the limitations of experimental evidence on this subject, economists have increasingly utilized observational or quasi-experimental analyses to explore the associations between dairy production and child nutrition outcomes in less developed settings. In economic history studies, Baten (2009, 2014) tests a “protein proximity” hypothesis with 19th Century European military recruitment data from Central Europe. Utilizing the idea that fresh milk in these economies could not be traded over large distances, he finds that adult men in closer proximity to dairy production were substantially less likely to be too short for military recruitment. Still other studies hypothesize that trends in milk consumption explain longer term secular improvements in heights at later stages of economic development, such as 20th Century Japan (Takahashi, 1984) and India (Mamidi et al., 2011). A recent paper also examined adult heights in 42 European countries with varying levels of development. Even after controlling for genetic factors, they found that the national supply of protein from dairy products was the single strongest predictor of adult stature (Grasgruber et al., 2014). A related study of 105 countries from different continents also found strong associations between average milk consumption levels and adult male heights (Grasgruber et al., 2016). In contemporary developing countries several studies have examined associations between household livestock ownership and child growth outcomes, though not all studies focus on milk-producing animals specifically. Like Baten (2009, 2014) these studies assume (often implicitly) that fresh milk is generally non-tradable and not a perfect substitute for powdered milk. Hoddinott et al. (2015) use two large surveys from Ethiopia to specifically explore the association between cattle ownership, dairy consumption and HAZ scores. They cite the fact that 90% of milk produced in rural Ethiopia is consumed by the household producing it, implying that cattle ownership ought to be a very strong predictor of regular dairy intake. Consistent with that conjecture they find strong positive associations between cattle ownership and HAZ (as high as 0.47 standard deviations in the 12–23 month age-range). They also implement placebo tests to explore the concern that cattle ownership proxies for generic wealth effects on child nutrition. Rawlins et al. (2014) evaluate Heifer International’s dairy cow and goat ownership programs in Rwanda, albeit in a non-randomized quasi-experimental design with a small sample of 217 children aged 0–59 months (precluding the possibility of detailed age disaggregation). They find that children from households who received a goat 12 months prior to the time of the survey saw no growth differential over controls, whereas transfers of pregnant cows (high-productivity foreign breeds) improved height-for-age Z scores by 0.57 standard deviations, a large but imprecisely estimated effect. Similarly, Kabunga et al. (2017) use matching methods to gauge the impacts of adoption of improved dairy cow varieties on HAZ of children aged 6–59 months. They find HAZ impacts of 0.48-0.49 standard deviations, though also some evidence of larger impacts for household with greater herd sizes or larger acreage.2 Overall, there is fairly consistent evidence that dairy cow ownership is associated with child growth in poorer populations, although there are several limitations and caveats surrounding this evidence. First, the evidence is confined to East African localities where cattle ownership is relatively common, so external validity is a concern. Second, this literature potentially suffers from several internal validity issues, including the confounding role of livestock as a source of imperfectly measured rural wealth, and potential concerns over associations between livestock ownership and ethnicity.3 Another outstanding concern not addressed in the previous literature is that the availability of cow’s milk leads to premature cessation of breastfeeding by mothers. Exclusive breastfeeding is strongly recommended for the first 6 months of life, especially in developing country settings, because of its critical role in preventing diarrhea and respiratory infections (Horta and Victora, 2013), and because cow’s milk can stress a newborn’s immature kidneys and irritate the lining of the stomach and small intestine, leading to blood loss and iron-deficiency anemia (FAO, 2013). In light of these limitations, this paper utilizes a unique dataset to attempt a more comprehensive assessment of the nutritional implications of dairy production and consumption in Bangladesh. Bangladesh is a particularly important case study in the context of dairy production. In addition to its high rates of stunting (36%), Headey and Hoddinott (2016) emphasize that Bangladesh has an under-diversified food supply, with FAO data suggesting that ASFs account for less than 5% of total calories supplied (Headey and Hoddinott, 2016). This situation partly stems from exceptionally low levels of milk consumption, which in per capita terms is less than half that of neighbouring India (Headey and Hoddinott, 2016). A likely explanation of this is the country’s exceptionally severe land constraints (and hence feed constraints), with average farm sizes in Bangladesh averaging just half a hectare, and rural landlessness widespread. It may also be that cultural norms – historical unavailability of milk – has kept demand for milk relatively low. In this paper we use the nationally representative Bangladesh Integrated Household Survey (BIHS) of rural areas, which was conducted over two rounds in 2011/2012 and 2015. Uniquely for such a large survey, this dataset contains rich information both on nutrition outcomes, individual food consumption, agricultural assets and production, and a range of other potential determinants of nutrition. Methodologically, we propose a novel difference-in-difference approach to assessing the impact of dairy cow ownership on child nutrition outcomes, by distinguishing between households with lactating dairy cows that have produced milk over the past 12 months (treatment), households with cows that have not produced milk in the past 12 months (placebo), and households that do not own any dairy cows (control). We note that this is not a placebo in the medical definition (according to which a person consumes a treatment of no intended therapeutic value), but in the sense that non-lactating cows might have a similar long run economic value any direct milk supply to the household. This distinction between the treatment and placebo emerges from the fact that smallholder dairy producers in Bangladesh typically only own a few cows because of the extreme land and feed constraints mentioned above. Specifically, 80% of Bangladeshi farmers in our nationally representative sample own just 1–2 cows and no farmers in our sample own more than 4 animals. Given that at any given time all or some of these cows will not be lactating – since there is a minimum 12-month inter-calving cycle for each animal even among the most technologically sophisticated dairy producers – there is a non-trivial proportion of dairy cow owners in Bangladesh who would be unable to produce milk on a continuous basis for exogenous biological reasons.4 In effect, then, the combination of small herds and a biologically determined component of the lactation cycle potentially creates a valid placebo group of children who are treated with cows that have not produced any milk. We therefore test three hypotheses: Children in treatment group will be taller than children in control; Children in the placebo group will not be taller than the control; and Children in treatment group will be taller than placebo group children. In addition to these tests we also examine whether livestock ownership or milk production is associated with other observable potentially confounding factors, such as maternal nutritional knowledge and empowerment, and overall child diversity, exclusive of milk. And unlike previous studies in this literature we explore the policy-relevant question of whether access to a stable household level supply of dairy products leads to substitution between breastfeeding and dairy milk intake. We find that milk production is strongly associated with linear growth, but only for children in the crucial first 1000 days of life (particularly the 12–23 month range). The effects we observe are very close in magnitude to those observed in the aforementioned quasi-experimental study by Rawlins et al. (2014) for Rwanda and Kabunga et al. (2017) for Uganda, but larger than the more observational study by Hoddinott et al. (2015) who analyse the impacts of owning any cow, rather than milk-producing cows specifically (rendering their results more like an intent- to-treat analysis). Null results for the placebo group also lend credence to the identification assumptions underlying our approach, as do additional placebo tests which rule out systematic differences in nutritional knowledge and women’s empowerment. However, we do find some evidence of potentially harmful effects of household dairy availability on breastfeeding in the first year of life, suggesting dairy-oriented nutrition strategies need to proactively promote exclusive breastfeeding in the first six months to prevent premature substitution into dairy. The remainder of this paper is organized as follows. Section 2 describes the data and the methods used to analyse them. Section 3 tests associations between different ASF production and various nutrition outcomes. Section 4 provides some important sensitivity tests and extensions, and Section 5 concludes with a discussion of the implications of these findings for programs and policies, as well as for future research.","As outlined above, our objective in this paper is to test for significant differences in milk consumption and child growth between household groups that are defined by dairy production and cow ownership. Previous papers in this literature have tended to focus on a comparison between a “treatment group” of households that own any dairy cow and a “control group” of households that do not own any dairy cows. In our data we instead narrow the definition of treatment households to those that owned cows that actually produced milk in the past 12 months (hereafter treatment). We then define what can be thought of as a “placebo group” of children exposed to cows that had not produced any milk in the past 12 months (note that we think of this group as a placebo because the treatment is not milk per se - in which case the placebo would be a milk substitute - but milk-producing cows). In an ideal experimental design children would be randomly assigned across groups, but in observational settings a significant concern is that there may be systematic nutrition- relevant differences between treated and non-treated children (e.g. wealth, nutritional knowledge, women’s empowerment). Achieving more experimental conditions might therefore require extensive control for potential confounding factors. The conceptual model described in Hoddinott et al. (2015) is a useful starting point for thinking about the various factors that might influence household decisionmaking processes with respect to dairy production, dairy consumption and child nutrition. They posit a household utility model in which child nutrition is one argument. Nutritional status is itself a function of nutrient (food) intake, as well as nutritional knowledge, culture, healthcare, genetic endowments, and locational characteristics (such as the prevalence of disease; access to information about good child care practices). In a world of perfectly functioning markets, nutrient intake would be primarily influenced by income, and households could sequentially maximize farm and nonfarm income before deciding how to spend that income so as to maximize nutrition outcomes subject to other arguments in the utility function. However, the perishability of milk in poorly developed value chains renders household production and consumption decisions non-separable. In other words, if households struggle to access affordable milk via markets, they could opt to own dairy cows. This implies that the decisions to own dairy cows and/or produce milk may be endogenous, influenced as it is by nutrition knowledge and farm production parameters such as the availability of capital (income, savings, wealth), access to land (feed), access to input and output markets to obtain feed and sell produce, household labour supply, farm management skills, and the role of women in household decisionmaking, including dairy production and feeding practices. Since omission of these kinds of factors could lead to biased coefficients on the impacts of cattle ownership or milk production on child growth, our empirical models need to control for these factors as extensively as possible. Fortunately, the Bangladesh Integrated Household Survey (BIHS) not only contains detailed data on children’s food intake and nutrition outcomes, but also an exceptionally rich array of data on income, wealth, agricultural production and assets, access to markets, women’s empowerment and women’s nutrition knowledge (IFPRI, 2016). BIHS is also a large survey representative of rural Bangladesh that has been implemented in two rounds (2011/2012 and 2015) and constitutes a panel for the majority of households. However, because we are interested in child growth in the first 5 years of life – particularly the 12–23 month period – we treat both rounds as repeated cross-sections rather than a panel.5 The combined rounds make up a sample of 11,796 households (some surveyed twice), which includes 4268 pre-school children aged 0–59 months. Height-for-age Z scores (HAZ), using the World Health Organization’s global child growth reference standards (WHO, 2006), constitutes our primary outcome of interest. As noted above, from 6 months to around 24 months growth faltering tends to be particularly pronounced in developing country populations due to prolonged nutritional deficiencies associated with inappropriate complementary feeding and repeated or chronic infections (Victora et al., 2010). It is also common to define children as stunted if HAZ falls below -2, though statistical epidemiologists have strongly argued against using dichotomous dependent variables, as it unnecessarily discards valuable information and reduces precision (Royston et al., 2006). However, we report stunting results as an extension to our main HAZ results. In this paper our interest is in dairy production-dairy consumption pathways, rather than dairy production-income/wealth pathways (in principle, income from any source could improve diets). Our regression models therefore control for household expenditure and wealth, but our dataset also allows us to examine whether dairy consumption is likely to be the main mechanism linking cow ownership to child growth by using additional data on children’s consumption of various foods as well as household data on how different foods were obtained. In terms of the former we primarily focus on children’s consumption of dairy products in the past 24 h, defined as a dichotomous indicator. To help rule more generic income-based pathways we also use a dietary diversity score (0–6 food groups) that excludes dairy, as well as estimates of children’s total calorie consumption (excluding breastmilk). Our expectation is that dairy production influences dairy consumption, but not non-dairy dietary diversification or total calorie intake. We can also explore how households sourced different foods since the BIHS asks respondents to estimate the proportion of each food provided through market purchases, provided by other sources, or provided by home production. We also note that, in principle, these consumption data might also be used to examine the impacts of dairy consumption on child growth. However, a critically important limitation of consumption data is that they are based on short recall periods (24-hour or weekly recall), meaning that they are potentially quite poor indicators of regular consumption of milk in the past 12 months or more (Thorne-Lyman et al., 2014). This measurement problem with short-recall consumption suggests that longer recall questions on milk production may be a much better indicator of regular access to dairy products in settings where markets for perishable products are highly imperfect. However, since long-recall production quantity indicators also suffer from bias we use a simpler dichotomous indicator of whether or not milk was produced in the last 12 months – along with cow ownership - to define our treatment, placebo and control groups. Clearly these groups are not the result of random assignment, although we can use multivariate regressions to reduce the biases of confounding factors that influence cow ownership or lactation decisions. We first assess the determinants of milk production, with the expectation that cattle herd size (female and males) is a key observable driver that we can subsequently control for in our main HAZ regressions. We then use multivariate reduced form regressions to control for a broader range of potential confounding factors. In addition to dairy herd size, we were also concerned that cattle ownership may simply reflect more generic livestock wealth, so we extensively control for other forms of livestock (bullock/buffalo, goat, sheep, chicken, duck and other birds) and aggregate livestock into an index of Tropical Livestock Units (TLU), which can be thought of as a measure of aggregate livestock wealth. The remaining control variables are more common to most nutrition specifications, and to estimation of health production functions, such as Todd and Wolpin (2007) and Hoddinott et al. (2015). This includes child characteristics (sex, age, breastfeeding status), parental characteristics (age and schooling), household characteristics (per capita monthly expenditure, the aggregate value of 26 household assets, hectares of cultivable land owned, household toilet and water access, access to electricity, exposure to NGO services) and several community characteristics (distances to the nearest weekly/periodic outdoor market, and to the nearest town and to the nearest health centre). Our regressions also include fixed effects for all 65 districts in which the BIHS was conducted. Clearly the main coefficient of interest is that pertaining to the treatment group, which we interpret as the effect of dairy availability on child growth net of any impacts of dairy production on other inputs into the health production function, such as income, or changes in breastfeeding. However, we also test for significant differences between the coefficients for treatment and placebo, and whether the coefficient for placebo is significantly different from zero (i.e. from the control, the omitted control group). A significant coefficient on placebo would suggest that cattle ownership influence HAZ through channels other than dairy consumption. A biological issue of paramount importance is the need to explore age- specific variation in the sensitivity of children’s growth to exposure to dairy production, an issue emphasized in Hoddinott et al. (2015). For the HAZ analysis we primarily focus on children 6–23 months and 24–59 months, as well as smaller age intervals. The biology of growth identified in Victora et al. (2010) suggests that most growth faltering takes place in the 6–23 month window, so dairy consumption in this period ought to be critical. We do report results for older children (24–59), although it is not clear that our 12-month dairy production indicator should predict stronger growth because of misclassification errors. That is, some 24–59 month children who may have consumed dairy in the past 12 months (according to our indicator) may not have consumed dairy in their critical 6–23 month window. In our extensions to the basic model we also examine two indicators that were not collected for all households and would therefore entail sample restrictions: maternal nutrition knowledge score and a maternal empowerment score based on women’s control over and ownership of various agricultural assets. We use these indicators as dependent variables to test whether dairy producing households are significantly more likely to have mothers with better nutrition knowledge or greater empowerment. Here we test the null hypotheses that the coefficient on treatment is equal to that of placebo and control. Rejection of this null would cast doubt might suggest that part of the estimated effects of milk production on HAZ pertains to greater nutrition knowledge or empowerment. We also estimate alternative HAZ specification where production quantities of milk are used in place of the dummy variable for any milk produced. This is not our preferred indicator because of concerns over measurement error, related to the challenges of accurately recalling production over a long period, but we nevertheless consider it a useful alternative test. Descriptive statistics ~~~~~~~~~~~~~~~~~~~~~~ Table 1 provides descriptive statistics for the key variables for a sample of children 0–23 months of age. Fig. 1 also reports a local polynomial smoother curve (LPOLY) of HAZ scores against child age to reveal the dynamics of growth faltering in rural Bangladesh. There are several broad inferences to be made from these results. First, the sample of children is highly undernourished, consistent with other nationally representative surveys of Bangladesh. Mean HAZ scores are −1.37, and one third of children are stunted (by age two fully half are stunted). However, consistent with previous research (Victora et al., 2010), most of the growth faltering in Bangladesh occurs in the 6–23 month window, as shown by the red vertical lines in Fig. 1. This accelerated period of growth faltering could partially be due to poor diets. Notably, the percentage of all children who consumed dairy in the past 24 h is just 22%, which is particularly low given that in more developed societies many children would consume milk on a daily basis. Consistent with low milk consumption is the low ownership of milk producing cows (14%), while a further 8% own a cow that has not produced milk in the past 12 months. Determinants of milk production ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The higher socioeconomic status of treatment might imply that any apparent benefits of dairy production partially reflect the benefits of greater socioeconomic status. This points to the importance of multivariate regression models saturated with a wide array of controls, as well as the importance of placebo tests. However, we can also examine the determinants of milk production across among households that own at least one cow (treatment and placebo) to assess the relative importance of herd size versus other socioeconomic indicators. On biological grounds one would expect milk production to be strongly associated with herd size, including the number of both female animals and male animals. Owning more female animals obviously reduces the risk that the herd as a whole will not have produced any milk in the past 12 months. However, without male animals, producers would need to either rent in bulls, or access artificial insemination services. While the latter are common in Bangladesh, previous research points to poor farm management practices reducing the success of artificial insemination services (see footnote 3). Table 2 reports the results for those variables that statistically explain whether or not a cow-owning household has produced milk in the past 12 months. With the exception of maternal age, the only significant predictors of dairy production status are indicators of herd size; coefficients on the range of other indicators of household socioeconomic status are all insignificant, individually and jointly. The results suggest that milk production status is non-linearly related to herd size: owning 2 dairy cows or 1 bullock greatly increases the probability of producing milk in the past year, but additional animals do not much alter these probabilities. Fig. 2 explores the relationship between herd size and annual milk production on the y-axis and the number of cows owned on the x-axis. However, we plot a curve for households that own at least one bullock, as well as those that do not, in order to examine interaction effects. The results reveal the expected finding that owning just one cow with no bullock results in very low levels of milk production because there is a high likelihood that this single cow may not have been lactating at any time in the past 12 months. Owning more cows greatly improves milk production. Moreover, the returns to owning one cow and at least one bullock are fairly high, and not greatly increased by owning more cows. Associations between dairy production and dietary indicators ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Fig. 3 shows local polynomial smoother plots of the relationship between 24-hr dairy consumption and child age, with 90% confidence intervals (CIs). We use 90% CIs in order to implement a one sided test at the 5% level that treatment status is associated with higher HAZ. Panel (i) compares treatment to placebo, while Panel (ii) compares treatment to control. The 90% CIs do not overlap in either panel, indicating that treatment children have significantly higher levels of dairy consumption compared to the placebo or control groups throughout the 0–59 month age range. The magnitude of the difference between treatment and placebo and control varies between 15–25 percentage points depending on the age of the child. Table 3 examines this relationship in a multivariate regression model with a full set of controls, but also looks at whether milk production has any impact on non-dairy dietary diversity and child calorie intake. Results in Regression (1) suggest that milk production leads to approximately a 14-point increase in milk consumption, although some of the difference in milk consumption across groups observed in Fig. 3 is likely driven by differences in socioeconomic status (household expenditure, maternal education) across groups.6 Another striking result from Fig. 3 is that many children under the age of 12 months consume cow’s milk, even though recommendations (albeit based more on developed country samples) recommend milk consumption be initiated only at 12 months (FAO, 2013). Moreover, previous research using the same dataset suggests that children are often given the lion’s share of a household’s milk supply in Bangladesh (Sununtnasuk and Fiedler, 2017). Finally, Table 3 also examines whether there are systematic differences in non-dairy dietary diversity across groups, as well as total calorie intake (exclusive of breastmilk). We find no significant associations between treatment and these two dietary indicators, although the placebo group has higher calorie intake than the treatment or control groups. The lack of any impact on non-dairy dietary diversity suggests the results may not be confounded by socioeconomic differences between groups (Hoddinott, Headey and Dereje 2014). The lack of a significant impact on calories suggests that milk consumption is not primarily operating through increasing a child’s overall calorie intake in this context. Associations between dairy production and child growth ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 4 presents least squares regression results with a full set of control variables, stratified by 6–23 months, 24–59 months, and then by series of overlapping 12-month age brackets used to further corroborate the importance of milk in this 6–23 month window. The most striking result is the large 0.52 standard deviation (SD) difference between treatment and control children in the 6–23 month window; a difference which entirely disappears in the 24–59 month window. The latter result is likely explained by the fact that there may be low serial correlation between milk production in the past year and milk production in earlier years, precisely because of variations in lactation cycles among small-scale dairy producers. In columns (3) and (4) we see that the results are consistent across the 6–17 month and 12–23 month windows, though column (4) shows a relatively large but insignificant coefficient on the placebo group coefficient, while column (5) confirms that the benefits of milk production are no longer apparent once we move above the 23 month threshold. We interpret this as evidence that milk consumption has its largest impact in the first 1000 days; as the age range moves beyond ∼23 months the 12 month recall becomes a more imprecise indicator of whether the child actually consumed milk in the 6–23 month period. Further confirmation that the results are strongest in the 6–23 month period is provided by Wald tests of significant differences between the treatment and placebo coefficients in the 6–23 month, 6–17 month and 12–23 month ranges. This suggests that it is milk production, not cattle ownership per se, that yields sizeable benefits for linear growth in early childhood. Extensions ~~~~~~~~~~ In addition to the results above we also engaged in a series of extensions designed to explore some additional complexities in the associations examined above. We first tested for differential impacts of treatment on boys and girls, but found no statistically significant differences in results for the age ranges above. We also tested for interactions between treatment status and maternal empowerment scores and maternal nutritional knowledge on the grounds that these might be mediating factors, but all interactions were insignificant. We also included empowerment scores and knowledge scores as dependent variables to see if these might be potential confounding factors, but treatment status had no significant impact on either variable (results available on request). In Table 5 we used stunting status (HAZ<-2) as the dependent variable. Stunting is a widely used public health measure, although using a dichotomous indicator rather than a continuous indicator effectively discards information and is likely to reduce precision. The pattern of results in Table 5 are very similar to those reported in Table 4, although the Wald tests no longer report statistically significant differences across the treatment and control groups (seemingly due to the expected increase in imprecision). That caveat aside, the results imply that regular dairy consumption has strong impacts on stunting, although treatment-control and treatment-placebo comparisons yield quite different inferences. Among children 6–23 months the model predicts a 10.4-point reduction in stunting relative to the control group. However, the placebo group also has a large, negative but statistically insignificant coefficient that – interpreted literally – would imply only a 2.4-point reduction in stunting from exposure to treatment. Among children 12–23 and 18–29 months the point estimates on treatment are even larger, implying 14 and 22-point reductions in the risk of stunting relative to control, and 8.4-point and 11.3-point reductions relative to the placebo group. Overall, then, these results for stunting status are broadly similar to the HAZ results in Table 4, although it is no longer possible to establish statistically significant differences between treatment and placebo. An alternative to modelling a dichotomous indicator of whether the household produced any milk is to specify the household’s estimate of the quantity of milk it produced in the past 12 months, which we measure as the log of litres per child. OLS coefficients estimates for this indicator are reported in Table 6. These coefficients are significant in the 6–23 month and 12–23 month brackets, and marginally insignificant in the 6–17 month bracket. In the 6–23 month range the coefficient implies that increasing milk production by 10% would reduce stunting by 0.08 percentage points. The coefficients are imprecisely estimated, however, and likely suffer from attenuation bias related to the significant challenges that respondents have in accurately answering 12-month recall questions. Overall, though, the results are broadly consistent with the results from Table 4.","One concern with the results reported in Fig. 2 is that many children in the treatment group consume dairy at young ages (Fig. 3) when it may be harmful to the infant digestive system (FAO, 2013), or may substitute for breastmilk, which has been linked with a range of desirable health outcomes (PAHO/WHO, 2003). In this section we explore whether there might be substitution between breastmilk and household supplies of dairy milk. Fig. 4 plots breastfeeding status by child age with comparisons between treatment and placebo (Panel i) and treatment and control (Panel ii). The results show that, from birth to around 8 months of age, dairy-producing households are significantly less likely to breastfeed their children. Above this age range there is no significant difference in breastfeeding rates. This suggests that access to dairy milk may have a negative spillover on breastfeeding practices in the critically important 0–5 months age range when it is strongly recommended for infants to be exclusively breastfed. In Table 7 we estimate a linear probability model with current breastfeeding status as the dependent variable, with the usual battery of control variables included to see whether the results in Fig. 4 are robust to a multivariate model. This is indeed the case: children 0–11 months from treatment households are around 21.7% less likely to be breastfed than children from control households, and there is a similar statistically significant difference between treatment and placebo. The evidence therefore suggests that easy access to dairy milk greatly reduces the incentive for mothers to breastfeed.","Despite strong biological evidence on the links between dairy consumption and child growth, and substantial empirical evidence from developed country populations, surprisingly little research has documented the impacts of regular consumption of dairy products on child growth in developing countries. Recent economic research has instead examined associations between cattle ownership and child growth, but only looked at East African populations. And to our knowledge none of this research has examined substitution of dairy milk for breast milk. In this paper we examined these associations in Bangladesh where we were able to distinguish between cows that produced milk in the past 12 months and those that did not. This dichotomy served two purposes. First, by focusing more specifically on herds that have actually produced milk our estimates may more closely approximate the growth benefit of the latent variable of interest, the regular consumption of dairy products. Second, “treating” children with cows that have not produced milk offers a potentially meaningful placebo test. We find results broadly consistent with the findings of East African settings. Similar to Hoddinott et al. (2015), we were able to disaggregate results by age and show that the benefits of cattle ownership (or regular supply of dairy products) emerges primarily in the 6–23 month critical window of child growth. However, Hoddinott et al. (2015) find an estimated impact of owning at least one cow of 0.21 standard deviations, without knowing whether the cow produced milk or not. When we replicate that approach we find an impact of 0.35 standard deviations for owning any cow (results available on request), whereas the results reported above suggest an estimated impact of 0.52 SD for owning at least one cow that produced milk in the past 12 months. Hence the associations estimated in this paper are substantially larger and partially pertain to the use of a better proxy for regular milk consumption. Our point estimates are very similar in magnitude to those of Rawlins et al. (2014) from Rwanda, and Kabunga et al. (2017) from Uganda, even though both of those studies focus on improved (high-yielding) cattle varieties rather than ownership of any type of dairy cow. This literature therefore corroborates existing evidence on the importance of cow’s milk for linear growth, which mostly stems from more developed settings (Iannotti et al., 2013; de Beer, 2012; Hoppe et al., 2006). Given that less than a quarter of rural Bangladeshi children consumed dairy products over the previous 24 h, and that almost half of rural Bangladeshi children are stunted, increasing dairy consumption among children and women of childbearing age should be a central priority for nutritional strategies in Bangladesh. The best means of doing so is unclear, however. With exceptionally high population densities even in rural areas, Bangladesh has no clear comparative advantage in large- scale dairy production and may ultimately need to rely more on milk powder imports, which are still heavily taxed with a tariff of 25%. Additional constraints may be more cultural in nature. Like many East Asian countries, Bangladesh has no strong tradition of milk consumption. However, several East Asian countries, such as Thailand and Vietnam, have been extremely successful in increasing dairy consumption through combinations of imports and rapid growth in domestic production, as well as marketing campaigns and school feeding programs aimed at increasing nutritional knowledge and consumer demand for dairy products (FAO, 2008). However, our results also provide a further rationale for utilizing campaigns aimed at improving nutritional knowledge; that there is a need to reduce the perceived substitutability between dairy products and breastmilk."],["Population-level analysis of dietary influences on nutritional status is challenging in part due to limitations in dietary intake data. Household expenditure surveys, covering recent household expenditures and including key food groups, are routinely conducted in low- and middle-income countries. These data may help identify patterns of food expenditure that relate to child growth. Objectives We investigated the relationship between household food expenditures and child growth using factor analysis. Methods We used data on 6993 children from Ethiopia, India, Peru and Vietnam at ages 5, 8 and 12y from the Young Lives cohort. We compared associations between household food expenditures and child growth (height-for-age z scores, HAZ; body mass index-for-age z scores, BMI-Z) using total household food expenditures and the “household food group expenditure index” (HFGEI) extracted from household expenditures with factor analysis on the seven food groups in the child dietary diversity scale, controlling for total food expenditures, child dietary diversity, data collection round, rural/urban residence and child sex. We used the HFGEI to capture households’ allocations of their finances across food groups in the context of local food pricing, availability and pReferences Results The HFGEI was associated with significant increases in child HAZ in Ethiopia (0.07), India (0.14), and Vietnam (0.07) after adjusting for all control variables. Total food expenditures remained significantly associated with increases in BMI-Z for India (0.15), Peru (0.11) and Vietnam (0.06) after adjusting for study round, HFGEI, dietary diversity, rural residence, and whether the child was female. Dietary diversity was inversely associated with BMI-Z in India and Peru. Mean dietary diversity increased from age 5y to 8y and decreased from age 8y to 12y in all countries. Conclusion Household food expenditure data provide insights into household food purchasing patterns that significantly predict HAZ and BMI-Z. Including food expenditure patterns data in analyses may yield important information about child nutritional status and linear growth. --------------------------------------------------------------------------------","Globally, 165 million children are stunted and 50 million children are wasted (Black et al., 2013a). Stunted and wasted children suffer short- and long-term consequences (Walker et al., 2005; Crookston et al., 2011; Victora et al., 2008; Behrman, 2014); therefore, improving children’s nutrition is a global priority (United Nations, 2014; United Nations, 2016). Food intake is one of the causes of undernutrition. It is difficult to identify through population-level analyses what aspects of food intake drive poor nutritional status, in part because information on influences on the food choices that determine consumption is often lacking. Heterogeneities in food consumption may be considerable across households because of variations in preferences, food prices, food availabilities and resource constraints. Investigating patterns in food expenditures at the household level may provide a novel tool for assessing child and household nutritional risk. Furthermore, accurate quantitative measures of dietary intake are time-consuming to obtain, and require extensive food composition databases and nutritional expertise for data collection and analysis (Willett, 1998; Gibson, 2005; Magarey et al., 2011; Fiedler et al., 2013). Additionally, when researchers, program planners and evaluators collect data on foods and liquids consumed in the previous 24 h, this information may not reflect usual intake (Sempos et al., 1985). Because policy makers lack access to information on usual patterns of dietary intake, it is challenging to determine best approaches for improving individuals’ consumption of food. Data on household food expenditures (defined as market purchases, gifts and foods drawn from own production or stocks consumed by the household) reflect periods longer than 24 h (often 2 weeks) and are consistent, for example, with Living Measurement Surveys conducted by The World Bank and national statistical bureaus (Long et al., 2007). Field workers collecting household expenditure data need to be trained. Similarly, individuals who collect and analyze data using 24-h recalls need specialist training. However, in contrast to use of expenditure data, analysis of dietary intake data requires regular, time-consuming updates of food data bases with information on food preparations, corrections for cooking, waste and portions (Gibson, 2005). Additional advantages and disadvantages of 24-h recall data and expenditure data are outlined in Box 1. Many governmental and non-governmental organizations conduct household consumption and expenditure surveys (HCES) that include information on how much households spend on key food groups. They conduct these surveys every 3–5 years in more than 125 low- and middle-income countries (Fiedler et al., 2012a). Several studies demonstrate significant relationships between nutrient intake based on dietary intake and nutrient intake derived from food expenditures (Naska et al., 2001; Jariseta et al., 2012) after converting food expenditure data to estimates of individual dietary intake (Fiedler et al., 2012a). The HCES surveys are important sources of information on family food choices. Some researchers have converted HCES data to estimates of dietary intake to identify implications for nutritional policy and planning (Jariseta et al., 2012; Fiedler et al., 2012b; Bermudez et al., 2012). One of the important benefits of utilizing HCES data is the lower cost of data collection. According to one estimate, the cost of collecting and analyzing 24-h recall data is estimated to be 75 times that of utilizing HCES data on household food expenditures (Fiedler et al., 2013). We hypothesize that household dietary patterns, as reflected in food expenditure data, are important drivers of family and individual dietary quality which will be manifested in measures of child growth. Household food expenditure surveys have been used to explore determinants of child nutritional status (Bermudez et al., 2012; Campbell et al., 2010; Sari et al., 2010; Torlesse et al., 2003; Mauludyani et al., 2014). Associations have been observed between specific food expenditure patterns and nutritional status. For example, in Indonesia, Sari and colleagues documented that higher household expenditures on non-grain and animal- source foods reduced the risk of stunting among children ages 0–59 months (Sari et al., 2010). In these studies, researchers used information on expenditures for specific food groups and combinations of food groups, selecting those food groups that have been shown most frequently to be associated with nutritional status. Data reduction approaches such as factor analysis and principal components analysis (PCA) have been used by some investigators to identify patterns in food consumption based on dietary intake data (Becquey et al., 2010; Bedard et al., 2015; Devlin et al., 2012; Moskal, 2014; Pisa et al., 2015). We extend the use of data reduction approaches to household food expenditures to determine whether there are underlying drivers of food expenditure patterns that are associated with child growth. Fig. 1 represents the conceptual framework guiding this research. We posit that household characteristics (preferences, resources, demographics) and community characteristics (food prices and availability, urbanization) underlie three important indicators of food consumption: 1) allocation patterns of household expenditures across food groups, 2) dietary diversity (an indicator of likelihood of achieving necessary micronutrient intakes), and 3) total household food expenditures (an indicator of dietary quantity). We test the hypotheses that each of these three is associated with children’s nutritional status. In this study, we examine a cohort of children from four diverse low- and middle-income countries (Ethiopia, India, Peru and Vietnam) using food group expenditure data and a variety of analytic strategies. We characterize household food group expenditure patterns and estimate their associations with children’s nutritional status ((height-for-age (HAZ) and body mass index (BMI-Z)) during early and middle childhood. Our analyses are innovative because 1) our data cover children from ages 5–12 years, 2) we assess the broad applicability of findings across four diverse settings, and 3) we describe associations between food group expenditure patterns and children’s nutritional status. This work contributes to the ongoing conversation about using HCES data for nutrition policy (Fiedler et al., 2012a; Fiedler, 2013), while also exploring a novel approach to identifying patterns in HCES data that is less expensive and less dependent on assumptions than converting HCES expenditure data to estimates of individual level dietary intake. Study design and participants ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We used data from the Young Lives (YL) study younger cohort, which is comprised of ∼8000 children in Ethiopia, India, Peru, and Vietnam. The YL study team recruited ∼2000 children from each country in 2002 (round 1) at approximately 1y of age with subsequent data collection at age 5y (round 2: 2006), age 8y (round 3: 2009) and age 12y (round 4: 2013). YL used a multistage sampling design which was pro-poor, with the first stage consisting of selection of 20 sentinel sites. In Ethiopia, the sampling universe included the most food-insecure areas. In Peru, the richest 5% of districts were excluded from the sample. While poor clusters were moderately oversampled, the final samples represented a variety of social, geographic, and demographic groups. The sample in India consisted of households from Andhra Pradesh and Telangana, while the three other countries used nationwide samples (Humphries et al., 2015). Children’s ages at each round ranged from 6 to 18 months (hereafter “1y”, round 1), 4.5 to 5.5 years (“5y”, round 2), 7.5 to 8.5 years (“8y”, round 3) and 11.5 to 12.5 y (“12y”, round 4). Sampling methods have been reported previously (Humphries et al., 2015). Additional study methods are described elsewhere (Barnett, 2012), and are provided at http://www.younglives.org.uk (Young Lives., 2015). From age 1y to age 12y, the YL country cohorts lost between 1.5% and 5.7% of participants to attrition (Ethiopia 114/1999; India 81/2011; Peru 106/2052; Vietnam 36/2000). Household food expenditures were collected starting with round 2; consequently, this study included data from rounds 2, 3 and 4. Children were excluded from each round if they were missing information on anthropometry, rural/urban residence or household food expenditures. Final sample sizes for rounds 2–4, respectively, were: Ethiopia (1744, 1742, 1733); India (1804, 1806, 1801); Peru (1795, 1788, 1775), Vietnam (1788, 1754, 1684). Household food expenditures The respondent (usually the mother) was asked to report on consumption of between 21 and 33 food categories in the previous two weeks. For each food group, the respondent estimated (a) total expenditure on all individual foods in that food group consumed in the previous 2 weeks; (b) value of gifts received or food paid in lieu of wages in that food group; and (c) value of own stores used in that food group (whether from the household’s own production, shop or stocks). We aligned food groups across rounds and across countries to generate 18 groups that were included in rounds 2, 3 and 4. We then aggregated food groups in two different ways in preparation for factor analyses. Approach 1: For food groups for which many households reported no expenditures, we aggregated groups into combinations of similar foods to make, for example, a fruit and vegetables group and an animal source foods group (meat, fish, eggs, and dairy). Approach 2: We aggregated food group expenditures to align with the seven food groups in the WHO child dietary diversity measure (World Health Organization Department of Child and Adolescent Health and Development et al., 2010). Based on the YL household census conducted as part of data collection, we generated the number of adult equivalents (AE) in the household (Glewwe and Twum-Baah, 1991) and converted food group expenditures to per AE values. We adjusted food expenditures independently in each country based on country-specific consumer price indices (Young Lives., 2015). These values were adjusted to 2006 local currency and adjusted again to facilitate international comparability using purchasing power parity conversions (Taylor, 2002; Taylor and Sarno, 2004). We summed total food expenditures across food groups to give a total food expenditure per adult equivalent. We winsorized expenditures in each food group to address outliers (Reifman and Keyton, 2010) for each of the four countries, replacing values below the first percentile with the 1st percentile, and replacing values above the 99th percentile with the 99th percentile. Ethiopia (15.5, 16.2, 17.4) had the lowest total food expenditures per adult equivalent at each time point, and Peru (41.1, 47.3, 55.3) had the highest (Table 1). Ethiopia was the only country with the preponderance of food expenditures in the whole cereals category (Table 1). Peru (10.56–15.48) and Vietnam (8.54–12.0) had similar median levels of expenditures on animal source foods (ASF). In Peru ASF expenditures included dairy, meat and fish and in Vietnam ASF expenditures were predominantly meat and fish. Median egg expenditures were similar in Peru and Vietnam, lower in India, and minimal in Ethiopia (Table 1). The median level of expenditures on ASF in India (3.64–4.23) was about a third of the Peru and Vietnam medians, and the median in Ethiopia (0.59–0.86) was less than a quarter of the Indian expenditures (Table 1). Median fruit and vegetable expenditures were highest in Vietnam (3.35–4.45), followed by Peru (2.39–4.41), India (2.60–3.71) and Ethiopia (0.42–0.56) (Table 1). Child anthropometry Field workers measured height using locally-made stadiometers with standing plates and moveable head boards accurate to 1 mm. We calculated HAZ using WHO 2006 standards for children 0–59 months (de Onis et al., 2006) and WHO 2007 standards for older children (de Onis et al., 2007). Field workers measured weight using calibrated digital balances (Soehnle) with 100 g precision. We calculated body- mass-indices for age (BMI-Z) using WHO growth curves. All measurements were taken according to WHO guidelines (World Health Organization Department of Nutrition for Health and Development and Development, 2008; United Nations Department of Technical Co-operation for Development and Statistical Office, 1986). Birth dates were drawn from children’s health cards when available, and mothers’ reports otherwise. Our analyses focused on HAZ and BMI-Z as the standard indicators of malnutrition in children (Black et al., 2013b). Dietary diversity We assessed individual dietary diversity by summing the number of standard dietary diversity food groups that the child was reported to have eaten the previous day (World Health Organization Department of Child and Adolescent Health and Development et al., 2010). We combined food groups into the seven recommended categories at age 5 y, including (1) starches (cereals, roots and tubers), (2) meat (meat, fish), (3) eggs, (4) legumes and nuts, (5) dairy, (6) fruit and vegetables, and (7) fats and oils. At ages 8 y and 12 y, vitamin A rich fruits and vegetables were recorded as separate categories as in the adult version of dietary diversity. Control variables Other measures included round of data collection, sex of child, and rural/urban residence. Statistical methods ~~~~~~~~~~~~~~~~~~~ We used Stata (version 14.0, 2013. Stata Corp) for all analyses. Results were considered statistically significant for p values < 0.05. We present associations between household food expenditures, dietary diversity and child growth (HAZ, BMI-Z) using (a) total household food expenditures and (b) factor analysis extracted from household expenditures on the dietary diversity food groups, controlling for total food expenditures, round of data collection, rural/urban residence and sex of child. Factor analysis: we pooled food expenditures at ages 5y, 8y and 12y in each country then analyzed food expenditures at the country level. We ran both factor analysis and PCA with one, two and three factors for each country, with the upper limit of three factors based on scree plots and eigen values >1.0. We assessed each PCA and factor solution for loading >0.40 for at least three food groups on each component or factor. None of the PCA solutions in any of the four countries met those criteria. The one-factor solution in all four countries met these requirements, so we utilized the one-factor solution for this analysis. The one-factor solution represents the latent driver of households’ allocations of their food finances across food groups, and we refer to the latent variable as the ‘household food group expenditure index’ (HFGEI), which represents food choices within the constraints of household preferences, household resources, and local food pricing and availability. We used multivariable ordinary least squares regressions for HAZ and BMI-Z to examine associations between food expenditures and anthropometry.","Mean HAZ increased from ages 5y to 12y in India (−1.65 to −1.45), Peru (−1.53 to −0.97) and Vietnam (−1.35 to −1.06), although it remained constant for Ethiopia (−1.48 to −1.47) (Table 2). BMI-Z decreased from ages 5y to 12y in Ethiopia (−0.62 to −1.82), India (−1.16 to −1.35) and Vietnam (−0.31 to −0.65). Mean HAZ was negative in all countries across all rounds, and mean BMI-Z was negative across all rounds in all countries except Peru, where mean BMI-Z ranged from 0.52 to 0.67 (Table 2). The percentages of individuals who lived in rural areas remained relatively constant in Ethiopia (59–60%), India (72–74%) and Vietnam (80–81%), although Peru experienced a decrease from 45% to 27% because of migration to cities (Table 2). In all countries there was an increase in mean dietary diversity from age 5y to age 8y, and a decrease from age 8y to age 12y (Table 2). Peru had the highest dietary diversity at each round, and Ethiopia had the lowest. Five food groups loaded on the first factor, HFGEI, in all of the countries (starches, fruit and vegetables, meat and fish, eggs, and fats), contributing significantly to the common variance across all of the food groups (Table 3). However, we found that relative loadings for each of the food groups were quite different across countries. Correlations of HFGEI with food group expenditures also varied across the four countries (Fig. 2). For example, the correlation between meat and fish expenditures and HFGEI were lowest in India (0.54), slightly higher in Ethiopia (0.57), and highest in Peru (0.76) and Vietnam (0.80). Households in the lowest quartile of HFGEI in Ethiopia spent on average 57% of their food expenditures on starches (Table 4). Adjusted expenditure values allow cross-country comparisons. Based on mean percentages of total food expenditures on specific food groups, we present a picture of how households allocate their food budget. Across all four countries, mean percentages of food expenditures on meat and fish were lowest for households in the lowest quartile of HFGEI, and highest in the highest quartile of HFGEI (Table 4). We observed the same pattern for fruit and vegetable expenditures for Peru and Vietnam, although in Ethiopia and India there were minimal changes in percentage expenditures on fruit and vegetables across the four quartiles of HFGEI (Table 4). Our findings suggest that total food expenditures were positively associated with HAZ (p < 0.05) and BMI-Z (p < 0.05) in all four countries after adjusting for study round, rural residence, and whether the child was female (Table 5). With adjustments for dietary diversity and HFGEI, total food expenditures were no longer a significant predictor of HAZ in Ethiopia and India. HFGEI was associated with significant increases in child HAZ in Ethiopia, India, and Vietnam with adjustments for data collection round, total food expenditures, dietary diversity, rural residence and whether the child was female (Table 5). In the fully-adjusted models that included study round, HFGEI, dietary diversity, rural residence, and whether the child was female, total food expenditures remained significantly associated with BMI-Z in India, Peru and Vietnam. Dietary diversity was inversely associated with BMI-Z in India and Peru; namely a higher diversity score was associated with lower BMI-Z. Coefficients for both countries were essentially the same (−0.048 for India and −0.047 for Peru) as were p values (0.04 and <0.01). It is important to note that mean BMI-Z in India was negative at all three ages, and mean BMI-Z in Peru was positive. Total food expenditures were not significantly associated with BMI-Z in Ethiopia. In Peru, the only country with a positive mean BMI-Z, there was a significant positive association between total food expenditures and BMI-Z.","We examined data from Young Lives, a cohort of children from Ethiopia, India, Peru and Vietnam, spanning three ages (5, 8 and 12 yrs), to assess whether total food expenditures (dietary quantity), dietary diversity (dietary quality), and household food group expenditure patterns were associated with measures of undernutrition. We found that total food expenditures were positively associated with HAZ (p < 0.05) and BMI-Z (p < 0.05) in all four countries with adjustments for study round, rural residence, and whether the child was female. Similar to other studies (Campbell et al., 2010; Sari et al., 2010; Torlesse et al., 2003; Mauludyani et al., 2014; Rosinger et al., 2013), we found associations between specific food expenditure patterns and nutritional status. For example, Rosinger and colleagues (Rosinger et al., 2013) note that in a nomadic Amazonian population, men (but not women) living in households with high monetary expenditures on market foods (top third) had significantly higher BMI, weight, percentage body fat, and probability of being overweight or obese. However, their study focused on adults, not children. A recent analysis of household food expenditures and inflation in Mozambique during 2008/2009 (Arndt et al., 2016) found higher rates of acute child malnutrition (lower weight-for-age z scores, WAZ), in periods with higher quarterly rates of inflation. When we included dietary diversity and the HFGEI in multivariate HAZ models, total food expenditures were not significant predictors of HAZ in Ethiopia and India. According to these results, how households allocate expenditures across food groups, may be a more important driver of child linear growth than total food expenditures, at least for Ethiopia and India. In multivariate BMI-Z models that included study round, HFGEI, dietary diversity, rural residence, and whether the child was female, total food expenditures remained significantly positively associated with BMI-Z in India, Peru and Vietnam. We found that dietary diversity was inversely associated with BMI-Z in India and Peru. In Peru, where mean BMI-Z is positive, the inverse relationship suggests that higher dietary diversity is associated with a lower risk of higher BMI-Z, and potential overweight. In India, where mean BMI-Z is negative, the inverse association between dietary diversity and BMI-Z is more challenging to interpret. Previous analysis by this team found that dietary diversity was not a significant mediator between food security and child anthropometry in India (Humphries et al., 2015), although another recent study concluded that poor dietary diversity is an important predictor of chronic undernutrition in India (Corsi et al., 2016). The HFGEI was associated with significant increases in child HAZ in Ethiopia, India, and Vietnam with adjustments for data collection round, total food expenditures, dietary diversity, rural residence and whether the child was female. All countries showed an increase in mean dietary diversity from age 5y to age 8y, which is likely an artifact of the addition of one additional food group (vitamin A rich fruits and vegetables) at age 8y, and a decrease from age 8y to age 12y with the same number of food groups in those two rounds. Five food groups loaded on HFGEI in all countries (starches, fruit and vegetables, meat and fish, eggs, and fats), contributing significantly to the common variance across all of the food groups, though relative loadings for each of the food groups were quite different. While a number of researchers (Becquey et al., 2010; Bedard et al., 2015; Devlin et al., 2012) have used factor analysis, PCA, and other approaches to identify groups of foods based on dietary intake data, use of household food expenditure data is less common (Campbell et al., 2010; Sari et al., 2010; Torlesse et al., 2003). In their examination of the relationship between food expenditures and nutritional status, Rosinger and colleagues identified five categories of market foods in a nomadic Amazonian population. These foods are known to reflect the nutrition transition and included dairy, oils, market meats, refined carbohydrates, and sweets. However, Rosinger et al. (Rosinger et al., 2013) did not conduct factor analysis but rather summed total expenditures for each food group. In contrast, Fan and colleagues (Fan et al., 2007) found eight clusters of food expenditures ranging from “balanced” meals eaten largely at home to “fast food” and “full service” meals consumed outside the household.","Our study had several limitations. The intra-household allocation of food is unspecified, so we do not know the relationship between household food expenditures and food consumption of individual children. YL obtained information on child dietary diversity by asking the mother or caregiver at ages 5y and 8y, and by asking the child directly at 12y. In order to fully interpret the relationship between maternal vs. self-reported dietary diversity for children we need to better understand patterns of dietary diversity as children age. This study focuses primarily on current food consumption as a predictor of child growth, although extensive literature (Dangour et al., 2013), including several recent studies, have noted the importance of non-food influences on child height (Griffen, 2016; Krishna et al., 2016; Puentes et al., 2016; Dearden, et al. 2017). As noted previously, the challenges associated with using dietary intake data are numerous and include the importance of research staff with dietary data expertise, substantial interview time, the difficulty of capturing such information, and complex analytic methods. The conversion of food expenditure data into estimates of nutrient intake requires details about food items purchased, local and seasonal costs and regularly updated country-specific tables of food composition. In addition, assumptions regarding intra-household food distribution are required. Use of food expenditure data as an indicator of latent household food group expenditure patterns is simpler, though still requiring assumptions about intra-household distribution, and our study indicates that this measure provides insight into dietary patterns that are associated with child growth. Importantly, food choice patterns may reflect a modifiable component of household behavior, which could be addressed through behavior-change or market price interventions.","In this study, we 1) provide a new way to consider food preferences, relative prices and availabilities by observing household food group expenditure patterns that are relevant for nutritional status, 2) contribute to the literature on how household dietary quantity and quality predict undernutrition, and 3) demonstrate the utility of food group expenditure data as an indicator of potential nutrition risk. We used factor analysis to identify an underlying pattern of consumption across food groups that is related to child growth. Factor analysis and other data reduction techniques may add important information over and above what is provided by total food expenditures alone (as reflected in their significance in our estimates). This is useful because although disaggregated expenditure data may be an important proxy of the types of foods households consume, in the aggregate, their complexity may pose a challenge for policy makers, program planners, and managers seeking to understand associations between food expenditures and children’s nutritional status. This is particularly important because household expenditure surveys are routinely conducted in many low- and middle-income countries. These surveys provide an important but often neglected source of information about how much money households allocate to food, which foods they prioritize, and whether such prioritization provides a diversity of healthy foods. Thus, household expenditure surveys may also shed light on children’s growth. More studies, especially those that are longitudinal in nature, are needed to validate our findings and further explore changes in dietary diversity as children age. We intend to continue examining Young Lives data to explore potential longitudinal effects, sibling effects, and the relationship of HFGEI to other household characteristics such as parental schooling attainment. In summary, our study has shown the potential for using disaggregated household food expenditure data to explore patterns of household food consumption as predictors of child growth. We report significant differences in these patterns across four diverse low- and middle-income countries.","Mary E. Penny has received research funding from the food industry for studies unrelated to this research. The other authors declare no conflict of interest.","This study is based on research funded by the Bill & Melinda Gates Foundation (Global Health Grant OPP1032713), Eunice Shriver Kennedy National Institute of Child Health and Development (Grant R01 HD070993) and Grand Challenges Canada (Grant 0072-03). The study uses data from Young Lives, a 15-year survey investigating the changing nature of childhood poverty in Ethiopia, India (Andhra Pradesh and Telangana), Peru and Vietnam (www.younglives.org.uk). Young Lives is core-funded by the UK Department for International Development (DFID) and was co-funded from 2010 to 2014 by the Netherlands Ministry of Foreign Affairs. The authors are responsible for all the findings and conclusions: they do not necessarily reflect positions or policies of the Bill & Melinda Gates Foundation, the Eunice Shriver Kennedy National Institute of Child Health and Development, Grand Challenges Canada, Young Lives, DFID or other funders.","The University of Oxford Ethics Committee and the Peruvian Instituto de Investigación Nutricional IRB approved YL study protocols. Approval for these analyses was obtained from the University of Pennsylvania. Written parental consent was obtained at the beginning of the study and confirmed verbally at each round. Assent was obtained from children."],["We consider the market for a risky asset with heterogeneous valuations. Private information that agents have about their own valuation is reflected in the equilibrium price. We study the learning externalities that arise in this setting, and in particular their implications for price informativeness and welfare. When private signals are noisy, so that agents rely more on the information conveyed by prices, discouraging information gathering may be Pareto improving. Complementarities in information acquisition can lead to multiple equilibria. --------------------------------------------------------------------------------","We study the market for a risky asset for which agents have interdependent private valuations. Heterogeneous valuations may arise for various reasons. For example, agents may differ with respect to the uses they have for the asset, their liquidity needs, their investment opportunities, or the regulatory constraints they face. Diversity in valuations can be thought of as an indirect way to capture idiosyncratic preference or endowment shocks.1 It can also be interpreted in purely behavioral terms – for example, agents could “agree to disagree” about the distribution of the asset payoff, or a subset of traders could be subject to psychological biases or misperceptions. Each trader is uncertain about his own valuation, and has the opportunity to acquire private information about it prior to trade. Equilibrium prices reflect some of this information. We use a standard competitive rational expectations setup, with Gaussian shocks and constant absolute risk aversion, that nests the classical models of Grossman and Stiglitz (1980) and Hellwig (1980). Essentially the only difference with respect to the classical framework is that we allow agents' valuations to be imperfectly correlated. This gives us a tractable model of partial revelation without resorting to exogenous noise trade, with a unique linear equilibrium price function for any allocation of private information. The model highlights the role played by learning externalities in determining the information content of prices and the welfare of market participants. The welfare analysis is complicated by the fact that price informativeness is a multidimensional object in an economy with heterogenous valuations and, moreover, there is no unambiguous link between price informativeness for a given type and the welfare of that type. Agents can make better portfolio decisions if prices are more informative about their valuation. But more informative prices are also closer to their true valuation, reducing profitable trading opportunities. We find that increasing the cost of information acquisition for agents of the highest cost type leads to a reduction in the proportion of these agents who acquire information, lowering price informativeness for them and improving their welfare. Price informativeness for other types is higher, on the other hand, while the effect on their welfare depends on how precise their private information is. When their private signals are noisy, so that they have more to gain from learning from prices, they are better off. This is the case in which discouraging information acquisition by the highest cost type makes all types better off. Notice that it is precisely when prices have an important role to play in aggregating and transmitting private information that curtailing the collection of private information (by a subset of agents) is Pareto improving. A more general takeaway is that private information collection can impact different groups of agents differently, both in terms of the information conveyed by prices and welfare. Across-type complementarities, wherein information gathering by one type interferes with learning from prices by other types, are an important ingredient of our welfare result. Within-type complementarities play no role here, but are crucial when we consider equilibrium multiplicity. It can turn out that there is an equilibrium in which no agent of type i (for some i) acquires information and another equilibrium in which all of these agents do. In fact, in the equilibrium in which no type i agent is informed, prices are more informative for all types, including type i. Both across-type and within-type complementarities are at play here. Related literature ~~~~~~~~~~~~~~~~~~ Vives (2014) studies a competitive rational expectations equilibrium model with private valuations. As in our paper, there is no equilibrium with a high correlation of types, and when an equilibrium does exist the price function is partially revealing. However, price informativeness does not depend on the mass of informed agents (as long as this mass is positive) – the price reveals the average type of all agents regardless of how many are informed. This in turn implies that the information acquisition decisions of agents are independent. Our stochastic environment shares some features with that of Rostek and Weretka (2012, 2015), insofar as they allow heterogeneity in the correlations between the private valuations of traders. They impose an “equicommonality” assumption, namely that the average correlation between the valuation of a trader and those of the remaining traders is the same for all traders. We do not impose any restriction on the correlation structure for our results on the characterization of equilibrium and price informativeness with exogenous private information (though we do impose symmetry in our analysis of information acquisition). The aims of the Rostek–Weretka papers are different from ours – they study the effect of an exogenous increase in the number of traders on price informativeness (which, in contrast to our setting, is the same for all traders) and on market power. Our framework generalizes the models of Grossman and Stiglitz (1980) and Hellwig (1980), as well as several later extensions of these models. We go beyond this literature in looking at information acquisition by different groups of traders, and analyzing the learning externalities that arise both within and across groups. While the social value of a public signal in a symmetric information economy has been the subject of a voluminous literature going back to Hirshleifer (1971) (see Gottardi and Rahi, 2014 and the references cited therein), not much research has been done on the welfare properties of private information production when prices reflect some of this information. In particular, the literature gives little guidance on the circumstances in which policies that affect private information collection can improve market outcomes in a rational expectations economy. In Vives (2014), information acquisition is socially efficient provided the marginal cost of information is sufficiently low. This efficiency result is not surprising given that there are no learning externalities in this model. Allen (1984) shows that imposing a tax on information gathering in a variant of the Grossman–Stiglitz model can make all agents better off. But the welfare analysis is compromised by the presence of noise traders.2 In fact, most of the rational expectations literature relies on exogenous noise trade and hence does not provide a suitable framework for welfare analysis. Usually a proxy for welfare is employed, such as price informativeness, price volatility or some measure of liquidity. There are a few papers that feature fully optimizing traders but, with the exception of Vives (2014) cited above, they do not address the question of the optimality of the equilibrium allocation of private information. There is a large literature on complementarities in information gathering. The closest to the present paper are competitive models in which these complementarities arise because prices become less informative as more agents acquire information.3 Stein (1987) provides an early example of the entry of informed speculators reducing price informativeness for existing traders, in a setting where agents seek to forecast shocks to the supply of the underlying in a futures market. In an environment closer to that of Grossman and Stiglitz (1980), but with different assumptions on preferences and distributions, Barlevy and Veronesi (2008) find that a complementarity can arise because the asset payoff and noise trader demand are negatively correlated. Their mechanism has a similar flavor to our within-type complementarity which is due to a negative correlation between the valuations of traders from different groups. Price informativeness can be decreasing in the incidence of informed trading in Ganguli and Yang (2009) and Manzano and Vives (2011) because agents have access to two sources of information (about the asset payoff and the asset supply), in Goldstein et al. (2014) because agents with different trading opportunities in segmented markets may trade on the same information in opposite directions, and in Breon-Drish (2012) due to non-normality of shocks. Relative to this literature, our model admits a more pronounced multiplicity of equilibria, including equilibria in which agents who collect information have a higher cost of information acquisition than those who do not, in an otherwise symmetric economy. The paper is organized as follows. We describe the economy in Section 2. In Sections 3–5 we take the information acquisition decisions of agents as given. We characterize the unique linear equilibrium price function in Section 3. Then, in Section 4, we provide several examples in which this characterization can be employed. In Section 5 we analyze the information content of the price for each type. We endogenize information acquisition in Section 6, and discuss learning externalities within and across types. Section 7 is devoted to welfare and Section 8 to equilibrium multiplicity. Section 9 concludes. Most of the proofs are in the appendices.","Optimal portfolios Equilibrium price function","Grossman–Stiglitz Grossman–Stiglitz with optimizing liquidity traders Hellwig Multiple hedgers This example is fully worked out in Appendix B; we provide the salient details here.","Price informativeness","Utilities Utilities of informed vs uninformed Partial revelation An equilibrium vector λ has at least two elements that are strictly positive. Price informativeness vs cost Lower bound on ρ Ranking by price informativeness The incentive to acquire information depends on the value of the common correlation coefficient ρ. Our next result characterizes the values of ρ for which either all agents acquire information or none do. More precisely, in the latter case, all agents have an incentive to free ride on the information gathering of others, and hence there is no equilibrium. Information acquisition: polar cases λN- equilibrium","Welfare The learning externality for the other types goes in the opposite direction. Information acquisition by type N agents interferes with learning from prices by agents of types other than N (and this is true regardless of the value of ρ). This across-type complementarity is responsible for the somewhat counterintuitive result that discouraging information acquisition is Pareto improving precisely when prices have an important role to play in information aggregation.","Multiple equilibria Multiple equilibria II The complementarity result in Goldstein et al. (2014) has a similar flavor to ours: it is driven by a sufficiently strong hedging motive that makes a subset of informed investors trade in the opposite direction to others who only have a speculative motive. Another plausible scenario that can generate negatively correlated valuations, described by Barlevy and Veronesi (2008), is one where some agents have access to a private technology the returns on which are higher in good times, when the asset fundamental is also high. These agents sell the asset in order to free up resources for other projects.","We study competitive rational expectations equilibria in an economy in which agents have interdependent private valuations for the risky asset. For any given allocation of private information, there is a unique linear equilibrium price function that takes a very simple form. We characterize the endogenous distribution of private information when agents can choose whether or not to pay for it. We highlight the role of learning externalities within and across types of agents. When private signals are noisy and agents rely primarily on the information transmitted by prices, raising the cost of information collection for the highest cost type, and thereby curtailing their information gathering activities, can make all types better off. When valuations across types are negatively correlated, multiple equilibria can arise."],["Healthy lifestyle choices and doctor consultations can be substitutes or complements in the health production function. In this paper we consider the relation between the number of doctor consultations and the frequency of patient physical activity. We use a novel application of the Dose-Response Function model proposed by Hirano and Imbens (2004) to deal with treatment endogeneity under the no unmeasured confounding assumption. Our application takes account of unobserved heterogeneity and uses dynamic non-linear models for the treatment and outcome variables of interest. Using seven waves of the British Household Panel Survey, we find that higher treatment intensity and frequency of physical activity are inversely related. We show that accounting for both treatment selection and unobserved heterogeneity halves the size of this relationship. An additional doctor consultation is associated with a 0.5 percentage point reduction in the probability of undertaking vigorous physical activity. Our results hold for a sub-sample visiting the doctor for health check-ups, and are shown to be robust using instrumental variables. --------------------------------------------------------------------------------","Within the World Health Organisation (WHO) European Region, almost 77 percent of the disease burden is due to five major non-communicable diseases (NCD): diabetes, cardiovascular diseases, cancer, chronic respiratory diseases and mental disorders. Amongst its nine global targets to combat these diseases, the WHO has included a reduction of physical inactivity and tobacco consumption, and an increase in treatment and prevention of NCD by primary care doctors (World Health Organization, 2014). There is a wide range of activities that primary care doctors can undertake in treating and preventing NCD, including testing, prescribing and providing lifestyle advice to their patients. A large literature has investigated the determinants of lifestyle behaviours and contacts with primary care doctors (see for example, Manning et al., 1991; Kenkel, 2000; Chaloupka and Warner, 2000; Cawley and Ruhm, 2011; Fernandez-Olano et al., 2006; Morris et al., 2005). Both forms of health investments have common determinants, including socio- economic and demographic factors, preferences, social networks and information. However, little is known about the interaction between these investments. Our aim is to bring together the literature on the determinants of lifestyle behaviours and healthcare utilisation by examining the association between contacts with primary care doctors and healthy lifestyle choices. There is a substantial literature showing that health status is positively affected by the supply of doctors (see for example, Aakvik and Holmảs, 2006; Auster et al., 1969; Gravelle et al., 2008; Or et al., 2005; Robst, 2001; Robst and Graham, 1997). Evidence from the U.S., U.K., Norway and a cross-section of OECD countries shows that increasing the number of doctors per capita decreases mortality rates and improves health-related quality of life. In a Becker-type economic framework, the effect of contacts with doctors on healthy lifestyle choices is ambiguous (Becker, 2007). Individuals invest in their health to equate marginal utility of this investment with its marginal cost. However, there is a trade-off between current costs of healthy lifestyle behaviours (e.g. diverting time and resources away from other activities) and future increased life expectancy. In an application of this model Kaestner et al. (2014) identified two offsetting effects that are applicable to the present study. On the one hand, there is a “competing risk of death effect” as more contacts with doctors might increase the quantity and productivity of health investments which in turn increase life expectancy and the benefit of investments in health. This leads to a positive association between contacts with doctors and healthy lifestyle choices. On the other hand, Kaestner et al. (2014) pointed out that a “technological substitution effect” might occur if healthy lifestyle choices and contacts with doctors are substitutes in the health production function. This leads to a negative association between contacts with doctors and healthy lifestyle choices because more doctor contacts lower the marginal benefit of other health investments. Although the direction of this association could have important implications for policies that aim to increase access to health care professionals, only one paper has explicitly investigated this empirical question. Schneider and Ulrich (2008) used two waves of the German Socio-Economic Panel Study (GSOEP) to examine the relation between a patient’s health-related behaviour and the probability of visiting a doctor. Patients’ health-related behaviours were measured by an indicator that took a value of one if the respondent was smoking and overweight. They used a recursive bivariate probit model with the exclusion restriction that stress directly affects patients’ health-related behaviour and does not directly affect visits to the doctor. As patients who are overweight and smoke were more likely to visit the doctor, they found evidence of substitutability between visits to the doctor and healthy lifestyle choices. Doctors can affect patients’ health behaviours by providing lifestyle advice and treatment. Whilst we would expect healthy lifestyle behaviours and lifestyle advice to be either complements or independent of each other, treatment and health behaviours could be substitutes, complements or independent of each other. The only three papers investigating this relationship focused on different target populations and treatment regimens, and found mixed results. Kaestner et al. (2014) used the Framingham Heart Study spanning between 1983 and 2001 to examine the relationship between the introduction and widespread diffusion of statins and health behaviours. They found evidence that statin use is a substitute for healthy diet with a particularly large increase in female obesity (33% of the mean). They also found evidence of an increase in moderate alcohol drinking of about 15% of the mean and a decrease in sedentary activity among men. Using pooled cross- sectional data from the Health Survey for England, Fichera and Sutton (2011) found that prescription of lipid-lowering drugs complemented quitting smoking behaviour in patients with cardiovascular diseases, but smoking cessation advice was not effective in reducing smoking. Fichera et al. (2014) used a unique linkage between three waves of the English Longitudinal Study of Ageing and practice-level data on the volume of treatments delivered by doctors. They decomposed doctors’ effort into an element induced by the payment system and a discretionary element, using an exogenous change in doctors’ remuneration that led them to increase rates of prescription and disease control. They found that increases in the rates of disease control decreased patients’ cigarette consumption. In this paper we examine the association between the “intensity” of treatment and the level of effort that individuals exert in protecting their own health. We measure treatment intensity as the number of contacts with a primary care doctor and individuals’ health behaviours as the frequency of their physical activity, their smoking and alcohol consumption in seven waves of the British Household Panel Survey. This is a new empirical application of the relation between treatment and healthy lifestyle choices as Kaestner et al. (2014), Fichera and Sutton (2011) and Schneider and Ulrich (2008) did not examine the intensity effect of treatment and Fichera et al. (2014) could only focus on practice-level treatment rates. This is the first methodological application combining the continuous treatment approach with dynamic panel data models. Identification is provided by comparing individuals with different numbers of contacts with the doctor, but the same predicted “intensity” of contacts based on their personal characteristics. The dose-response function uses the GPS to capture the confounders that affect both visits to the doctor and healthy lifestyle choices. It controls for confounding by (complex functions of) observable factors but does not deal with unobserved confounding. We test the robustness of the results to this limitation using fixed effects models and instrumental variables. The rest of the paper is structured as follows. Section 2 describes the data and the summary statistics. Details of our econometric methodology are examined in Section 3. Section 4 discusses the results. Section 5 concludes. The British Household Panel Survey (BHPS) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The BHPS is an annual survey of each adult (16 years of age and older) member of a nationally representative sample of more than 5000 households, making a total of approximately 10,000 individual interviews. In this survey individuals are asked “Since [last 12 months], approximately how many times have you talked to, or visited a GP or family doctor about your own health? Please do not include any visits to a hospital” with the possible answers being: none; one or two times; three to five times; six to ten times; and more than ten times. Individuals are not asked for reasons for their GP visits. In the main analysis, we consider physical activity as the proxy for individuals’ investments in their health. All individuals in the survey are asked about the frequency of their physical activity in one of a succession of questions that ask about things people do in their leisure time. As this question is asked every other year from 1996 to 2008, we select seven of the 18 waves of the BHPS. From the question: “Please […] tell me how frequently you: Play sport or go walking or swimming?” individuals can choose any of the following: “At least once a week; At least once a month; Several times a year; Once a year or less; Never/almost never”. We define physical activity in increasing level of frequency, or effort. We also consider, as supplementary analysis, smoking and alcohol consumption. Smoking is measured as the average number of cigarettes per day and alcohol drinking is a four scale variable (from drinking at least once a week (1) to once a year or less (4)). We consider a number of dummies indicating whether the respondent is white, black, Asian (Indian, Pakistani, Bangladeshi or Chinese) or other ethnic background. A set of dummy variables is included to indicate whether the respondent has obtained a university degree, a high school diploma (the U.K. A-level or O-level), the Higher National Diploma, a semi-professional qualification in the U.K or no qualification at all. We consider the number of children under the age of two years in the household. We dichotomise employment status to indicate whether the respondent is employed (either be employee or self-employed) as opposed to retired, unemployed, on maternity leave or on other employment status. We have taken the natural logarithm of the equivalised value of household monthly income and deflated it by the consumer price index with 1995 as base year. We also considered the number of rooms in the house as an indicator of wealth. A set of dummy variables is included indicating the geographical region of the UK in which the respondent lives. Respondents are asked to identify the physical health problems and disabilities they are currently suffering from a list of 15 physical health conditions. We group these conditions in an homogenous set of eight dummies: musculoskeletal (e.g. arms, legs, feet and back problems); cardiovascular diseases (e.g. heart problems and high blood pressure); diabetes; skin, head or sight problems; respiratory problems; stomach problems; depression and other conditions. In addition to the type of conditions, we construct a series of five dummies indicating the number of health conditions between 0 and two, three, four, five and over six. In order to mitigate reverse causality both the number and type of health conditions are measured at the first wave in which individuals are interviewed. In supplementary analyses we focus on a subsample of individuals who have visited the doctor at least once for preventive purposes. This is intended to alleviate concerns that unobserved health conditions affect both the propensity to engage in physical activity and doctor visits. Unfortunately, there is no longitudinal data in the UK that asks patients for the reasons why they have visited the doctor. The English Longitudinal Study of Ageing, like the BHPS, reports the type of conditions diagnosed by the doctor and whether the respondent has visited the doctor, but does not contain the number of doctor visits or the reason for visiting the doctor. The Health Survey for England, a cross-sectional survey held since 1991, asks for the number of doctor visits but not the reason for visiting the doctor. Therefore, we restrict one of our supplementary analyses to BHPS respondents who have undergone at least one health check in the last year. In a series of questions about which preventive health check-ups respondents have undertaken, we select the National Health Service (NHS) check-ups that are most likely done in a primary care practice. These are blood pressure measurement, cholesterol measurement, cervical screening, breast screening and blood tests. In the supplementary analysis, we restrict the sample to individuals who reported having had at least one of these check-ups in the last year. Not all of the visits in this sample would have been for preventive purposes, but this supplementary analysis is focused on a sub-set of the full sample for which a greater proportion of their visits were for preventive purposes. Nearly 2000 individuals per year (about 94% of the sample) reported having one of these tests, leading to a combined sample of 11,736 observations. For our instrumental variables analyses we generate two instruments: i) the number of times that the individual’s spouse has visited to a doctor; and ii) the average number of consultations with the doctor in the individual’s Local Authority District (LAD) of residence. Summary statistics ~~~~~~~~~~~~~~~~~~ In the main analysis we consider the population aged between 30 and 59 years because their need for medical consultations and their health effort is expected to differ substantially from the older population. Summary statistics are reported in Table 1. More than a quarter of this population group have not been to the doctor in the past year. Approximately 37% of respondents visited the doctor once or twice a year and 19% went to the doctor between three and five times a year. More frequent visits to the doctor are rarer, with almost 9% of people going to the doctor six to ten times a year and about 8% of people going more than ten times a year. Whilst 48% of the sample reported playing sport, walking or swimming at least once a week, 26% of people reported that they never or almost never undertook these forms of physical activity. About 12% of people do these forms of physical activity at least once a month and 9% several times a year. On average this sample has an equivalised household monthly income of £679 and lives in a house containing five rooms. About 78% of the population is either an employee or self-employed. Approximately 16% of people report having at least a university degree and 51% report to have obtained a high school diploma. The population aged between 30 and 59 years is relatively healthy, with 93% having at most two health conditions and only 2% of people reporting having six conditions or more. About 16% of people report having a type of musculoskeletal problem and 17% report skin, head or sight problems.","Our empirical strategy has two main features. Firstly, we predict the propensity score from a (panel) grouped count data model to account for selection of the intensity of treatment, as the number of visits to the doctor depends on individuals’ previous behaviour and socioeconomic characteristics. Secondly, we also account for non-linearities and persistency in the effort that individuals exert on their health investments with panel data ordered probit models. Matching methods have been widely used in the programme evaluation literature of the last two decades (see Augurkzy and Kluve, 2007 for an overview). This is largely due to their ability to mimic experimental settings ex post. As many observational studies involve non-binary treatments, recent literature has extended propensity score methods to the cases of multi-valued treatments (Imbens, 2000; Lechner, 2001), and, more recently, continuous treatments (Behrman et al., 2004; Hirano and Imbens, 2004; Imai and van Dyk, 2004). Hirano and Imbens (2004) apply a generalisation of the binary treatment propensity score, namely the generalised propensity score (GPS), to a population of individuals winning the Megabucks lottery in Massachusetts in the mid-1980. They estimate a dose-response function (DRF) for the amount of lottery prize wins on subsequent labour earnings using the propensity score to adjust for differences in pre- treatment characteristics. In this section, we build on the approach developed by Hirano and Imbens (2004). As in Hirano and Imbens (2004) application, the “intensity” of treatment depends on pre-treatment characteristics. We therefore compare individuals with similar pre-treatment characteristics and similar GPS, i.e. predicted levels of treatment, but different actual treatment levels. Implementation and estimation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The choice of covariates depends on the behavioural factors that affect healthcare utilisation. Education, a proxy for human capital, is related to both health knowledge and self-management (Goldman and Smith, 2002; Cutler and Lleras-Muney, 2006). Income and employment status are proxies for the opportunity cost of time of visiting the doctor. As doctor consultations do not attract user charges in the UK, this is the only cost to the individual. There is also some evidence that ethnicity affects the utilisation of primary care in England (Goddard, 2008). Finally, we consider the number and type of health conditions measured in each individual’s first year of observation. As there might be interactions between these characteristics, we include interactions between types of conditions, age and education. This categorisation of these pre-treatment variables achieves good balance between the treatment and comparison groups. The estimation of the dose-response function (DRF) consists of three stages, which will now be explained in turn. First stage – treatment model and the balancing test Within each cut, we compute the GPS from Eq. (2) at the median of each cut of the treatment. Then, we divide each cut into blocks defined by tertiles of the GPS evaluated at the median, considering only the GPS distribution of individuals in that particular cut of medical consultations. Within each block we calculate the mean difference of each covariate between individuals who belong to a block of the cut and those who belong to other cuts. We combine all the mean differences by using a weighted average with weights given by the number of observations in each tertile of the GPS. This procedure is repeated for each of the cuts and for each pre-treatment characteristic. The key assumption of the first stage is that, conditional on the GPS, there are no statistically significant differences between the characteristics of individuals belonging to different treatment intervals. This does not necessarily imply that there are no differences in their unobserved characteristics. Supplementary analyses ~~~~~~~~~~~~~~~~~~~~~~ We also undertake a range of supplementary analyses. Some of these check the robustness of the results to the model specification; others focus on an alternative sample of individuals who have visited the doctor at least once for preventive purposes and implement alternative econometric techniques. Alternative model specifications We modify the main model specification in four ways. First, we modify the RE ordered probit model to include all the covariates used to estimate the GPS directly in the outcome model. Second, we modify the definition of treatment in the outcome models to include the number of doctor visits: i) treated as count variable; ii) with a set of dummies for each interval of visits; and iii) the prediction from the interval regression rounded to the closest integer. Third, we adopt the stratification method suggested by Imai and van Dyk (2004) by estimating the outcome model separately for each tertile of the GPS and then take the weighted average of the coefficients. Finally, we repeat the analysis described in Section 3.1 including the past frequency of physical activity in the treatment model. In this analysis we lose one year of data for the outcome model. Other specifications provide wider evidence to support the plausibility of the findings. We repeat the analysis for two other health behaviours, smoking and alcohol drinking. We consider the same years used for estimating the physical activity outcome models (i.e. every other year between 1996 and 2008) and focus only on those participating at some level in these behaviours in the first year of observation. We estimate treatment and outcome models for the older population aged 60 and over. Refinement of treatment There is a concern that reverse causation might still bias our results even after controlling for past physical activity, carefully selecting the timing between outcome and treatment, and analysing other health behaviours. We attempt to address this concern by selecting a sub-sample of individuals who we know have visited a doctor for preventive activity by undertaking at least one check-up. There is no other micro-level longitudinal data in the UK that asks patients the reason for visiting the doctor. So whilst this is only an attempt to check the robustness of our estimates for a sub-sample of BHPS respondents who access preventative health services, we acknowledge that they might have visited the doctor for other reasons as well. We have not found any recent aggregate figures for the whole of the UK reporting summary statistics on the reasons why people visit the doctor. However, we found that a 2013 study by the Information Service Division in Scotland reported that amongst the ten activities that attract most of the consultations in primary care practices are blood testing (500 consultations per 1000 population), blood pressure monitoring (350 consultations per 1000 population) and general diagnostic tests (210 consultations per 1000 population). Prescription or medication review account for about 60 consultations per 1000 population. This might suggest that most consultations are for preventive or monitoring purposes. Using Kaestner et al. (2014) theoretical model, for this group of people we can think of visits to the doctor and healthy lifestyle choices as two preventive activities in their health production function. A “technological substitution effect” might prevail if doctor contacts lower the marginal benefit of other health investments. For the sub-sample of people who have at least one check-up we re-estimate a RE ordered probit model including all the covariates in the treatment model. Main analyses ~~~~~~~~~~~~~ We report the results of each of the three stages for estimating the DRF, namely, the treatment model and the balancing test, the outcome model and the DRF plot. Tables 3 and 4 report the balancing tests for each of the three cuts of the treatment, unadjusted and adjusted for the GPS, respectively. We report a more conservative significance value at the one percent level because of multiple comparisons over 70 variables. Table 3 shows a high level of statistical imbalance in most of the pre-treatment covariates for each of the cuts. Imbalance is especially high when considering socio-economic characteristics, and the initial types and number of health conditions and visits. This indicates that BHPS respondents who visit the doctor more or less frequently differ in their observed characteristics. In Table 4 we show that after adjusting for the GPS we obtain a very good balance for all of the pre-treatment characteristics. This indicates that conditional on the GPS, BHPS respondents visiting the doctor more or less often are similar to each other. However, we cannot assert that they have similar unobserved characteristics. As we cannot directly interpret the magnitude of the coefficients in non-linear models, we report marginal effects in Table 6. We compare the marginal effects from Model (III), the static RE model with GPS, to the marginal effects from Model (IV), the dynamic RE model with GPS. We show that the size of the effect estimated by Model (III) is double that estimated by Model (IV) for frequency of physical activity between once a year and once a week. The marginal effects from Model (IV) show that treatment is associated with a shift of the distribution of physical activity to the left. An additional medical consultation is associated with a decrease in the probability of engaging with physical activity at least once a week by 0.5 percentage points while the probability of not doing any physical activity at all increases by 0.4 percentage points. The changes in moderate physical activity are smaller as, on average, an additional medical consultation is associated with an increase in the probability to engage in physical activity between once a year and once a month by almost 0.05 percentage points. Fig. 2 displays the linear predictions of physical activity from the second stage non-linear regression with a range between zero and one. The figure shows a general reduction in the level of physical activity as the intensity of treatment increases. The slope of the DRF is steeper for more than two visits to the doctor. This indicates a negative association between doctor visits and frequency of physical activity that becomes stronger at higher levels of intensity. Supplementary analyses ~~~~~~~~~~~~~~~~~~~~~~ In Table A1 we report estimates when including all the covariates from the GPS regression directly in the outcome equation. The association between physical activity and number of visits to the doctor remains negative and statistically significant and the size of the coefficient is quite close to the one in Model (IV) of Table 5. Including the GPS in the outcome model reduces the curse of dimensionality and improves efficiency because we can just use a term (i.e. the GPS) predicted from a treatment model that includes interactions between covariates and polynomial terms instead of individual covariates. As shown in Fig. 1 doctor visits have been predicted from a constrained model where only the bounds of the intervals map the actual visits. There might be a concern that this prediction is driving our results. In Table A2 we show this is not the case as the same negative association between visits to the doctor and physical activity holds when treatment is defined as a count variable (Model I), as a set of dummy variables (Model II) or as a prediction from the constrained regression rounded to the closest integer (Model III). Model II shows that the association between doctor visits and physical activity is steeper at higher intensity levels as reported in the DRF plot. Table A3 shows that the relationship between treatment intensity and physical activity is similar when we estimate the outcome model separately by GPS tertiles. In Table A4 we report the results of the outcome model where GPS has been obtained from a treatment model that includes physical activity. We show that our previous results were not driven by the omission of past physical activity from the treatment model as the coefficient of interest has a similar size and statistical significance. This should alleviate concerns about reverse causality. We find a statistically significant relationship between smoking or drinking and number of visits to the doctor. Table A5 indicates that more frequent visits to the doctor are associated with more drinking and smoking, a similar association to the one found for physical activity. In Table A6 we report alternative specifications to alleviate concerns of reverse causality and unobserved confounding. All models show that there is still a negative association between doctor visits and physical activity. We report in Table A6 the set of covariates that is shared across all models. Model (I) is a RE ordered probit model estimated on the sample of those who undertake at least one health check. The size of the coefficient is very similar to the one in our preferred specification of a dynamic RE model with the GPS in Table 5. We have also estimated Model (I) using all the specifications in Table 5 and results are very similar (available from the authors on request). Model (II) is estimated with a linear FE model where only time varying covariates have been included. The coefficient on doctor visits is also very similar to our previous specifications. Models (III–IV) report the second stage RE ordered probit coefficients of the 2SRI specification described in Eq. (5). The negative association between number of visits to the doctor and physical activity holds when using either spousal visits to the doctor (Model III) or area average visits (Model IV) as instruments. The similarity of the coefficients on doctor visits using these alternative instruments goes some way to alleviate concerns over their limitations. The coefficient on the residuals predicted from the first stage regression is statistically significant indicating endogeneity of visits to the doctor (Terza et al., 2008). It can be interpreted as evidence that those who have a higher propensity to go to the doctor have a lower propensity to engage in physical activity. As the first stage estimates of these models are similar to those reported in Table 2, we do not report them here, but they are available on request. Both instruments are relevant instruments as they are positively and statistically significantly (at the one percent level) associated with individual’s i visits to the doctor (with a coefficient of 0.02 and 0.14, respectively; and a value of the z statistics greater than 10). In Table A7 we show that the negative association between visits to the doctor and physical activity holds for the older population as well. The magnitude of this relation is higher than the one found for the sample aged between 30 and 59 years.","Healthy lifestyle choices and medical consultations can be substitute or complements in the health production function. Although previous literature (Kaestner et al., 2014; Schneider and Ulrich, 2008) has found evidence of substitutability, medical treatment was measured as a dichotomous variable in these applications. In this paper we have examined the effect of increasing treatment “intensity”, the number of doctor contacts, on frequency of physical activity using seven waves of the BHPS. We have found evidence of a negative association between treatment intensity and physical activity. This relationship is stronger the higher the intensity of treatment. An additional medical consultation is associated with a reduction in the probability of engaging in physical activity at least once a week by 0.5 percentage points while the probability of not doing any physical activity at all increases by 0.4 percentage points. This association is related to a shift of the distribution of physical activity to the left towards lower frequency of engagement. The changes in moderate physical activity are smaller as, on average, an additional medical consultation is associated with an increase in the probability to engage in physical activity between once a year and once a month by about 0.05 percentage points. We have also shown that a simple regression of the number of visits to the doctor on the frequency of physical activity suffers from selection bias and over-estimates the relation between medical consultations and investments in health. We have attempted to mitigate this selection bias problem with a novel application of the dose-response function developed by Hirano and Imbens (2004) that combines the continuous treatment approach with dynamic panel data models. Our novel methodological application has produced three insights in the modelling of the relation between treatment intensity and healthy lifestyle choices. Firstly, we have shown that selection bias accounts for part of the relation between treatment intensity and healthy lifestyle choices as there is a 14% reduction of the coefficient of treatment intensity when the generalised propensity score (GPS) is included in the regression. Secondly, the dose-response function with the GPS could lead to efficiency gains as it allows confounders to enter flexibly in the outcome model via the GPS that can then be stratified and modelled with higher polynomial orders. Our results suggest that accounting for non-linearities in the characteristics determining treatment selection is important as the second-order polynomial of the GPS is statistically significant in the outcome regression. Finally, combining dynamic models and a dose-response function with the GPS has the advantage to flexibly account for treatment selection, unobserved heterogeneity and the dynamic nature of healthy lifestyle choices. We have found the size of selection bias is lower as there is an almost 42% reduction in the coefficient of treatment intensity when we estimate the dose-response function in a dynamic random effects model (i.e. including both the GPS and the lagged values of the outcome variable). One limitation of our paper is that the measure of frequency of physical activity is only limited to playing sports, swimming or walking. Whilst this is the only type of physical activity that is consistently measured across the BHPS sample, we have shown evidence of a negative association between treatment intensity and other healthy lifestyle behaviours such as reducing cigarettes and alcohol consumption. A second limitation which we share with the study by Hirano and Imbens (2004) is the lack of exogenous variation in treatment. The application by Hirano and Imbens (2004) focused on a cross-section of lottery winnings which although exogenous belong to a particular selected sample of players. Although we combine dynamic panel data models with the GPS, we are cautious in asserting we are estimating a causal effect. We have attempted to mitigate this limitation by using 2SRI models with spousal and area-average visits to the doctor as instruments. Both instruments have advantages and disadvantages relating to the amount of variation and the potential for direct pathways to physical activity. Under the assumption of time-invariant unobserved heterogeneity we have estimated a FE model. Although our results are robust to both specifications, we note that time-varying unobserved heterogeneity and omitted variable bias might still be possible. A third limitation is that, in following Bia and Mattei (2008), Hirano and Imbens (2004) and Imbens (2000), we do not correct the standard errors for the inclusion of the GPS in the outcome model. A final limitation is that we cannot determine what elements of the treatment generate an inverse relationship with healthy lifestyle choices, as the dataset contains no information on the cause and content of doctor consultations. The non-linear association between doctor visits and frequency of physical activity might be concerning if an unobserved (to us as researchers) health problem has induced patients to initiate a doctor visit. This would generate a non-linear and reverse causal association between doctor visits and frequency of physical activity. There is no longitudinal survey data in the UK that asks respondents the reason for visiting the doctor. Instead, we have shown that our results are robust to restricting our sample to individuals who have had at least one preventative health check-up. Official statistics suggest that the majority of doctor consultations are for preventive purposes. These two pieces of information point to the direction of a substitution between two preventive activities in the health production function. However, we highlight that data do not allow us to ascertain the reason for visiting the doctor and therefore we cannot give a causal interpretation to the estimates produced in this study."],["The purpose of this study is to explore the main correlates of male height in 105 countries in Europe & overseas, Asia, North Africa and Oceania. Actual data on male height are compared with the average consumption of 28 protein sources (FAOSTAT, 1993-2009) and seven socioeconomic indicators (according to the World Bank, the CIA World Factbook and the United Nations). This comparison identified three fundamental types of diets based on rice, wheat and milk, respectively. The consumption of rice dominates in tropical Asia, where it is accompanied by a very low total protein and energy intake, and one of the shortest statures in the world (∼162-168 cm). Wheat prevails in Muslim countries in North Africa and the Near East, which is where we also observe the highest plant protein consumption in the world and moderately tall statures that do not exceed 174 cm. In taller nations, the intake of protein and energy no longer fundamentally rises, but the consumption of plant proteins markedly decreases at the expense of animal proteins, especially those from dairy. Their highest consumption rates can be found in Northern and Central Europe, with the global peak of male height in the Netherlands (184 cm). In general, when only the complete data from 72 countries were considered, the consumption of protein from the five most correlated foods (r = 0.85) and the human development index (r = 0.84) are most strongly associated with tall statures. A notable finding is the low consumption of the most correlated proteins in Muslim oil superpowers and highly developed countries of East Asia, which could explain their lagging behind Europe in terms of physical stature. --------------------------------------------------------------------------------","In our previous study (Grasgruber et al., 2014), we identified nutrition and genetics as the strongest correlates of height among contemporary young men from 42 European and three overseas countries (Australia, New Zealand and the USA). Out of all the socioeconomic factors that were examined, only children's mortality approached the significance of nutrition and genetics, which points to the importance of a disease-free environment. Improved nutrition and better healthcare are direct consequences of improving living standards that accompanied the process of the industrial revolution (Hatton, 2013). In the present study, we aim to extend this research to North Africa, Asia and Oceania. These regions include mostly developing countries, but Muslim oil superpowers (Bahrain, Brunei, Kuwait, Oman, Qatar, Saudi Arabia and the United Arab Emirates/UAE) and some developed countries of East Asia (Japan, Singapore, South Korea, Taiwan) currently belong to the wealthiest in the world, with the gross domestic product (GDP) per capita higher than 30,000 USD (World Bank, 2013), not to mention the semi-independent territory of Hong Kong. Interestingly, male height in some of these regions (Arab countries of North Africa and the Near East) was once similar or even higher than in Europe, but after the 1880s it started to lag behind considerably (Stegl and Baten, 2009). On the other hand, wealthy nations of East Asia are still known for their surprisingly small stature, despite very high values of the GDP per capita (Baten and Blum, 2014). Therefore, it would be important to identify the main factors that currently distinguish these regions from Europe. As we did similarly in our previous study, we plan to explore the correlation of factors such as nutrition, healthcare and national wealth with the height of contemporary young men. Genetic factors (frequencies of Y haplogroups) are included as well, but in a supplementary function, because no Y haplogroup is shared in appreciable frequencies across the whole area of Europe, North Africa, Asia and Oceania. Collection of anthropometric data ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The selected regions of North Africa, Asia and Oceania encompass 63 countries (including three semi-independent territories–American Samoa, French Polynesia, Hong Kong), for which data on body height were researched. In Oceania, only more populous countries with a population size exceeding 50,000 inhabitants were considered, because statistics were not available for small island nations. Preferably, we searched for anthropometric data on young, mature men aged 18–30 years (but ideally 20–25 years) from the time after the year 2000. Only surveys that incorporated at least 50 individuals were used, but whenever possible, nationwide surveys with more than 200 individuals were preferred. The only samples with fewer than 100 individuals were from Guam (n = 59) and Algeria (n = 55)1. By far the most representative data (regular measurements of recruits) were available from Israel, but paradoxically, they were the most problematic due to the inclusion of young immigrants, who were not born and raised in Israel2. The majority of surveys included in our study are nationwide health surveys incorporating all social groups. Some other surveys had certain limitations. Six of them (from Bhutan, Laos, Libya, Oman, the Maldives and Syria) incorporate an urban population from a single city. The sample from Libya consists of patients visiting a hospital in Derna. Three surveys (from Saudi Arabia, the United Arab Emirates and the Federative States of Micronesia) come from a specific region of the country. The sample from North Korea includes refugees. In the case of Afghanistan, Kyrgyzstan, Pakistan, Tajikistan and Turkmenistan, the male height was estimated, based on highly representative studies of local women. The height in Kazakhstan (175.6 cm) was computed from the data of Facchini et al. (2007). (See Appendix: Methods for a more detailed discussion.) Due to the scarcity of information from some countries, certain compromises needed to be made. This especially concerns the STEPS surveys (Noncommunicable Disease Risk Factor Surveys) performed by the World Health Organization (WHO), which rarely include people younger than 25 years and routinely start with the age category of 25–34 years. Nevertheless, it is unlikely that the inclusion of older subjects would markedly distort our results, because the pace of the secular trend in the majority of the examined countries is very slow or almost non-existent. Besides that, some means of male height from our previous study were updated. This update relates to Bosnia and Herzegovina, Bulgaria, the Czech Republic, Georgia, Moldova and Ukraine (see Appendix: Methods and Appendix Table 1). Recent anthropometric data were even obtained from Armenia, but they did not seem to be sufficiently representative. Nevertheless, they indirectly supported the accuracy of our male estimate based on DHS 2005 (171.9 cm)3. Altogether, information on body height was collected from 61 out of 63 targeted countries (Table 1). Only data from Macau and New Caledonia were missing4. Our list also includes the Maori, but considering that they make up a minority in New Zealand, this sample was not useable. Collection of nutritional and sociodemographic data ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The information about the average daily protein consumption (in grams) was computed from FAOSTAT.org5. The statistics on the gross domestic product (GDP) per capita (by purchasing power parity/PPP, in current international USD), health expenditure per capita (by PPP, in constant 2005 international USD), urbanization (% of urban population), children's mortality under 5 years (per 1000 live births) and total fertility rate (births per woman) were taken from the World Bank6, and the Gini index of social inequality from the CIA World Factbook7. In addition, we included the human development index (HDI) that is regularly calculated by the United Nations8. Since 2010, the HDI has been computed from the statistics of life expectancy, mean and expected years of schooling, and the GDP per capita by PPP. To obtain the most precise results, we extended the period, for which the average values of nutritional and sociodemographic data were calculated (from 2000 up to 1993), but this required the elimination of Montenegro from the European sample, because many key data from Montenegro would be missing. A further historical extension (before 1993) was not possible, because many statistics from the disintegrated countries of the former Communist bloc have been available only since 1993. The information on the GDP and health expenditure has been complete since 1995. The average daily protein consumption, children's mortality, total fertility and urbanization were therefore computed for the period 1993–2009 and the GDP and health expenditure for 1995–2009. The data of the Gini index were computed as averages of all the available values from 1990 to 2013. Considering that the historical HDI data were often incomplete, only the HDI values from 2011 were used. This limitation should not pose a problem, because HDI values largely summarize the societal development in the near past. Furthermore, the use of the HDI from 2011 enabled a mutual comparison with the newly introduced ‘inequality-adjusted HDI’ (IHDI), which takes the demographic variance in the HDI indicators into account. Altogether, information was collected on male height from 106 countries (including Montenegro), on urbanization from 105 countries, on total fertility from 102 countries, on children's mortality from 101 countries, on the HDI from 100 countries, on health expenditure from 97 countries, on the GDP from 96 countries, on protein consumption from 93 countries, but on the Gini index only from 80 countries. All the data were available for 72 countries. Collection of genetic data ~~~~~~~~~~~~~~~~~~~~~~~~~~ For the genetic comparison, we use frequencies of Y haplogroups (male lineages)9. Y haplogroups usually show a more refined geographical pattern than mtDNA haplogroups (female lineages), because males rarely leave their ‘clan’ (Rosser et al., 2000). This is also evident from the fact that the distribution of Y haplogroups correlates with certain language groups10. Therefore, it is more likely that Y haplogroups will be associated with the prevalence of certain physical traits, because they mostly evolved within separate patrilocal communities. As shown by the results of our previous study, typical European Y haplogroups correlate quite strongly both with male height and with the prevalence of lactose tolerance. The relationship between the combined frequency of Y haplogroups I-M170 & R1b-U106 and male height in 34 countries reached r = 0.75 (p < 0.001). It is true that a recent study by Robinson et al. (2015) attributed only 24% of the height differences among 14 European countries to genetic differences, but this percentage may be underestimated, because this study apparently used very low height values of the whole male population and it did not include any samples from the Western Balkans. In the present study, two major regions with a relatively high frequency of particular Y haplogroups are examined. In North Africa and the Near East (including Turkey and the Caucasian republics), six major Y haplogroups were selected: E1b1b1a1-M78 (E1b-M78), E1b1b1b1a-M81 (E1b-M81), G-M201, J1-M267, J2-M172 and R1a-M420. Data on the frequency of these Y haplogroups were available from 21 countries (Appendix Table 2a). Only representative information on Bahrain and Israel was missing. In South, Southeast & East Asia and Oceania, four Y haplogroups were selected: O1-MSY2.2, O2-P31, O2a-PK4 and O3-M122 (Appendix Table 2b). Data from this region are still quite scarce and often confined to specific areas or tribal groups; hence information from 8 out of 35 targeted countries was missing. Data on the phenotypic prevalence of lactose tolerance are taken from three main sources: the Global Lactase Persistence Association Database, Ingram et al. (2009) and Flatz and Rotthauwe (1977). Only studies incorporating ≥50 individuals were used, which limited the number of countries to 20 (see Appendix Table 3). When combined with the countries examined in our previous study, the prevalence of lactose tolerance was available for 47 countries. Statistical analyses ~~~~~~~~~~~~~~~~~~~~ Statistical analyses were conducted by the software Statistica 12. A standard comparison via Pearson linear correlations was performed between male height and all the data that were available (Tables 2 and 3a–3b). Because complete information was not available for all the countries, the number of countries differed from variable to variable. To make the results comparable, we undertook an additional, separate analysis with only 72 countries, for which all the information was available (Table 3b and Appendix Table 4). These 72 countries were also subsequently used in a multiple regression. The drawback of this limited sample was the complete absence of Oceania and many Muslim countries. Therefore, an additional regression analysis with 83 countries (without the Gini index) was performed, but as we will show below, there were only small differences in the results. More sophisticated statistical procedures such as the use of instrumental variables were also considered, but we failed to find a variable that would be useful for such models. Distribution of male height ~~~~~~~~~~~~~~~~~~~~~~~~~~~ The data collected from 61 countries/regions are summarized in Table 1 and the geographical comparison of all 106 countries is displayed in Fig. 1. The range of values is very wide, from 161.6 cm in Timor-Leste up to 183.8 cm in the Netherlands (22.2 cm). Very tall statures above 180 cm are typical only of Europe, especially the areas with the highest frequency of Y-haplogroups I-M170, R1b-U106 and R1a-M42011, and the highest quality of nutrition. The lowest values (below 165 cm) can be found in Yemen and Southeast Asia (Timor-Leste, Cambodia, Bangladesh etc.). Highly developed countries in East Asia (Japan, Singapore, South Korea, Taiwan), China and Muslim countries in North Africa and the Near East are positioned roughly in the middle, with heights fluctuating around 170 cm. The Lebanese (175.5 cm) and most likely even the people of Kazakhstan (ca. 175.6 cm) are by far the tallest in Asia. Somewhat surprisingly, the average height in highly- developed Israel (174.5 cm) is shorter than that in Lebanon, and this is true even after the exclusion of 19-year old recruits, who were sons of recent immigrants and were not born in Israel. The small sample of Algerians (174.6 cm) is the tallest in North Africa. However, their stature is still only comparable to that of the shortest nations of Europe. Remarkably, it is Polynesians, who reach the highest values among the 61 new samples, more specifically the inhabitants of French Polynesia (178.6 cm)12, Tonga (176.7 cm)13 and American Samoa (175.9 cm)14. Marked geographical changes in male stature are hidden in the national averages of the largest countries—China and India. Zhang and Wang (2011) recently analysed data from the nationwide health survey in China (2005) and found that the height in 30 Chinese provinces (except Tibet) ranged from 165.6 cm for men in rural areas of the Guizhou province to 175.4 cm for men in urban areas of the Liaoning province. If the data from urban and rural areas from this study are pooled, the tallest statures can generally be found in northeastern provinces around the Yellow Sea such as Beijing (174.5 cm), Liaoning and Shandong (both 174.2 cm), while small statures are typical of southern regions, particularly the provinces of Chongqing (167.1 cm) and Guizhou (166.5 cm) (see Appendix Fig. 1). The geographical changes of male stature in India are comparably large. According to the most recent Indian nationwide survey (NFHS 2005–2006; Mamidi et al., 2011), the highest male averages in the age category 20–29 years were documented in northwestern states—Punjab (168.4 cm), Haryana (168.1 cm) and Jammu and Kashmir (168.0 cm). The shortest statures are typical of northeastern states—Meghalaya (157.5 cm), Sikkim (160.0 cm) and Arunachal Pradesh (161.0 cm) (Appendix Fig. 2)15. GDP per capita ~~~~~~~~~~~~~~ The correlations of male height with the GDP per capita and other socioeconomic indicators are summarized in Table 2. Although the growing GDP per capita has been the fundamental trigger of the positive height trend during the last century, it fails to be a good correlate of male height in 96 countries (r = 0.30; p = 0.003) (Fig. 2). This is primarily due to two reasons: First, as we showed in our previous study, the relationship between height and GDP in Europe was distorted by the historical division into the “Western” and “Communist” blocs. Communist countries lagged behind in terms of economic development, but historically, the wealthiest of them were characterized by a diet associated with tall statures, which is based on milk products, pork and fish. In contrast, the dietary customs in some “Western” countries such as Portugal, Spain, Italy and Greece were different and only began to change in a similar way during the most recent decades. Second, the high GDP per capita in Muslim oil superpowers, as well as in the wealthy countries of East Asia, is not accompanied by the same type of nutrition or public expenses on healthcare that we find in the tallest nations. This will be the subject of the upcoming paragraphs in this section. Health expenditure per capita ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Health expenditure turns out to be a much better correlate of male height than the GDP (r = 0.60; p < 0.001 in 97 countries) (Appendix Fig. 3). This is mainly due to the fact that the public health expenses in Muslim oil superpowers are much lower than their GDP would predict and are more in line with the unimpressive height of the local young men. Such a striking discrepancy points to a strongly uneven redistribution of yields from the oil industry. Children's mortality under five years ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The most significant socioeconomic factor out of the six that we examined in our previous study is again children's mortality (r = −0.73; p < 0.001 in 101 countries) (Fig. 3). This result remains valid even in the sample of 72 countries, for which all the data were available (r = −0.78; p < 0.001) (Table 2). The range of values is enormous, from ∼4 in Finland, Iceland, Singapore and Sweden up to 132 deaths/1000 live births in Afghanistan. Interestingly, the low children's mortality rate in tropical Asia, East Asia and Muslim oil superpowers has no influence on the stature of local males. Total fertility ~~~~~~~~~~~~~~~ Total fertility (r = −0.64; p < 0.001 in 102 countries) follows children's mortality as the second strongest socioeconomic correlate (Fig. 4). Although total fertility plays only a marginal role in Europe (r = −0.26; p = 0.09), where fertility rates are almost unanimously very low, it reaches statistical significance in the global context, because the fertility level in developing countries is still high. It is logical to assume that the lower the number of children per family, the higher the financial expenditures per child would be, and hence the living conditions of the children would improve. Again, it is noteworthy that this assumption does not apply so well to tropical and East Asia, where the birth rates are similarly low or even lower than those in Europe, but the difference in body height is still 10–20 cm. On the other hand, the birth rates in Muslim oil superpowers are unusually high for the standards seen in developed countries, reaching almost 4.0 in Oman and Saudi Arabia, and exceeding other Arab countries with a substantially lower GDP. This could serve as further indirect evidence that their high GDP does not translate into adequately high living standards for most of the population and/or the degree of societal development. Urbanization (% of urban population) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The correlation between urbanization and height is almost as strong as in the case of total fertility (r = 0.58; p < 0.001 in 105 countries) (Appendix Fig. 4). In general, the highest rates of urbanization are typical of Europe, Muslim countries and the highly developed countries of East Asia, but they have a much weaker correlation with height in the latter two. Gini index ~~~~~~~~~~ The Gini index, as the indicator of social inequality, does not show any statistical relationship with height in non-European countries (r = −0.04; p = 0.80), because the levels of social inequality are universally very high. This factor reaches significance only in Europe (r = −0.36; p = 0.017) and in the total sample of 80 countries (r = −0.51; p < 0.001) (Appendix Fig. 5). The Gini index tends to decrease with a growing GDP per capita (r = −0.29; p = 0.010), and countries in tropical Asia and Central Asia are the most affected by the combination of extreme poverty and deep social inequality (Appendix Fig. 6). It is symptomatic that data on the Gini index are not available from wealthy oil superpowers such as Saudi Arabia, Kuwait or Bahrain. Human development index (HDI) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ This comparison proved to be a very interesting addition to our analysis, because it shows one of the highest correlations with male height (r = 0.80; p < 0.001 in 100 countries) (Fig. 5). Height and the HDI seem to be largely interchangeable as indicators of human well-being. They are both related to the GDP per capita. High life expectancy is a factor that reflects the quality of healthcare and nutrition. Furthermore, better education of women translates into a lower fertility rate (Basu, 2002). The relationship between male height and the IHDI (inequality-adjusted HDI) (r = 0.87; p < 0.001 in 72 countries) (Appendix Fig. 7) is even stronger than in the case of the HDI, but if we consider only this sample of 72 countries, for which both these indicators are available, we find no difference in r-values. The difference between the HDI and the IHDI (a benchmark for social equality) is a good correlate of male height as well (r = −0.64; p < 0.001) (Appendix Fig. 8). Nutrition ~~~~~~~~~ Because our present study incorporated FAOSTAT data from a longer period of time (1993–2009), it is not surprising that the correlation coefficients in our European sample of 44 countries are often higher (Table 3a). Three additional protein sources (meat total, pelagic marine fish, eggs) reach significance and the position of milk products as key stimulators of physical growth is further strengthened (r = 0.50; p < 0.001). As already mentioned in our previous study, these findings have a very solid rationale, because milk products, red meat, eggs and some species of fish have the highest protein quality out of all common foods. The situation in 49 non-European countries is very different. The main correlate of height is not protein quality, but total protein consumption (protein quantity) (r = 0.71; p < 0.001). Fig. 6 shows that total protein consumption in tropical Asia and some other countries such as Tajikistan, Iraq, Afghanistan or the Solomon Islands is extremely low. The main socioeconomic factor boosting total protein intake in 72 countries is the HDI (r = 0.81; p < 0.001), followed by urbanization (r = 0.78; p < 0.001) and the GDP per capita (r = 0.76; p < 0.001). When these variables were assessed individually and the number of countries examined was thus greater, the association between total protein and GDP per capita in 85 countries decreased (r = 0.65; p < 0.001), because Muslim oil superpowers such as the UAE, Kuwait and Brunei emerged as striking outliers that consume less protein than their high GDP would predict (Appendix Fig. 10a). With the increasing consumption of total protein, there are large differences in height at the same consumption level (10+ cm), which points to the importance of protein quality. Other highly significant nutritional correlates of male stature in 49 non-European countries are total energy intake (r = 0.70; p < 0.001) (Appendix Fig. 9), rice protein (r = −0.65; p < 0.001) and wheat protein (r = 0.62; p < 0.001) (Fig. 7). Rice is the main source of protein in tropical Asia (particularly in Laos, Cambodia and Bangladesh; >30 g/day), which is where we encounter extremely small height means of ∼162–168 cm. The level of rice protein consumption is quite high even in Japan, China and both North and South Korea (>10 g/day). Remarkably, wheat protein correlates with male stature negatively in Europe (r = −0.68; p < 0.001), but positively outside Europe. When rice consumption decreases, wheat consumption increases (Appendix Fig. 11a) and so do the values of male stature (Fig. 7 and Appendix Fig. 11b). The intake of total energy and total protein increases as well. The positive height tendencies tied to the increasing consumption rates of wheat reach a peak of ∼174 cm in Muslim countries of North Africa and the Near East (Turkey, Tunisia, Azerbaijan). The consumption of vegetables and legumes in this region is very high as well and the proportion of animal proteins is moderate. As a result, some of these countries (Turkey, Egypt, Morocco) consume the largest amounts of plant protein in the world (Appendix Fig. 12). In taller nations, the intake of total protein and total energy is no longer rising fundamentally, but plant proteins are substituted by animal proteins (particularly dairy proteins) (Appendix Fig. 13). Consequently, the relationship between most plant proteins and male height is curvilinear (Appendix Fig. 12), similar to that between GDP per capita and plant protein (Appendix Fig. 10b). The animal protein intake in 72 countries rises very strongly with the HDI (r = 0.89; p < 0.001) and GDP per capita (r = 0.87; p < 0.001). Muslim oil superpowers consume much less animal protein than their high GDP per capita predicts, which again decreases the r-values, when 85 countries are considered (r = 0.67; p < 0.001) (Appendix Fig. 10c). The intake of protein from milk products (dairy proteins) emerges as the most significant nutritional correlate of stature not only in Europe, but in all 93 countries examined in this study (r = 0.79; p < 0.001) (Table 3b; Appendix Fig. 14), followed by total protein (r = 0.74; p < 0.001) and animal protein (r = 0.73; p < 0.001). The most negative nutritional correlate in the total sample is again rice (r = −0.74; p < 0.001). Dairy proteins are most frequently consumed in the Netherlands, Sweden and Finland (almost 30 g/day), while in countries from the eastern half of Asia such as Cambodia, Laos or North Korea, their intake is virtually zero. Remarkably, this low intake of milk products is in accordance with the low prevalence of lactose tolerance in Southeast and East Asia (Appendix Fig. 15a) and the widespread undernutrition in this region. These facts confirm the fundamental evolutionary significance of lactose tolerance in a world, where the scarcity of valuable nutrients is an everyday reality16. Similarly to our previous study, we tried to find ratios or combinations of protein intake that would further improve the predictive power. The ratios between animal proteins and plant/cereal protein were not useful in this regard, but combinations of proteins were. Pork protein has the biggest additive effect, when combined with dairy proteins (r = 0.82; p < 0.001), followed by protein from eggs (r = 0.81) and potatoes (r = 0.81). The correlation reaches r = 0.84, when potatoes are added to dairy and pork. The highest r-value (r = 0.85) was achieved via the combination of proteins from dairy, pork, beef, eggs and potatoes (‘highly correlated proteins’). The strength of this relationship is visually impressive (Fig. 8). ‘Highly correlated proteins’ are also the strongest correlate of male height (r = 0.85), when all the nutritional and socioeconomic variables from 72 countries are considered (Table 3b). Interestingly, in this sample of 72 countries, ‘highly correlated proteins’ are strongly associated with the HDI (r = 0.82) and GDP per capita (r = 0.77; p < 0.001), but the addition of wealthy Arab countries again decreases the relationship with the GDP per capita very noticeably (r = 0.49; p < 0.001) (Appendix Fig. 10d). While Europeans often consume more of these proteins than their GDP per capita would predict, wealthy, short-statured Muslims and East Asians consume less. ‘Highly correlated proteins’ are also strongly associated with lactose tolerance (r = 0.80; p < 0.001), although in this case, the number of countries was limited to only 44. The relationships of nutrition with various socioeconomic factors and lactose tolerance are presented in Appendix Tables 5a–5e. These results confirm that red meat and eggs are the most height-related components of the human diet after milk, which primarily stems from the complete amino acid spectrum of their proteins (Appendix Table 6). Our calculations show that at least in Europe and overseas, the correlation of eight main food items with height accords better with their amino acid scores based on the older FAO standard 1985 than the newer FAO standard 2007 (see Appendix Tables 7a and 7b). More concretely, the new FAO standard 2007 produces relatively higher amino acid scores for beef and fish, and a relatively lower amino acid score for pork. This is mainly due to the marked decrease of tryptophan content in the new standard17. The role of pork in our analysis is weakened because of its absence from the diet of Muslim countries (Appendix Fig. 17), but its statistical relationship with stature is still stronger than that of beef and fish, when all 93 countries are considered. Therefore, we think that in the light of the admitted imperfection of the current amino acid scores18, the results of our ecological study should be taken into account seriously. On the other hand, the significance of potatoes, even as an independent item (r = 0.68; p < 0.001), is unexpected because of the poor quality of potato proteins, their low consumption rate and a very low ‘nutrient density’19. Therefore, we cannot exclude that the correlation between potatoes and height is only spurious. Nevertheless, it is visually quite persuasive (Appendix Fig. 18) and the rationale behind this finding deserves further discussion20. In contrast with Europe, fish consumption does not contribute positively to male stature in the total sample of 93 countries (r = −0.15; p = 0.16) and freshwater fish even correlate slightly negatively (r = −0.21; p = 0.047). This result is largely deceptive, because fish consumption shows some positive relationship with male height, when the examined countries are divided according to regions (Appendix Fig. 19). In Southeast Asia, Japan and many countries in Oceania, fish remain the main source of animal protein. Other protein items with a high height correlation are not consumed in large quantities. Furthermore, there are big differences in protein quality among various species of fish. When the food items with the most negative r-values (rice and legumes) are combined, a small additive effect is observed, when compared with rice alone (r = −0.75; p < 0.001) (Appendix Fig. 20). The proportion of energy intake from protein (assuming 4.1 kcal per 1 gram of protein) clearly highlights the low ‘nutrient density’ of the diet in tropical Asia (Appendix Fig. 21). Current nutritional trends in Asia, North Africa and Oceania ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The evidence presented in the previous paragraphs can persuasively explain, why many wealthy Asian countries have a smaller physical stature than European countries. A high GDP, high rates of urbanization, low children's mortality rate, high indices of human development and mostly above-average health expenditure and below-average fertility rates are not sufficient to compensate for the persisting low intake of the proteins that are most strongly associated with tall stature in the present study. In fact, the intake of ‘highly correlated proteins’ in Kuwait, the UAE, Japan and South Korea is on the level of the poorest European countries such as Georgia and Moldova (Fig. 8). Furthermore, the level of social inequality is apparently higher than it is in the European nations with the tallest people. We can also justly suspect that the data on the Gini index from rich oil superpowers are not published for a good reason, because the disparity in the distribution of national wealth may have no parallel to the rest of the world. Therefore, the consumption of high-quality animal proteins could theoretically predict future development. A particularly fast increase in stature can be expected in Kazakhstan, Myanmar and Vietnam, where the annual animal protein intake during the decade 2001–2011 rose by >15 g (Fig. 9). A fast upward rate (>8 g/decade) can even be observed in China, Morocco, Turkmenistan and South Korea. In contrast, the amount of animal proteins consumed has more or less declined in Afghanistan, Japan, Lebanon, Mongolia, Palestine, the Solomon Islands and the UAE. An extremely low level of consumption (∼10 g/day) combined with a stagnating or only negligibly growing trend line, is typical of Bangladesh, India, Iraq, Laos, Nepal, North Korea, Tajikistan and Yemen. Genetics: North Africa and the Near East ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As already noted in the introduction, testing the relationship between height and genetic markers plays only a supplementary role in the present analysis, because the distribution of certain Y haplogroups is geographically limited. In 21 countries of the targeted regions of North Africa and the Near East, the only haplogroup that shows a significant, negative relationship with male height is J1-M267 (r = −0.68; p < 0.001) (Fig. 10a). J1-M267 is a signature of human populations that expanded from the Zagros Mountains during the Holocene and its dominant sublineage J1a2b-P58 (formerly J1e) has been connected with the spread of pastoral nomadism in the Arabian Peninsula (Chiaroni et al., 2010). Today, the frequencies of J1-M267 peak in Yemen (73%), followed by other countries of the Arabian Peninsula (∼35–60%) and speakers of Northeast Caucasian languages (>50%). The role of J1-M267 remains robust (p < 0.01) even after controlling for protein consumption and all the socioeconomic variables, except total fertility (p = 0.14). Among the four remaining haplogroups, G-M201 approaches significance as a correlate of taller statures (r = 0.41; p = 0.07) (Fig. 10b). This lineage reaches the highest frequencies in Georgia (39%) and is much rarer elsewhere (18% in Azerbaijan, 12% in Iran, 11% in Turkey). The positive effect of G-M201 becomes statistically significant, when North Africa is excluded (r = 0.55; p = 0.028). The Anatolian Neolithic lineage J2-M172 (r = 0.21; p = 0.36), as well as the East African haplogroup E1b-M78 (r = 0.21; p = 0.36) and the Northwest African (Berber) haplogroup E1b-M81 (r = 0.20; p = 0.38) have a non-significant, but still somewhat positive association with height. Interestingly, the combination of G-M201 and E-M78 correlates significantly positively with height (r = 0.48; p = 0.027) and three presumably ‘agricultural’ haplogroups G-M201, E1b-M78 and J2-M172 approach statistical significance (r = 0.40; p = 0.07) (Fig. 10b)21. When E-M81 is added to these three haplogroups, the r-value markedly rises to r = 0.57 (p = 0.008). The addition of the Indo-Iranian genetic signature R1a-M420 further increases the correlation coefficient to r = 0.62 (p = 0.003) (Fig. 10c), although R1a-M420 by itself shows no relationship with height (r = 0.11; p = 0.63), which must primarily be ascribed to its low frequencies across the examined regions. The significance (p < 0.05) of these five combined Y haplogroups persists after controlling for all the socioeconomic variables and nutrition, but again, total fertility decreases it the most relatively (p = 0.04). These results indicate that J1-M267 is the major genetic correlate of short stature in the Near East and North Africa, whereas G-M201, E1b-M78, E1b-M81, J2-M172 and R1a-M420 appear to have a positive effect, although it manifests significantly only when the frequencies of these five haplogroups are combined. According to these data, the greatest potential for height could be expected in Moroccans, Tunisians and Georgians, and the lowest in Yemenites. The only other factor that markedly decreases the significance of these results is total fertility. The independent evolution of lactose tolerance in Europe and Arabia ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The polarity between the ‘pastoral’ haplogroup J1-M267 and the ‘agricultural’ haplogroups G-M201, J2-M172 and E1b-M78 does not seem to be limited to body height. J1-M267 namely lies at the epicentre of another global peak of lactose tolerance, as evidenced by relatively high prevalence values from Jordan (51%), Kuwait (53%) and Saudi Arabia (56%). Indeed, lactose tolerance has a high positive correlation with J1-M267 (r = 0.87; p = 0.023) (Appendix Fig. 22a), but a high negative correlation with E1b-M78, G-M201 and J2-M172 (r = −0.89; p = 0.019) (Appendix Fig. 22b). Although these results are based on only six countries, they are definitely intriguing, because the evolution of lactose tolerance within pastoral communities of J1-M267 makes good sense. In addition, in our previous paper we found that the combined frequencies of E1b-M78, G-M201 and J-P209 correlated strongly negatively with lactose tolerance even in Europe (r = −0.73; p < 0.001). This finding could potentially question the calculations of Itan et al. (2009), who propose that the genes of lactose tolerance in Europe were inherited from early Central European farmers. At the same time, it is important to note that the alleles determining lactose tolerance in Europe (−13910 * T) and in Arabia (−13915 * G) are different, which shows that this genetic trait developed independently (Ingram et al., 2009). The data from our previous study show that the spread of lactose tolerance in Europe is most closely tied to Y haplogroups I-M170 & R1b-U106 and the area of the North German plain22. The first appearance of this genetic trait can be dated to the Late Neolithic/Early Bronze Age (fourth to third millenium BC), which is already supported by paleogenetic data (Allentoft et al., 2015; Mathieson et al., 2015a). This period is characterized by the expansion of the Corded Ware culture in Northern and Central Europe, and a sudden dramatic increase in male stature of ∼7 cm. However, the subsequent dissemination of lactose tolerance in Western Europe and the British Isles must be connected with some other cultural circles, most probably the Bell Beaker culture (in the late third millenium BC) and Y haplogroup R1b-S11623. This assumption can already be supported by the first concrete paleogenetic evidence (Cassidy et al., 2015). Interestingly, other paleogenetic studies using autosomal DNA (Allentoft et al., 2015; Haak et al., 2015; Mathieson et al., 2015a, 2015b) suggest a more complex scenario. The populations carrying I-M170 & R1b-U106 may have originally acquired lactose tolerance from the pastoral Yamnaya people, who migrated to Central Europe from the Pontic/Caspian steppe around 3000 BC and formed the basis of the Corded Ware culture. This assumption is based on the strikingly high frequency of lactose tolerance in the Yamnaya culture and even other related steppe cultures of Central Asia (Afanasievo, Mezhovskaya, Karasuk; fourth to second millenium BC) (Allentoft et al., 2015), of which at least the latter two could be connected with Indo-Iranian speakers. Indeed, the allele that is responsible for the relatively high lactose tolerance in Pakistan (47%) and India (38%) is identical to −13910 * T from Europe (Ingram et al., 2009) (Appendix Fig. 15b). Furthermore, a positive geographical connection appears to exist between lactose tolerance in Central Asia/India and Y haplogroup R1a-Z93—a specific Indo-Iranian subbranch of R1a-M420 identified by Underhill et al. (2015) (Appendix Fig. 16). An eastern migration to Central Europe during the Late Neolithic/Bronze Age is also reflected by the very high (50%) frequencies of R1a-M420 in the available samples of Corded Ware males (Mathieson et al., 2015b), but as we emphasized in our previous article, R1a-M420 correlates slightly negatively with lactose tolerance in today's Europe (r = −0.10; p = 0.62) (see Appendix Fig. 15a). Only the combined frequency of two minor subbranches typical of Germanic speaking nations (R1a-Z284 and R1a-M417*) shows a certain positive relationship (r = 0.45; p = 0.045 in 20 countries). Furthermore, the problem with the findings of Allentoft et al. (2015) is that lactose tolerance frequency was determined only indirectly (based on the presence of mutations that accompany −13910 * T today). Mathieson et al. (2015a) could not find any trace of −13910 * T in the available Yamnaya samples and the oldest sample containing −13910 * T belonged to a man from the Bell Beaker culture (ca. 2300 BC). In any case, irrespective of the origin of −13910 * T, its frequencies started to increase markedly as late as the period after its emergence in Central Europe and all the above mentioned models are not mutually exclusive. When the available data on lactose tolerance from Europe are combined with those from North Africa, Asia and Oceania, they correlate highly with dairy proteins (r = 0.80; p < 0.001 in 44 countries) (Appendix Fig. 23a), but much less with milk protein (r = 0.39; p = 0.008 in 44 countries) (Appendix Fig. 23b). This counterintuitive finding is in line with the results from Europe. It shows that lactose tolerance is not a good predictor of the contemporary rates of milk intake, but it is a fundamental prerequisite for the long-term incorporation of milk products into the diet. In developing countries, milk usually serves as the main source of high-quality proteins, irrespective of the lactose tolerance of the local population. Furthermore, it is often consumed in the fermented form (yoghurt) that retains the biological quality of milk. The highly developed, lactose tolerant nations of Europe have gradually replaced milk with more expensive milk products such as cheese. Although cheese and curd contain only casein (the predominant form of protein in milk), which is of a somewhat lower quality than the complete milk protein, the advantage of these products lies in the much higher protein concentration, relative to their volume and energy intake. Not too surprisingly, lactose tolerance also has a strong relationship with the intake of ‘highly correlated proteins’ (r = 0.80) and animal protein in general (r = 0.75; p < 0.001) (Appendix Table 5e). On the other hand, it has the most negative relationship with rice protein (r = −0.60; p < 0.001) and protein from rice and legumes (r = −0.62; p < 0.001). Similar results are reported by Blum (2013), who suggests that lactose tolerance is a driving force of high animal protein intake. In Europe, lactose tolerance by itself strongly predicts male height (r = 0.71; p < 0.001) and its predictive power remains significant even after controlling for all the other factors that we examined, except its own genetic signature (I-M170 & R1b-U106). This suggests that the genes for lactose tolerance are associated with some genes that determine tall stature in Europeans. Indeed, the sudden introduction of ‘tall genes’ into Central Europe during the Late Neolithic/Early Bronze Age is apparent even in the paleogenetic analysis conducted by Mathieson et al. (2015a). In our present extended sample, lactose tolerance also has a strong positive correlation with male height in 47 countries (r = 0.80; p < 0.001) (Fig. 10d). Its significance as a correlate of male height is retained even after controlling for all the other variables, except ‘highly correlated proteins’ (p = 0.11). Nevertheless, it is paradoxical that Y haplogroup J1-M267 is associated with both lactose tolerance and shorter stature. All we can say is that with the exception of Kuwait and the UAE, the consumption of milk and other dairy products in the Arabian Peninsula is currently low and lactose tolerance thus does not bring any practical benefits. Genetics: South, Southeast and East Asia, and Oceania ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Another analysis of Y haplogroups was performed for South, Southeast and East Asia, and Oceania. We found that the correlations of Y haplogroups with male height in these regions are generally weak, but they markedly increase, after countries with negligible frequencies (2% >) are excluded. In this case, O1-MSY2.2 (r = −0.33; p = 0.20 in 17 countries), O2-P31 (r = −0.26; p = 0.37 in 14 countries) and O2a-PK4 (r = −0.43; p = 0.17 in 12 countries) tend to correlate with shorter statures. Only the combination of O1 & O2a reaches significance (r = −0.53; p = 0.017 in 20 countries) (Fig. 11a). In contrast, O3-M122 is significantly associated with taller statures (r = 0.42; p = 0.042 in 24 countries) (Fig. 11b). Interestingly, it is O3, not O1 that dominates in the tall Austronesian speakers of Oceania. Regression analysis of the total sample ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As already stated above, only 72 countries with the complete data for the 38 variables were used for the multiple regression analysis (Table 3b). To obtain the simplest predictive model, only four nutritional variables that were significant in the bivariate correlations were selected: ‘Highly correlated proteins’, poultry, rice (or rice & legumes), and total energy intake. Protein sources with daily intakes below 5 g/day were eliminated, similarly like largely duplicit items strongly related to ‘highly correlated proteins’ (total protein, animal protein, meat total) and plant proteins with a curvilinear relationship with height. Among the socioeconomic variables, we excluded only the human development index, because it is characterized by the strongest degree of collinearity. A separate regression with nutritional variables (Table 4) shows that when merely two items are considered, by far the highest percentage of variation (adj. R2) is explained by the combination of ‘highly correlated proteins’ with rice (79.5%), or with rice & legumes (79.0%). Only total energy intake slightly improves this model (to 80.9% and 81.5%, respectively)24. Out of all the socioeconomic variables, total fertility (86.3%) and children's mortality (85.1%) have the biggest additive effect, while urbanization influences the model only marginally (82.1%). The GDP per capita, health expenditure per capita and Gini index decrease it very slightly. Nevertheless, the best model explaining 87.2% of the variation was achieved via the combination of 7 variables. When we conducted a similar analysis with 83 countries, excluding the Gini index (for bivariate correlations, see Appendix Table 8), we obtained practically the same results, with only slightly lower adj. R2 values (Appendix Table 9). In all these models, rice (or rice & legumes) is the only variable that always retains a very high level of significance (p < 0.001). These findings point to rice as the most negatively correlated dietary factor, more so than wheat, which is somewhat unexpected. Although rice is a source of low-quality protein, similar to other cereals, it has a higher amino acid score, that is, a higher protein quality than wheat flour (0.63 vs. 0.48 according to the FAO standard 1985). Since rice protein correlates negatively not only with wheat protein (r = −0.69; p < 0.001), but even with many sources of high-quality animal proteins, especially total dairy (r = −0.66), ‘highly correlated proteins’ (r = −0.62; p < 0.001), and even total protein (r = −0.56) and total energy (r = −0.51, p < 0.001), it could be assumed that high rice consumption symbolizes a diet with a low content of milk and other important foodstuffs, and reflects general malnutrition. This conclusion could also indicate that rice reflects general poverty, but the preference for rice or wheat apparently bears no close relation to the national GDP per capita (Appendix Figs. 24a and 24b) and rice is even more expensive to harvest than wheat25. Furthermore, the same strong polarity between height and rice/wheat that was documented in our sample of 93 countries exists in India and China (Appendix Figs. 25a–25d and 26a–26d). Rice and wheat do not show any significant association with the GDP per capita in 29 Indian states and 30 Chinese provinces, but rice correlates significantly negatively with male stature both in India (r = −0.62; p < 0.001) and China (r = −0.41; p = 0.024). In contrast, wheat shows a positive relationship with male stature in India (r = 0.53; p = 0.003) and tends to have the same effect in China (r = 0.25; p = 0.19). Wheat and rice correlate strongly negatively with each other, especially in China26. Therefore, it is not poverty per se, but mainly geography that influences rice consumption, because rice-producing regions are unsuitable for the cultivation of wheat (and vice versa). On the other hand, we should also explain, why the intake of protein and energy in poor, wheat-consuming nations is much higher than in comparably poor, rice-consuming nations (Appendix Figs. 10a and 24c). In Vietnam, living in a farming community and having a lower socioeconomic status are the main determinants of a low energy intake and high carbohydrate consumption, the latter being directly related to rice (Nguyen et al., 2013). This points to limited food alternatives in isolated, poor farming regions of tropical Asia, which are not suitable for large-scale production of both wheat and livestock. Indeed, another important factor from the World Bank database, arable land (% of land area, 1993–2009)27, has the most positive relationship with the consumption of protein from beef (r = 0.42; p < 0.001) and milk (r = 0.41; p < 0.001) in 93 countries, and it is also positively tied to ‘highly correlated proteins’ (r = 0.34; p < 0.001), total protein (r = 0.23; p = 0.029) and wheat protein (r = 0.21; p = 0.048). On the other hand, it correlates most negatively with proteins from rice & legumes (r = −0.31; p = 0.003), rice protein (r = −0.27; p = 0.008) and legume protein (incl. soy) (r = −0.27; p = 0.009). In accordance with this finding, the consumption of milk and and the consumption of wheat in India are strongly associated with each other (r = 0.81; p < 0.001 in urban areas, r = 0.73; p < 0.001 in rural areas). Furthermore, when the complete data from 72 countries are considered, the socioeconomic factor with the strongest (negative) relationship with rice is urbanization (r = −0.57; p < 0.001), which also indicates the influence of narrow food choices. Still, the lack of other food alternatives in the farming communities of tropical Asia cannot explain, why the maximum consumption rates of rice protein in the FAOSTAT database (∼30 g/day) are much lower than those of wheat protein (∼50 g/day). One possible explanation, already outlined above (see Appendix Fig. 21), is that rice is characterized by a very low nutrient density. The content of protein in rice is much lower than that in wheat28 and when protein digestibility (PDCAAS score) is taken into account, roughly 24% more energy from cooked white rice must be consumed per gram of complete protein, when compared with wheat flour, in addition to the weight being nearly 4-times greater29. The data from the Czech Nutridatabaze.cz indicate even greater differences (43% more energy) between white bread and husked parboiled rice30. Therefore, rice may not only further exacerbate the negative effect of a low total protein and energy intake, but its high consumption may also directly contribute to it31. Residuals of observed and predicted height in the total sample ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The comparison of observed and predicted height (Table 5 and Fig. 12a and b), based on model (10) from Table 4, shows that the region of tropical Asia is characterized by a clear tendency towards shorter heights than the model predicts (−0.5 cm). Furthermore, after the exclusion of the outlier urban sample from Laos, the difference rises to −0.8 cm. In contrast, China and South Korea are above the predicted value. Apparently, these results are compatible with the presumed relationships between height and various subbranches of Y haplogroup O-M175 in East and Southeast Asia, although the supposed genetic impact would be relatively small. The influence of genetics in North Africa and the Near East also appears to be plausible, when only nutrition is considered, but it diminishes markedly after socioeconomic factors are taken into account. Perhaps the most interesting observation is thus the position of five Altaic-speaking nations of Central Asia, which are consistently below the predicted value (−2.0 cm on average). Not too surprisingly, countries from Western and Southwestern Europe have negative residuals, whereas positive residuals are generally the most prevalent in the Western Balkans and Central/Northern Europe (as much as +6.3 cm in Bosnia and Herzegovina, +3.6 cm in Croatia and +3.1 cm in the Netherlands). Without the Gini index, we again get very similar results (Appendix Table 10 and Appendix Figs. 27a and 27b), but this time we can even assess Oceania and more Muslim countries. Interestingly, three countries from Remote Oceania (Fiji, Kiribati, Samoa) are considerably above the predicted height, and the positive residual in Samoa reaches +4.2 cm. Two countries from Near Oceania (the Solomon Islands, Vanuatu) do not come close to these values.","The current study extends our previous data from Europe and enables a better understanding of the enviromental determinants of physical growth in the developing world. The most fundamental finding is that the nutritional correlates of male height in North Africa, Asia and Oceania are very different and primarily depend on protein quantity, not protein quality. Furthermore, three basic nutritional styles can be distinguished, depending on the major source of protein: The first nutritional style (in tropical Asia) is based on rice and is also characterized by a very low consumption of protein and energy. It is accompanied by very small statures between 162 and 168 cm. The second one (in the Muslim countries of North Africa and the Near East) is based on wheat and the consumption of plant protein reaches the highest values in the world. The intake of total protein and total energy is relatively high as well and comparable with Europe, but the average height of young males is still rather short and does not exceed 174 cm. The third one is based on animal proteins (particularly those from dairy) and is typical of Northern/Central Europe. This region is characterized by the tallest statures in the world (>180 cm), being matched only by the inhabitants of the Western Balkans, in which we can presume extraordinary genetic predispositions. These world patterns in protein consumption have already been described in detail by Grigg (1995). Although our present study is based on a non- experimental, ecological comparison, its findings can provide useful insights into the relationship between these nutritional styles and the final adult stature. Most importantly, our results indicate that plant-based diets are not able to provide the optimal stimuli for physical growth, even if the intake of total protein and total energy poses no problem. In fact, we observed a difference of 10 cm (174 cm vs. 184 cm) between nations relying on the surplus of plant and animal proteins, respectively32. A low consumption of proteins that correlate highly with height can explain the seemingly perplexing, small stature in the developed countries of East Asia and the Muslim oil superpowers. Besides low protein quality, a frequently forgotten limiting factor of plant- based diets is their low nutritional density, with a disproportionate load of ‘empty calories’ from starch and oils that must be consumed per unit of a key nutrient. The countries with the highest plant protein intake already belong to the most obese in the world, particularly among females33, so it is not likely that the deficit of protein quality could easily be compensated by protein quantity. Last, but not least, our study can potentially question the current dietary recommendations regarding the intake of essential amino acids, because some foods that score highly according to the new FAO standard 2007 do not appear among the best correlates of height. In fact, recent studies indicate that even the contemporary total protein requirements for children are underestimated (Elango et al., 2011)34, which also agrees with our data, because we do not observe any levelling-off in many graphic comparisons of male height and protein consumption. Of all other variables examined in this study, the human development index (HDI) is the only factor that shows a comparably strong relationship with male height like to nutrition. This indicates that the factors leading to the increase in the average height intertwine with public policies that improve the overall quality of life. As in our previous study, children's mortality (i.e. a disease free environment) is the strongest correlate of stature among all the remaining socioeconomic indicators, but the forward stepwise regression also highlights the role of a lower total fertility rate (i.e. the amount of resources that can be expended per child) and partly urbanization as additional factors that can be targeted, when trying to speed up the pace of the positive height trend. Besides that, our study shows that, similar to the situation in Europe, the final height in non-European regions may be influenced by genetic factors. Their role in North Africa and the Near East appears to be similarly strong like in Europe, and the inverse relationship between height/lactose tolerance in this region is intriguing. The results are less persuasive concerning the southeastern part of Asia and Oceania, but genetic, socioeconomic and nutritional data from many local countries are still lacking. In any case, the verification of these findings is possible only via studies of autosomal DNA."],["We develop an empirical search-matching model which is suitable for analyzing the wage, employment and welfare impact of regulation in a labor market with heterogeneous workers and jobs. To achieve this we develop an equilibrium model of wage determination and employment which extends the current literature on equilibrium wage determination with matching and provides a bridge between some of the most prominent macro models and microeconometric research. The model incorporates productivity shocks, long-term contracts, on-the-job search and counter-offers. Importantly, the model allows for the possibility of assortative matching between workers and jobs due to complementarities between worker and job characteristics. We use the model to estimate the potential gain from optimal regulation and we consider the potential gains and redistributive impacts from optimal unemployment benefit policy. Here optimal policy is defined as that which maximizes total output and home production, accounting for the various constraints that arise from search frictions. The model is estimated on the NLSY using the method of moments. --------------------------------------------------------------------------------","Labor market imperfections may justify labor market interventions. Within a competitive framework regulation would be welfare reducing, as it would typically reduce employment and increase insiders' wages. By contrast, any friction constraining the allocation of workers to jobs inevitably allows some agents to appropriate a greater share of the rent than a central planner would deem fit. Indeed, if there are important complementarities in production, mismatch may produce substantial welfare losses relative to the first best or a constrained planner. We develop a search-matching model in which workers with different abilities are assigned to different tasks. Our model combines elements from key papers in the equilibrium search literature. Thus we allow for endogenous job destruction because of productivity shocks, drawing from the seminal paper of Mortensen and Pissarides (1994). To this framework we introduce two sided heterogeneity with potential complementarities between job and worker productivity following Shimer and Smith (2000). In this way we can investigate sorting in the labor market, one of our motivating interests. To account for job-to-job transitions and to better explain wage growth we allow for on-the-job search, which is not present in Shimer and Smith (2000). Wage determination is drawn from Postel- Vinay and Robin (2002), Dey and Flinn (2005) and Cahuc et al. (2006). In other words we combine bargaining over the surplus (for workers out of unemployment or moving to a new job) with Bertrand competition when a poaching firm is involved. In this way, our model provides a bridge between an essentially theoretical literature on allocation of heterogeneous workers to jobs,1 and the large microeconometric literature on job mobility and wage dynamics.2 The framework that we develop allows for frictions and inefficiencies in the labor market. The frictions are due to the time it takes to locate jobs, which means that individuals will spend time searching for a job while unemployed and if working they are likely to be mismatched (if there are complementarities in production). One may argue that such frictions are inevitable; nevertheless it is important to understand the welfare loss that they cause relative to the benchmark of a frictionless economy, because this provides a measure of how dominant search frictions are in determining economic outcomes.3 More importantly, our model can quantify the welfare loss from inefficiencies that can potentially be addressed by labor market regulation: first, the number of job seekers cause congestion making it harder for others to find jobs – this is a standard externality in models with endogenous arrival rates. Beyond that the potential complementarities between worker and firm productivities, the possibility of on-the-job search and the lack of commitment on the worker side allows for the possibility of another inefficiency: workers and jobs are sometimes willing to form matches whose flow output is lower than the combined cost of a vacancy and the lost out-of-work benefit. On the one hand the job, with its local monopsony power manages to extract sufficient surplus to make it worth hiring the worker; on the other hand the worker prefers the resulting current loss to the increased flow of income when out of work because being in a job provides a better outside offer to negotiate a wage once an alternative offer arrives. These features may be important in the labor market and our model can quantify their importance for welfare. However, Eeckhout and Kircher (2011) strongly argue against the possibility of identifying assortative matching in labor markets by this approach, bias-corrected or not. This is because the surplus of a match is not a monotonic function of worker and firm characteristics in general. Depending on the distribution of matches around the optimal, Beckerian allocation, a positive or a negative AKM correlation can be estimated irrespective of the sign of the correlation between workers' and firms' true unobserved characteristics in the population of active matches. In the previous working paper we also find that the AKM correlation is misleading and demonstrate that positive sorting with respect to unobserved characteristics may induce a negative correlation in the worker and firm effects estimated on a panel of wages. Lopes de Melo (2009), Hagedorn et al. (2012), Bagger and Lentz (2014) reach similar conclusions with different models. Note that this argument invalidating a structural interpretation of AKM is implicit in Gautier and Teulings (2006), who estimate a regression model of log wages on a quadratic function of worker and employer types. These types are calculated by projecting log wages separately on workers' and employers' observed characteristics. Gautier and Teulings's estimation rests on various parametric restrictions, but nevertheless convincingly claims 1) that the wage equation is nonlinear in workers' and employers' types, and 2) that they are positively correlated. Identification of complementarity and sorting, given that we use a panel of workers' wages and labor market transitions extracted from the NLSY,5 is also hard to prove or disprove theoretically. However, we argue that sorting can be seen in the way wage and employment mobility vary as a function of the length of time spent working following an unemployment spell. Due to search frictions, the cohort of workers entering the labor market following a spell out of work will start off mismatched, but through on- the-job search they will become better and better matched with time spent in the labor market. If the tendency to sort in equilibrium is strong this will result in wages spreading out as workers sort themselves. Monte Carlo simulations for the Simulated Method of Moment estimator that we have implemented here seem to suggest that the set of moments we match do identify the degree of sorting. We find that the NLSY data are best fitted by our model assuming no complementarity and zero sorting for workers with high-school education or less and positive complementarity and sorting for college-graduates.6 Finally, our model offers an empirical framework for understanding employment and wage determination in the presence of firm–worker complementarities, search frictions and productivity shocks. As a result it offers a way for evaluating the extent to which regulation may be welfare improving and can evaluate the impact of specific policies such as unemployment benefit. In a search framework with match complementarities unemployment benefit can have ambiguous effects on employment and total output. On the one hand, it allows workers to be more picky and form better matches, which comes at the cost of longer unemployment spells and higher unemployment. On the other hand, the fact that higher quality matches will be formed may induce firms to create more jobs, increasing the contact rate and potentially reducing unemployment duration. Our framework allows this effect to be quantified (see also Acemoglu and Shimer, 2000). And, in addition, it allows us to analyze the effect of such policies on the distribution of welfare thus showing who pays and who benefits from such a policy in this non-competitive environment. We find that the degree to which labor market interventions are justified depends on whether we are looking at the low or high skilled markets. Our finding that the market for unskilled labor (high school graduates or less) is characterized by an extremely low degree of complementarity implies that mismatch is not very costly. If we reallocated the employed workers in this group optimally the difference in (steady state) welfare would be 1.1 percent. On the other hand, among the college graduates, optimally reallocating the employed workers produces a difference in (steady state) welfare of 6.8 percent. We also find that search frictions are significant. If we do the same experiment with full employment the welfare changes are 7.8 and 19.6 for the high school and college groups respectively. The interaction of search frictions and the cost of mismatch (the degree of complementarity) differ between the two labor markets, implying that a social planner who is constrained by these friction could attain a welfare increase of 2 percent for the high school group but only 0.7 percent for the college group, largely by trading off the level of employment against the cost of creating vacancies. Finally, we find that if we limit the planner to an optimal unemployment benefit program, this can go a long way to realizing the potential welfare gains for the high school group. However, it is ineffectual for the college educated group as the resulting employment distortions outweigh the gains from improved match quality. The paper proceeds as follows. Section 2 describes the model, Section 3 and Section 4 present the estimation procedure. Sections 5 and 6 describe the data and the choice of moments. Section 7 presents the results of estimation, analyses the fit and discusses estimation of the degree of sorting. Section 8 presents the welfare analysis, the cost of search frictions, and analyses in detail the optimal unemployment benefit policy. Section 9 concludes.","We build a model of individual employment and wage dynamics, with heterogeneous workers and jobs and with productive complementarities at the match level. This model draws from Mortensen and Pissarides (1994), as far as the process of match creation and destruction is concerned, and from Postel-Vinay and Robin (2002), Dey and Flinn (2005) and Cahuc et al. (2006) in order to incorporate on-the-job search in the Mortensen–Pissarides model. In addition, we draw from Postel-Vinay and Turon (2010) the renegotiation mechanism for wages following firm-level productivity shocks. The wage dynamics follow from the process of search and matching and from firm-level productivity shocks, but entirely abstract from human capital accumulation and idiosyncratic ability shocks. While important, incorporating human capital accumulation would complicate matters beyond the scope of this paper.7 Meetings and match formation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Individuals and jobs are risk neutral and we assume efficiency, in the sense that any match where the surplus is positive will be formed when the worker and the job meet. Under these conditions we can characterize the set of equilibrium matches and their surplus separately from the sharing of the surplus between workers and jobs. Wages ~~~~~ Different wages are negotiated when leaving unemployment, upon poaching, or after a shock to the job productivity. Poaching Wages can only be renegotiated when either side has an interest to separate if they do not obtain an improved offer, assuming that the match remains viable for both parties. The events that can trigger renegotiation occur when a suitable outside offer is made, or when a productivity shock changes the value of the surplus sufficiently. We consider first the impact of an outside offer. In our approach there is an asymmetry between workers and firms because the latter do not search when the job is filled. As a result they do not fire workers when they find an alternative worker who would lead to a larger total surplus, nor do they force wages down when an alternative worker is found whose pay would imply an increased share for the firm. We decided to impose this asymmetry because in many institutional contexts it is hard for the firm to replace workers in this way. Moreover, we suspect that even when allowed firms would be reluctant to do so in practice. We do, however, allow firms to fire a worker when the current surplus becomes negative and then immediately search for a replacement. Productivity shocks Note that a firm hiring an unemployed worker offers the value of unemployment plus a share of the surplus, whereas after a positive shock to the surplus value there is no Nash bargaining. This may seem both ad hoc and inelegant. However, consider the other situation when the match surplus falls below the worker surplus at current wage. Under Nash bargaining the worker would get the same share of the new surplus as in the other triggering situation, wiping out any wage gains due to outside offers. That does not sound right. Intuitively, the worker should have more bargaining power in one case than in the other. Our assumption (zero bargaining power when the worker values falls below the value of unemployment; full bargaining power when the firm profit falls below the value of a vacancy) may seem a bit extreme, but it is motivated by the idea that the worker's bargaining power should depend on whether renegotiation is wanted by the worker or the employer, and it ensures that the value function of the worker is monotonically increasing in the wage. Furthermore, wages respond to job specific productivity shocks, but not always in an obvious direction. Separations and pay changes may happen following both good shocks that increase the value of productivity y and bad shocks that decrease it. It is all about mismatch: what matters is what happens to the overall surplus. A positive productivity shock, for example, can imply that the quality of the match becomes worse and the surplus declines, since the outside option of the firm has changed and it may be worthwhile to separate from the current worker and post a vacancy to find a better worker. Conversely a negative productivity shock can improve the surplus if this means the job type is now closer to the optimal one sought by the worker. A shock that reduces the surplus can still lead to a wage increase to compensate the worker who is now matched with a job with fewer future prospects of wage increases. Thus what really matters as far as the viability of the match and the possible options for renegotiation is whether a shock improves or worsens a particular match, measured by whether it leads to an increase or a decrease, respectively, of the surplus. Value functions ~~~~~~~~~~~~~~~ The next step in solving the model is to characterize the value functions of workers and jobs, which have been kept implicit up to now. These define the decision rules for each agent. We proceed by assuming that time is continuous. Employed workers On the right hand side the first term is the wage net of the flow value of unemployment (all in parentheses). The second term is the expected excess value to the worker of a productivity shock (times the probability that it occurs): in this case the worker either ends up with the entire new surplus or the new value or indeed nothing if the match is no longer feasible (see equations (7) and (8)). The third line is the expected excess value following an outside offer as in equation (5). The integral is over all offers that can improve the value (whether the worker moves or not). Steady-state flow equation This equation defines the steady-state equilibrium, together with the accounting equations (1) and (2). Free entry Appendix B provides a simple iterative algorithm that uses these equations and that of the surplus to compute the equilibrium objects. Measurement error ~~~~~~~~~~~~~~~~~ In our model wages are stochastic because of the type of firm that an individual may encounter, the outside offers or the layoffs that may occur and the productivity shocks. There will also be unexplained variation due to her productivity characteristic x. In the data an additional source of variation is measurement error that we need to account for, so as not to bias the other sources of variation. We use the monthly records of wages in the NLSY. Thus, while it may be reasonable to assume that measurement error is independent from one year to the next it may not be so within the year, as all records are reported at the same interview with recall. Having experimented with a number of alternatives, including a common equi-correlated component across all months, we settled on a measurement error structure that is common within year and independent across years. The variance of measurement error, which is assumed to be lognormal, is estimated alongside the other parameters of the model.","The simulated moments are not necessarily a smooth function of the parameters, although they would become so as the number of simulations increased to infinity. However, with any finite and relatively small number of simulations derivative based methods are not appropriate for finding the minimum. We thus use a method developed by Chernozhukov and Hong (2003), which does not require derivatives of the criterion function. They construct a Markov chain that converges to a stationary process of which the ergodic distribution has a mode that is asymptotically equivalent to the SMM estimator. Appendix C describes this procedure in detail.","We use the 1979 to 2002 waves of the National Longitudinal Survey of Youth 1979 (NLSY). The NLSY consists of 12,686 individuals who were 14 to 21 years of age as of January, 1979. It contains a nationally representative core random sample, as well as an over- sample of black, Hispanic, the military, and poor white individuals. For our analysis, we keep only white males from the core sample. We only include data for individuals once they have completed their education. We also drop individuals who have served in the military, and follow workers up to the point of a non-employment spell of 36 months or longer. We consider these workers to have left the labor force. We subdivide the data into two education groups: high school degree or less, and college graduate. The model is estimated separately on these subgroups. Individuals are interviewed once a year and provide retrospective information on their labor market transitions and their earnings. From this we construct histories at a monthly frequency aggregating the data as follows: we define a worker as employed in a given week if he worked more than 35 hours in the week. We define a worker's employment status in a month as the activity he was engaged in for the majority of the month, treating unemployment spells of two weeks or less as job-to-job transitions. After sample selection, we are left with an unbalanced panel of 2125 individuals (446,747 person months). We remove aggregate growth from wages based on average wage growth in years 10 to 20 from completion of education. At this point the main source of wage growth is due to aggregate productivity. Having removed this constant growth rate from wages we assume that the remaining growth is attributable to gains from job search.11 Finally, we trim the data to remove a very small number of outliers when calculating wage changes. When calculating month to month wage changes, we exclude observations where the wage falls by more than half or increases by more than a factor of five.","The model will be estimated based on worker level data recording transitions between jobs and between work and unemployment as well wages over time. We have no information on the firm side itself (such as for example productivity). As such identification is challenging. Lamadon et al. (2014) discuss the formal non-parametric identification of a model similar to this, albeit without productivity shocks and using matched employer employee data. Other insightful work with more formal discussion of identification is Hagedorn et al. (2012), who analyze a version of the model without on-the-job search, and Bagger and Lentz (2014), where sorting results from heterogeneous search strategies. The key result is that the complementarities in the worker–job match can be identified nonparametrically in their context, using job-side information on productivity and job duration. Here we present a heuristic description of the identification argument based on the simpler data at our disposal and using a specific parametric form for the match production function. Transitions in and out of work and between jobs play a key role in parameters controlling labor market mobility. In particular η (matching efficiency) and s (relative search intensity) that help determine the arrival rate of offers are identified by exit from unemployment and by mobility between jobs. Exit from employment is governed by two components of the model: productivity shocks to the job and the exogenous destruction rate. However, productivity shocks are also related to job-to-job transitions and wage growth and cannot be set to fit perfectly the exit rate from employment. The remainder is captured by the exogenous match destruction rate ξ. Changing the distribution of y and its shocks affects transitions both between jobs and in and out of unemployment. In addition, changing the distribution of x will affect the variance of wage growth in a very specific way, governed by the structure of pay setting and will also affect transitions in and out of work. Thus the distribution of these objects is intimately linked to observed transitions, which will limit the extent to which we can explain the variance of wages and their growth. We thus identify the variance of measurement error in wages from the variance of wages that the economic model is unable to reproduce. Identification of the model in practice requires a careful choice of moments that will be sensitive to the parameters we need to estimate. In particular, to summarize the employment dynamics, we use the long-run employment rate and the transitions between employment states and between jobs, calculated on years 16–20. To summarize wage dynamics, we include the level of (log)wages and their cross sectional variance as well as wage growth and its variance both within and between jobs. Each of these moments is calculated separately by year in the labor force (year since leaving school for the cohort). All the moments we use, their values in intervals of five years and the value produced by the model are shown in Table 4. In addition, we target the mean vacancy to unemployment rate (based on the mean and standard deviation from Hagedorn and Manovskii, 2008). The fit of the moments ~~~~~~~~~~~~~~~~~~~~~~ Summaries of the fit of the targeted moments by the model are presented in Table 4. As seen from the table the transitions rates are fitted remarkably well. In Figs. 1 and 2 we summarize the fit of the model for wages, wage growth and the corresponding variances, both overall and by type of transition for the lower and higher education individuals respectively. The model generally fits these patterns very well and certainly captures the qualitative features of the data.13 Any wage growth generated by the model is due to the job search process and reflects mobility towards better jobs and (for the higher educated people) improvements in sorting. For the lower educated people as we shall see there are very few complementarities; however as the workers move to higher surplus jobs (because of improved firm productivity) and by receiving outside offers they can improve their wage. This process of outside offers is responsible for the observed wage growth for the low skilled. Parameter estimates ~~~~~~~~~~~~~~~~~~~ The complete set of parameter estimates is presented in Table 5. Here we focus on a subset that have a direct economic interpretation; these are presented in Table 1. The key parameter of interest is the complementarity parameter (ρ). For lower skill worker this is close to and indeed not significantly different from one. This implies that there are practically no complementarities between worker and job characteristics, which appear to be perfect substitutes. However, for college graduates the elasticity of substitution is about 0.53 implying a high degree of complementarity and hence large gains from sorting. The gains from moving up the job ladder are thus much more important for college graduates than for those with lower levels of education. However, as mentioned above, the unskilled can still gain from outside offers: first, search frictions will imply a surplus, since the departure of the worker will mean the job will remain idle for some time and hence the firm will have an incentive to match outside offers from jobs with lower surplus. Second, the surplus is increasing in job productivity y but (almost) not in worker productivity x. This is because the cost of a vacancy is constant but the flow of out-of-work income is increasing in x (see equation (13)). Hence higher productivity firms can afford higher wages. The way the surplus is split is driven by the Nash bargaining parameter β. This is slightly higher for college graduates than for unskilled workers. The former obtain 27% of the surplus while the unskilled about 18%. Turning now to the parameters governing search, the search intensity (s) for the employed workers is a third of that for those out of work among the low education group and half that among the college graduates. Effectively this means that the rate of arrival of job offers is much higher for those out of work and this has an implication on what jobs the unemployed are willing to take. The implications of the parameter estimates of the search technology (η and s) are better understood by calculating the probabilities of a job contact when employed and unemployed, which are displayed in Table 2. The contact rates for the unskilled when unemployed are lower than for the higher skilled (17.1% compared to 23.1%). However the unskilled accept all offers when unemployed. Both the contact rate when working and the rate at which alternative offers are accepted declines with skill. Sorting in the labor market ~~~~~~~~~~~~~~~~~~~~~~~~~~~ The complementarities between worker and job characteristics for the higher education group imply perfect assortative matching in a first best world. However, the extent of sorting that occurs in practice depends on the importance of frictions. In Fig. 3 we summarize the sorting patterns in the decentralized equilibrium, as implied by our point estimates of the parameters. The lines in the figure are contours of the surplus function. On each contour the surplus is constant and non-negative, for various combinations of worker and job productivities. In the left hand panel, corresponding to the low education workers, such lines cover the entire support of the distribution of worker and firm characteristics. This reflects the fact that all matches are viable, leading to a non- negative surplus, because of the almost complete absence of complementarities.16 In the right hand panel, relating to college graduates, the upper-left and lower-right sections have no contours: these areas represent combinations of worker and job types where the surplus is negative and no matches ever occur.17 Waiting for a better match has higher value than starting to produce. For low skill workers over most of the space the contours are downward sloping. This points to a trade off between firm and job characteristics at a fixed match surplus. However, for college graduates the contours are mostly upward sloping (except at the highest levels of skills and job characteristics). This is because, when complementarities are very important an increase in the job productivity requires an increase in the human capital of the worker if the surplus is to remain constant, rather than decline. As an example of the matching process consider a college educated worker at the 50th percentile of the x distribution. The surplus initially increases in the type of the job, is maximized when matched to a job at the 50th percentile, and then declines again. This worker will initially match with any job above the 15th percentile, but will always move to a job that is closer to the 50th percentile, which may involve moving up or down the quantiles of y. Fig. 4 illustrates sorting within the College educated group by plotting the distribution of worker (job) characteristics conditional on various percentiles of the job (worker) productivity they are matched with. Distributions conditional on higher values of productivity of the counterpart to the match stochastically dominate those that are conditional on lower values. Moreover, the support of these distributions is limited to a strict subset of the entire support, reflecting the fact that some matches never occur. Finally, matches are not uniformly distributed over this support, even for the low educated. We illustrate the density of matches over the matching set in Fig. 5. For the college group, as we approach the main diagonal, of perfect sorting the density of matches increases. For the high school group, all workers above the 20th percentile move up toward the highest firm type at the same rate. Below the 20th percentile there is some sorting induced by the mild complementarity (for low x-types the surplus decreases when moving to higher y-type firms, although it is still positive. A caveat here is that our estimate is not distinguishable from linearity, in which case there would be no sorting at all for this group.","To provide a sense of the potential gains from policy, we consider three thought experiments. First we take the estimated search frictions as given and look at the Planner's constrained efficient solution. This experiment provides an upper bound on what can be achieved by policies that work to eliminate congestion externalities, taking the frictions as given. Second, we consider the thought experiment of ignoring the search frictions and solving the frictionless assignment problem. We do this both keeping the number of employed worker fixed to the estimated level, as well as for the full employment case, allowing us to separate the employment from the mismatch effects of frictions. This experiment provides an estimate of the upper bound to the benefits of finding technological solutions around the frictions (such as improved centralized matching) but has nothing to say about either the feasibility or costs of such a program. Finally, we consider an optimal Unemployment benefit program. This experiment provides an estimate of the potential gain from a feasible policy. The planner's constrained efficient solution ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The potential for welfare-enhancing labor market regulation arises from the job search frictions and the externalities they cause during the job allocation process. The externalities arise from the classic issue of “overcrowding” among job seekers, i.e. when an extra person or vacancy seeks a match it reduces the arrival rate for others as implied by the matching function. An extra dimension arises in our model because of heterogeneity and sorting: by having low quality jobs compete for workers they lengthen the time it takes to fill higher productivity ones, without adding much when they are filled (because they have zero or near zero surplus).18 This implies that because of complementarity, cutting some low productivity jobs may increase welfare, even if this means that some very low productivity workers never work. Any regulatory intervention in the labor market will improve welfare only to the extent that it can address the externalities discussed above, and to the extent to which they are significant. Thus, to provide a measure of the potential welfare gains from labor market regulation (such as in-work benefits, unemployment benefit, minimum wages, severance pay etc.) we solve the planners problem respecting the constraints arising from search frictions. Table 3 shows the breakdown of contributions to total welfare under different scenarios. The first column relates to the fully decentralized economy we observe from the data. The second column shows the results of the planner maximizing welfare as in (16). For the lowest education group the constrained planner is able to improve on the decentralized outcome by two percent. For the higher education group the planner can attain an increase in welfare of only 0.71 percent. The planner increases unemployment and reduces the number of jobs. To understand what is going on, note that increased match quality contributes nothing to the welfare increase. Output per match does not change. The planner increases welfare by reducing vacancies (which are very costly) and allowing more workers to engage in home production. For the high school or less group, where there are no complementarities in production, this reallocation achieves a relatively large increase in welfare. For college graduates, the welfare gains are much more modest because by reducing the number of vacancies some high surplus jobs are also eliminated and home production is less effective for this group. Nevertheless, the reduction of jobs for them is much lower. The costs of search frictions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In column (4), we run a similar counterfactual in which we assign all jobs to workers. Unemployment is drastically reduced and welfare increases a lot, largely because there are no recruiting costs, but also because of reduced unemployment, and a slightly higher match quality for the lowest education groups (see also the discussion in Subsection 8.3.2). This experiment indicates that there may be substantial gains from reducing search frictions: for the lowest skill group policies that improve search technology could improve welfare by up to 7.8% (19.6% for the college educated group). Part of this increase comes from reducing unemployment. But part also comes from improving sorting if we take the point estimate of the elasticity of substitution, which is 16.2. This can be seen from the third column of Table 3 where unemployment is kept equal to the level in the benchmark economy and we observe a rise in match quality. Overall efficiency gains For the low education group, optimal unemployment benefit can deliver 1.4% of improved welfare, corresponding to 68.8 percent of the potential gains attainable by the planner working under the same constraints. This involves increasing the baseline flow utility of being out of work (home production) for each individual by 11.09 percent of their expected output if employed, and financing this by a tax on output of 0.95 percent. The gain is effectively zero for the high education group. It is worth noting again that the improvement in steady state output comes from very different sources when comparing the elimination of frictions to the constrained planner or optimal unemployment benefit scheme. With the removal of frictions there is a direct increase in market production and a gain when netting out the costs of vacancy creation from home production. For the low education group, where the constrained planner can improve steady state output, this is implemented largely by raising unemployment and reducing the number of vacancies, resulting in lower vacancy creation costs and higher levels of home production, but without improving the average quality of productive matches. The redistributive effects of the policy As discussed above, we find that there is a potential aggregate gain from an optimal unemployment benefit scheme (although negligible for college graduates). In addition to overall efficiency gains, we are also interested in the redistributive effects of policy. In an environment with heterogeneous workers it is not necessarily the case that an increase in steady state output will benefit all workers the same, indeed it may harm some. In Fig. 6 we plot the difference in value, by worker type x, between being unemployed in an economy with and without the optimal unemployment benefit scheme. While the value of unemployment is higher for all worker types in the low education group, it is effectively zero for college educated workers above the second quintile. Thus, there are no losers from this optimal UI policy, but the gains are concentrated among the low educated as well as the lower productivity individuals among the college graduates.","We develop an equilibrium model of employment and wage determination, which builds on the work of Mortensen and Pissarides (1994), Shimer and Smith (2000) and Postel-Vinay and Robin (2002). In our model both workers and firms are heterogeneous and their productivity characteristics are potentially complementary in production creating the possibility of sorting. However, firms are subject to productivity shocks. Workers can search both on and off the job. This creates an environment where there may be potential for welfare improving labor market regulation. Moreover our framework is well suited to consider the redistributive (as well as efficiency) implications of policy. The scope for and impact of policy is thus an empirical issue in our model. We estimate the model based on NLSY data and find strong evidence of complementarities between worker and firm characteristics, leading to sorting for the college educated workers. For this education group these complementarities imply large efficiency losses due to mismatch between job and worker productivities caused by search frictions. The complementarities are much weaker for the low education workers, where the production function is effectively linear in individual an job productivities. Mismatch is a source of inefficiency that labor market regulation cannot correct; this would require changing the job search technology, improving job finding rates and enabling more mobility following shocks. However, we show that the potential welfare gains from eliminating mismatch and frictions can be as high as 20% for college graduates and 8% for the lower educated. Policies such as unemployment benefit can improve efficiency to the extent that they address the externalities induced by search frictions. We establish that optimal labor market regulation can improve welfare by up to 2% for low skill workers and 0.71% for college graduates. Some 70% of the improvement can be achieved with optimal unemployment benefit alone for the lower educated individuals, but such a policy can achieve nothing for the college graduate group. Our model opens up an empirical research agenda on which to build and address important issues. We demonstrate the importance of heterogeneity, sorting and search frictions. Among these are the welfare and labor market effects of risk and the role of assets in determining the wage offer distribution and the role of investment in human capital. Similarly, an important extension of such a model is considering investment decisions by firms and how this can affect productivity y which we took as given. Finally, this kind of model is well suited to interpreting matched employer employee data. Indeed such data could aid identification by providing direct information on firm level productivity, and on the distribution of worker types employed at the same firm type (see Lamadon et al., 2014). However, when we move to such data new, important and difficult questions arise when defining wage setting in an environment with sorting and multiple workers per firm. This is of course an important future area of research."],["Mendelian randomization methods, which use genetic variants as instrumental variables for exposures of interest to overcome problems of confounding and reverse causality, are becoming widespread for assessing causal relationships in epidemiological studies. The main purpose of this paper is to demonstrate how results can be biased if researchers select genetic variants on the basis of their association with the exposure in their own dataset, as often happens in candidate gene analyses. This can lead to estimates that indicate apparent \"causal\" relationships, despite there being no true effect of the exposure. In addition, we discuss the potential bias in estimates of magnitudes of effect from Mendelian randomization analyses when the measured exposure is a poor proxy for the true underlying exposure. We illustrate these points with specific reference to tobacco research. © 2013 The Authors. --------------------------------------------------------------------------------","Proving how exposures affect health outcomes can be problematic in observational studies. Even if an exposure and an outcome are associated, the direction of causality can be difficult to ascertain because health outcomes can lead to changes in behaviour which can affect exposures (Munafò and Araya, 2010). Mendelian randomization studies may help to shed light on these relationships by using genetic variants, such as single nucleotide polymorphisms (SNPs) (see Table 1 for definition), as instrumental variables for measured lifestyle exposures (Davey Smith and Ebrahim, 2003). Mendelian randomization studies can be used for two related purposes: (1) to provide evidence for the existence of causal associations, and (2) to enable accurate estimation of the magnitude of the effect of lifelong exposure to a risk factor on an outcome (Davey Smith and Ebrahim, 2004). As is the case for instrumental variable methods generally, for Mendelian randomization studies to be useful genetic variants must be robustly associated with the exposure of interest (Davey Smith and Ebrahim, 2005; Lawlor et al., 2008b). Despite this, recent Mendelian randomization studies conducted by Wehby et al. (2011a,b, 2012) have used genetic variants as instruments for smoking heaviness which were not shown to be associated with smoking phenotypes in large genome wide association studies. Whilst the authors acknowledge that these variants have not been consistently associated with smoking phenotypes, they suggest that the variants provide evidence of causal effects of smoking on body weight (Wehby et al., 2012) and smoking in pregnancy on birthweight (Wehby et al., 2011b) and risk of orofacial clefts in offspring (Wehby et al., 2011a). In addition, the authors use the genetic variants to estimate the magnitude of effect of smoking heaviness on their outcomes of interest (Wehby et al., 2011a,b, 2012). Even if the variants they use are truly associated with smoking behaviour, this is likely to produce incorrect estimates of the effect size of smoking on the outcome. Aims ~~~~ In this paper, we aim: (1) to illustrate, using a data simulation, why inferences based on the results of Mendelian randomization studies using genetic variants selected based on their association in a single sample are likely to be misleading and (2) to demonstrate why estimating the magnitudes of causal effects in cases where the measured exposure is not the same as the underlying exposure captured by the variant is problematic. We discuss these issues with reference to the specific case of tobacco as an exposure, but these principles can be applied more widely to Mendelian randomization and instrumental variable analyses. Assumptions of Mendelian randomization ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The principle of Mendelian randomization relies on the basic (but approximate) laws of Mendelian genetics (segregation and independent assortment). If these two laws hold, then at a population level, genetic variants will not be associated with the confounding factors that generally distort conventional observational studies (Davey Smith and Ebrahim, 2003; Davey Smith, 2011). In addition, genetic variants will not be affected by reverse causality (Davey Smith and Ebrahim, 2003). Epidemiological studies increasingly use Mendelian randomization to provide robust evidence of underlying causal mechanisms in a number of areas of health research including cardiovascular disease, cancer and mental health (Casas et al., 2005; Davey Smith et al., 2005; Benn et al., 2011; Scott et al., 2011; Interleukin-6 Receptor Mendelian Randomisation Analysis et al., 2012; Nordestgaard et al., 2012; Voight et al., 2012; Carslake et al., 2013). For a SNP to be a valid instrumental variable, the following assumptions must hold: (1) the SNP should be reliably associated with the exposure, (2) the SNP should only be associated with the outcome through the exposure of interest (the “exclusion restriction”) and (3) the SNP should be independent of other factors affecting the outcome (confounders) (Angrist et al., 1996; Lawlor et al., 2008b; Wehby et al., 2008; Clarke and Windmeijer, 2012). Moreover, to use Mendelian randomization for accurate estimation of effect sizes in mediation analysis using a measured exposure, the measured exposure should accurately capture the true causal exposure (Lawlor et al., 2008a; Pierce and VanderWeele, 2012). Genetic variants for tobacco research ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Large consortium-based genome wide association studies have found genetic variants robustly associated with smoking behaviours (Thorgeirsson et al., 2008; Furberg et al., 2010; Liu et al., 2010). One genetic variant that has been highlighted by these studies, amongst others, is located in the nicotinic receptor gene cluster CHRNA5–A3–B4 on chromosome 15. Two SNPs within this region, rs16969968 and rs1051730, which are in linkage disequilibrium and can be used interchangeably in studies on Europeans, consistently associate with measures of heaviness of smoking (e.g., cigarettes per day or biomarkers of nicotine exposure) (Freathy et al., 2009; Munafò et al., 2012). Smokers with a single copy of the smoking increasing allele smoke on average one extra cigarette per day compared to those with no copies. The effects of the SNP are additive, so people with two copies of the smoking increasing allele on average smoke two additional cigarettes a day (Ware et al., 2011). The strength and consistency of this association make these variants suitable instruments for use in Mendelian randomization studies. The second assumption of instrumental variable analysis, that the SNP should only be associated with the outcome through the exposure of interest, is rarely fully testable (Glymour et al., 2012). In Mendelian randomization, this assumption may be violated if the genetic variant has pleiotropic effects, is in linkage disequilibrium with another variant of differing function or if its effects are buffered by canalization (Davey Smith and Ebrahim, 2003). However, the biological function of the nicotinic receptor gene cluster and evidence from epidemiological studies suggest that this variant is likely to affect outcomes only through tobacco exposure (for a further discussion of this see Section 3). In addition, if the variant is associated with an outcome in smokers or former smokers but not never smokers, this is a good indication that the association is fully mediated through tobacco exposure (Freathy et al., 2011). The rs1051730 SNP has been used in Mendelian randomization studies to investigate the causal effect of cigarette smoking on body mass index, depression anxiety and birthweight of offspring (Freathy et al., 2011; Lewis et al., 2011; Bjorngaard et al., 2013; Tyrrell et al., 2012). Despite the identification of variants in the CHRNA5–A3–B4 gene cluster as suitable instruments, Wehby et al. (2011a,b, 2012) use other variants (in DRD2, MAOA, DRD4, 5HTT, GABBR2, CYP2D6) as instruments for smoking heaviness in their Mendelian randomization studies. The authors justify this approach by emphasizing the plausible biological roles of their chosen variants in smoking behaviour. However, this justification is questionable given that the candidate gene approach for finding functional genetic variants has had limited success, yielding few replicable associations and many false positives (Colhoun et al., 2003; Sleiman and Grant, 2010; Lawlor et al., 2008b). If these common variants are truly associated with the exposure, these associations should have been detected in the large genome wide association studies of smoking behaviour. We calculated that the largest of these studies, conducted by the TAG consortium, which included 74,000 smokers had 80% power to detect variants explaining as little as 0.05% of the variance in cigarettes per day (Furberg et al., 2010). Genetic variation in the CHRNA5–A3–B4 gene cluster explains about 1% of the variance in cigarettes per day (Munafò et al., 2012). Data simulation ~~~~~~~~~~~~~~~ We next expand our simulation to demonstrate how biases can occur if instruments are selected based on their observed associations with the exposure in the sample within which the Mendelian randomization experiment is being carried out. To simulate an example in which there is no effect of the exposure on the outcome, we set β1 = 0, so the outcome and exposure were only correlated (correlation = 0.6) due to the error terms. Thus the association of the exposure and outcome is confounded. This means that if our estimation model (estimator) is correct, then it should find no effect of the exposure on the outcome. If our estimator is incorrect and we find a relationship between the outcome and the exposure, then it suggests our estimator is biased. Next, to simulate the selection of genetic instruments within a sample, we randomly generated 1000 binary variables (Z) to simulate the SNPs (all had a frequency of 0.3). Since these instruments were randomly generated, there was no underlying effect of the SNPs on the exposure (α1 = 0). We used a binary instrument in a one instrument and one exposure example for simplicity, but these results are generalizable to additive genetic models or Mendelian randomization studies using multiple genetic variants (Pierce et al., 2011; Clarke and Windmeijer, 2012). We estimated the association of each SNP with the exposure, X, using robust linear regression. As expected, by chance, roughly 5% of these SNPs were associated with the exposure (using a p-value cut-off of 0.05). We selected the ten instruments most strongly associated with the exposure and ran a two stage least squares regression on the outcome using each of these instruments in turn. Table 3 presents the effect sizes and p-values for the association of the instrument with the exposure and the outcome along with the F-statistic (a measure of the strength of the association of instrument and exposure). Of the ten instruments selected, three had an F-statistic above the commonly used cut off point of 10, suggesting that the associations of instruments and exposure were strong enough for the instrumental variable estimates to be unbiased (Stock et al., 2002). Using two-stage least-squares regression, five of the instruments showed strong or moderate evidence for associations with the outcome (p values <0.01), and two further instruments were weakly associated (p values <0.1). However, we know that no “true” relationship exists, because of how we generated the data. Therefore, these instruments, and specifically how we selected the instruments biased the two-stage least squares estimates of the effect of the exposure on the outcome. Use of inappropriate genetic variants is not a problem specific to studies of tobacco research (Fletcher and Lehrer, 2011), but this example illustrates this problem well because of the availability of good instruments for smoking behaviour. The importance of this issue more generally in Mendelian randomization studies has been highlighted previously by Lawlor et al. (2008b) with reference to smoking- and obesity-related variants. The Beavis effect ~~~~~~~~~~~~~~~~~ Even when a variant discovered in a single sample is truly associated with an exposure, the effect sizes of variants identified within a single sample are, by the nature of their discovery, likely to be larger than in the overall population (the Beavis effect, or Winner's Curse) (Goring et al., 2001; Ioannidis, 2008; Burgess et al., 2011).","Mendelian randomization can provide very good estimates of the magnitude of effects of long term exposure to a risk factor on outcomes (Davey Smith and Ebrahim, 2005; Ference et al., 2012). However, when the phenotypic exposure of interest (e.g., cigarettes per day) does not adequately capture the “causal” exposure through which the genetic variant operates (e.g., lifetime exposure to tobacco), estimates from two-stage least-squares regression may be biased. In such cases, the second assumption of instrumental variable analysis (the exclusion restriction assumption) is violated. The genetic variant is still a valid instrument for the underlying phenotype of interest and can therefore still provide evidence of causality. However, it is not a valid instrument for the effect of the measured phenotype on the outcome and so magnitudes of effect are likely to be incorrect (Glymour et al., 2012). This principle also applies more widely to instrumental variable analyses using non genetic instruments, but this issue has not been well-developed in the econometrics or statistics literatures. In tobacco research, self-reported measures of smoking behaviour (such as number of cigarettes smoked per day) may be inadequate phenotypes because people smoke cigarettes differently. For example, there is variation in the number of puffs taken, volume of smoke inhaled or how far down the cigarette is smoked before it is discarded (Strasser et al., 2007; McNeill and Munafò, 2013). Objective measures of tobacco exposure (e.g., level of cotinine, the primary metabolite of nicotine) are likely to provide more valid assessment of actual biological exposure (i.e., the amount of smoked inhaled). For example, the rs1051370/rs16969968 variants are considerably more strongly associated with circulating levels of cotinine, than with self-reported daily cigarette consumption, explaining 4% and 1% of the variance in these phenotypes respectively (Keskitalo et al., 2009; Munafò et al., 2012). Researchers rarely have data on phenotypes such as cotinine, and often use a proxy measure such as self-reported cigarette smoking rates. This issue is illustrated in Fig. 1. We are particularly interested in the effect (a) of lifetime exposure to tobacco smoke (X) on an outcome measure (Y) (see Fig. 1A). Unfortunately, we may only have data on cigarettes smoked per day (X2), which is associated with but does not fully capture lifetime exposure (see Fig. 1B). The raw association of smoking on the outcome is confounded by the unobserved variable U (the error terms in our simulations). The genetic variant (Z), not only affects the total lifetime exposure (b), but also the number of cigarettes smoked (c). According to the second assumption of instrumental variable analysis, Z should only affect the outcome through its effect on the number of cigarettes smoked per day (X2) but in this case it also affects the outcome through lifetime exposure to tobacco smoke (X). In the example above, if we adjust the association of the variant (Z) with the outcome (Y) for the measured phenotype (X2) we would not expect the association to disappear because Z still affects Y through lifetime exposure to tobacco smoke (X). This issue has generated debate in the literature; the residual association observed between the CHRNA5–A3–B4 variants and lung cancer following adjustment for cigarettes per day has led to suggestions of a direct effect of the variant on lung cancer which does not operate though smoking (Lips et al., 2010; Wang et al., 2010a). However, Munafò et al. (2012) calculated that association between the variant and lung cancer was consistent with full mediation through tobacco exposure if cotinine were used as an intermediate measure of tobacco exposure rather than cigarettes per day. Therefore, the apparent direct association between these variants and lung cancer is likely to be a function of poor tobacco exposure measurement. This has important implications for the use of two-stage least-squares regression in Mendelian randomization analyses of smoking. If the measured exposure does not capture all dimensions of the relevant exposure domain, we can still infer a causal relationship, but cannot obtain an accurate estimate of the effect size of the underlying causal exposure. Thus the effect sizes presented in papers using cigarettes per day as the measured exposure of interest are likely to be subject to bias and should be interpreted with caution. It should be noted that this differs from the issue of random and systematic measurement error in the exposure phenotype, as discussed by Pierce and VanderWeele (2012). This is because even if cigarettes per day were measured perfectly, this phenotype would not adequately capture tobacco exposure. Whilst this is a particular issue for studies of tobacco use, this is also relevant for Mendelian randomization studies of other exposures. For example, estimates from Mendelian randomization studies using variants which affect caffeine consumption may be biased if the measured phenotype is number of cups of coffee consumed per day because this measure does not account for caffeine content of each cup. Glymour et al. (2012) also discuss this issue in relation to incorrect specification of the appropriate causal time period for an exposure, using body mass index as an example.","The results of Mendelian randomization studies, based on genetic variants chosen because of their association with the exposure in any one sample, do not contribute useful evidence of the effects of exposures on health outcomes. It is essential for Mendelian randomization studies to use genetic variants that are robustly associated with the exposure of interest. Fortunately, this is now possible for a number of exposures, including tobacco, generally because of variants identified in large genome wide association studies and replicated in independent samples (Timpson et al., 2005; Frayling et al., 2007; Hazra et al., 2008; Furberg et al., 2010; Wang et al., 2010b; Voight et al., 2012). Mendelian randomization studies, as well as establishing causal associations, can provide good estimates of the magnitudes of effect between exposures and outcomes as they are free from bias by confounding. However, estimates may be biased if the measured exposures are not the same underlying exposure as that represented by the genetic variant. Crucially, even if the underlying causal exposure is perfectly measured, if the variant additionally affects the outcome through a different pathway, neither causality nor strength of associations can be estimated. Mendelian randomization has the potential to be a valuable tool to further our understanding of the aetiology of disease. Researchers will only realize this potential if they base their studies on well-characterized variants and are cautious about making inferences about magnitudes of the relationships between observed phenotypes and outcomes.","Amy Taylor, Jennifer Ware and Marcus Munafò are members of the UK Centre for Tobacco and Alcohol Studies, a UKCRC Public Health Research: Centre of Excellence. Funding from British Heart Foundation, Cancer Research UK, Economic and Social Research Council, Medical Research Council, and the National Institute for Health Research, under the auspices of the UK Clinical Research Collaboration, is gratefully acknowledged. This work was supported by the Wellcome Trust (grant number 086684) and the Medical Research Council (grant numbers MR/J01351X/1, G0800612, G0802736, G0600705, MC_UU_12013/1-9). George Davey Smith and Neil Davies are supported by the European Research Council DEVHEALTH grant (269874). Jennifer Ware is supported by a Post-Doctoral Research Fellowship from the Oak Foundation. Tyler VanderWeele is supported by an NIH grant (R01 ES017876)."],["This paper proposes a hybrid monetary model of the dollar-yen exchange rate that takes into account factors affecting the conventional monetary model's building blocks. In particular, the hybrid monetary model is based on the incorporation of real stock prices to enhance money demand stability and also, productivity differential, relative government spending, and real oil price to explain real exchange rate persistence. By using quarterly data over a period of high international capital mobility and volatility (1980:01-2009:04), the results show that the proposed hybrid model provides a coherent long-run relation to explain the dollar-yen exchange rate as opposed to the conventional monetary model. © 2014. --------------------------------------------------------------------------------","Since the collapse of the Bretton Woods fixed exchange rate system in 1971, much attention has been paid towards finding a meaningful explanation of exchange rates. A wide range of models have been proposed to understand movements in the exchange rate, one of which is the monetary model (see Bilson, 1978; Frankel, 1979). Despite its rigorous theoretical underpinnings by linking the nominal exchange rate to its monetary fundamentals (e.g., money, income, and interest rates), the resulting reduced form has had limited empirical success until now. For example, although MacDonald and Taylor (1994) provided evidence of a long-run relation between monetary fundamentals and nominal exchange rates, the signs and magnitudes of estimated coefficients did not support the related monetary theories. Groen (2000), and Mark and Sul (2001) among others also found some evidence in a panel context, but this was under the assumption of a high order of heterogeneity across the country models. Similarly, Rapach and Wohar (2002) found some support for the theory using long time series, but this was related to different exchange rates and macro regimes, with some evolution in the composition of products in price indices. Taylor and Peel (2000) applied nonlinear methods to model a nominal exchange rate and monetary fundamentals (relative money supply and relative income), but such results are often sensitive to a small number of observations and become less robust as the sample evolves. Frömmel et al. (2005) estimated the real interest differential (RID) model of Frankel (1979) applying the Markov switching approach. However, the model was shown to relate to only one regime. Furthermore, the empirical failure of this model has been specifically found in regard of the US dollar–Japanese yen exchange rate. The evolution of this exchange rate has been much debated over the recent years with no consensus over the factors that drive the dynamics. For instance, Caporale and Pittis (2001) were unable to find a stable relation based on a monetary model of this exchange rate. Chinn and Moore (2011) also failed to find a long-run relation between the nominal dollar–yen exchange rate and its monetary fundamentals (money, industrial production, and interest rate differentials) even when they included cumulative order flow as opposed to the dollar–euro exchange rate. By contrast, MacDonald and Nagayasu (1998) only found that a simplified version of the RID model of Frankel (1979), that excluded the money demand functions, held for the yen–dollar exchange rate for the period 1975:Q3–1994:Q3. Tellingly, in a recent paper, Obstfeld (2009, p.1) comments that ‘the determinants of the yen's short- and even longer-term movements remain mysterious in light of the development of Japan's macro economy’. A possible explanation for the empirical failure of the dollar–yen exchange rate monetary model is perhaps the breakdown of its underlying building blocks; that is, stable money demand and purchasing power parity (PPP). Indeed, Hendry and Ericsson (1991) found that the conventional money demand equation for the US was not stable. Whereas, Friedman (1988) and McCornac (1991) confirmed the need for real stock prices to stabilise money demand equations using data from the United States and Japan, respectively.1 Sarno and Taylor (2002), on the other hand, found little support for the conventional notion of PPP by surveying a range of empirical studies. This corresponds well with the classic findings of Balassa (1964) and Samuelson (1964), which indicate that persistent deviations from PPP arise from productivity differentials. Chinn (1997, 2000), and Wang and Dunne (2003) among others showed that fluctuations in the nominal and real dollar–yen exchange rate are due to the impact of differentials in productivity and government expenditure along with real oil prices. This paper contributes to the existing literature by proposing a hybrid monetary model of the dollar–yen exchange rate that takes into account the breakdown of the aforementioned building blocks. That is, the proposed model captures both the monetary and the real aspects of the economy, thereby circumventing some of the potential pitfalls associated with earlier studies. More specifically, we examine the empirical performance of the standard RID model, developed by Frankel (1979), against this proposed hybrid version by employing the Johansen (1995) methodology and quarterly data from 1980:01 to 2009:04, a period characterised by high international capital mobility and volatility. The RID model has been widely used as it combines aspects of the sticky-price approach with the flexible-price one. Furthermore, this variant of the monetary approach is chosen because it is a realistic description when variation in the inflation differential is moderate as is the case between the US and Japan over the period under examination.2 Particularly, the theory underlines the role of expectations in different inflationary environments and the associated rapid adjustment in capital markets. The hybrid version, by contrast, is devised by using domestic and foreign money demand equations based on broader asset classes and also accounting for the factors that cause PPP to fail. That is, we incorporate real stock prices in the money demand equations,3 while we use the productivity differential, relative government spending, and real oil price to explain the persistence in the real dollar–yen exchange rate. The paper is organised as follows. Section 2 provides the theoretical framework for the exchange rate monetary model; Section 3 outlines the econometric technique used and describes the data; Section 4 explains the empirical results and the analysis; and finally Section 5 concludes.","The monetary model of the exchange rate is based on the assumptions that money demand equations are stable and that PPP holds. In this paper, we consider two forms of this model and place them under econometric scrutiny. The first is the RID model developed by Frankel (1979) and the second is a hybrid monetary model, proposed herein, that takes into account factors affecting the stability of the respective money demand equations and the validity of PPP. Otherwise, the RID model related to Eq. (8) hypothesises that an increase in the domestic money supply relative to the counterpart foreign one increases domestic prices and thus causes a one for one depreciation in the exchange rate (β1 = 1). An increase in domestic income or a decline in the expected rate of domestic inflation (proxied by the long-term interest rate) relative to the foreign one raises the demand for money and thus causes an appreciation in the exchange rate (β2 < 0, β4 > 0). An increase in the domestic nominal interest rate relative to the foreign one induces capital inflows towards the domestic economy and thus causes an appreciation in the exchange rate (β3 < 0). For further details the reader is directed to Frankel (1979). However, Friedman (1988) and subsequently McCornac (1991) and Caruso (2001) among others showed that the stability of the money demand functions used to specify the monetary model, Eqs. (3) and (4), depends on the inclusion of real stock prices. Furthermore, as Chortareas and Kapetanios (2004) pointed out, there is limited support for the conventional notion of PPP for Japan. Indeed, by visual inspection of Fig. 1, we find that the real dollar–yen exchange rate, calculated as the nominal exchange rate adjusted for the domestic and foreign price levels (see Eq. (6a) in Appendix A), does not appear to revert to mean. Balassa (1964) and Samuelson (1964) attributed the inadequacy of PPP to real economic shocks, in particular, to the unanticipated movement found in the productivity differentials between the traded and non-traded goods sectors across the economies. Financial variables also appear sensitive to the demand shocks associated with government expenditure (Chinn, 2000) and the supply shocks related to real oil prices (Amano and van Norden, 1998). In the context of the yen–dollar exchange rate, Chinn (1997, 2000), and Wang and Dunne (2003) found that real economic factors were responsible for any persistence in the real yen–dollar exchange rate during the post-Bretton Woods period. In addition to the coefficient restrictions discussed earlier (β1 = 1, β2 < 0, β3 < 0, β4 > 0), the HM model suggests that the sign of the coefficient on real stock prices, β5, depends on the extent to which the substitution effect (positive) dominates the wealth effect (negative) in the money demand equation. Based on the derivation provided in Appendix A, the sign of the coefficient on the productivity differential depends on the relative competitiveness of the traded goods sector. Specifically, an increase in the productivity of the traded sector relative to the non-traded sector in the domestic economy compared to the foreign one results in a fall in the domestic traded sector's goods prices relative to the foreign counterpart, and then an exchange rate appreciation (β6 < 0). The differential in government expenditure captures differences in demand side shocks (Chinn, 2000). As government expenditure is anticipated to be spent largely on non-tradable goods such as services, an increase in domestic government spending relative to the foreign counterpart should then increase the relative price of domestic non-tradable goods, leading to an exchange rate appreciation (β7 < 0). The sign of the coefficient on the real oil price is expected to be negative (β8 < 0) because oil price is given in the US dollar and higher real oil price should lead to an appreciation in the dollar (see Amano and van Norden, 1998). That is, the input costs in Japan are highly sensitive to the oil price because Japan is a net importer country and the third largest oil consumer and importer country after the United States and China.4 The econometric approach ~~~~~~~~~~~~~~~~~~~~~~~~ The results associated with the Johansen test are well-defined when the VAR model is well- specified (Johansen, 1995). The most appropriate lag length for the model is often selected on the basis of information criteria such as the Schwarz Bayesian information criterion (SBIC), the Akaike information criterion (AIC), and the Hannan–Quinn information criterion (HQIC). However, Burke and Hunter (2007) suggest that there can be substantial size distortion of the trace test relative to the null distribution when the selected lag order is sub-optimal.5 Therefore, we extend the model to include adequate lags to remove any serial correlation in case the lag selected based on information criteria does not capture the dynamics. As a result of sharp changes as well as differences in monetary policy between the United States and Japan throughout the sample period, we also include impulse dummies that remove the impact of extreme observations relating to 1980:4, 1982:3, 2002:2, and 2008:4. The corresponding known events for the first two dummies relate to the large short-term interest rate fluctuations in the United States and Japan in the late 1970s and early 1980s. Note that the fourth quarter of 1980 also corresponds with the end point of the fiscally liberal 60s and 70s that led to the election of Ronald Reagan as the US President and the Volker reforms at the Federal Reserve. The third dummy corresponds to the monetary expansions (now termed quantitative easing (QE)) adopted by the Bank of Japan from March 2001 to March 2003, while the fourth is due to QE in the United States as a result of the 2007–2008 banking crisis. We suggest that by investigating these two variable sets, we might be able to determine the key factors that identify the long-run monetary model of the dollar–yen exchange rate and explain the short-run behaviour of the different systems. Using series that are I(1), we can observe an exchange rate equation based on the model in question by finding a cointegrating relation and showing via a likelihood ratio test that this variable is neither long-run excluded (Juselius, 1995), nor weakly exogenous (Johansen, 1992).6 Data ~~~~ For this paper, we use quarterly seasonally unadjusted data, where available, for the United States vis-à-vis Japan over the period 1980:1–2009:4. We choose the start of the sample period in order to control for structural change in the Japanese financial system because by the end of 1979, the interbank rates in Japan were deregulated, capital controls were removed, and the certificate of deposit market developed (McCornac, 1991). We use quarterly data as GDP data are not available on a monthly basis. The short-term interest rates are represented by the official discount rates,7 whereas the long-term interest rates are represented by the 10-year government bond yields. Moreover, we use the consumer price index (CPI) to deflate the stock price indices represented by the S&P 500 in the United States and the Nikkei 225 in Japan. While government spending is defined as government consumption in proportion to GDP, the productivity is defined as industrial production divided by the corresponding employment level. The real oil price is the West Texas Intermediate (WTI) Cushing crude oil spot price (in dollars per barrel) deflated by the US CPI. The exchange rate (denoted as dollars per unit of yen), interest rates, national income, industrial production, and price levels (CPI) are sourced from the IMF's International Financial Statistics (IFS). Nonetheless, money supply (M1), oil price, and stock prices are from Thomson DataStream.8 Government spending and employment figures, on the other hand, are obtained from the OECD main economic indicators (MEI).","A prerequisite for conducting cointegration tests is to check the time series properties of the variables under investigation as to their order of integration. The null of non- stationarity is tested using augmented Dicky–Fuller (ADF) tests (Dickey and Fuller, 1981) and DF-GLS tests of Elliott et al. (1996). The results, as displayed in Tables 1 and 2, indicate that all the variables require first differencing to be stationary, hence they are integrated of order one (I(1)).9 Cointegration is then tested using the Johansen (1995) procedure. The first subsection presents the analysis of the RID model; the second subsection analyses the HM model; and finally validation of the hybrid model is reported in the third subsection. Long-run analysis of the RID model ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ For the data set XRID,t, the SBIC, HQIC, and AIC indicate that the VAR is first order (p = 1). However, in order to remove any serial correlation and enhance the specification of the model, we require p = 4. The Lagrange multiplier (LM) test for the presence of serial correlation and ARCH along with the Jarque–Bera test of non-normality is reported in Table 3 and suggests that, at the 5% level, there is no evidence of misspecification for the model. On the basis of this specification, the estimated eigenvalues and trace statistics are reported in Table 4. By inspection of the above results the estimated coefficient on the relative money supply has the sign expected by theory, even though it is large relative to the hypothesised magnitude of 1. Moreover, based on one-sided inference, we consider it significant at the 5% level. However, the coefficients on the rest of the monetary fundamentals have signs that are not consistent with the theory, although relative income and the long-term interest rate differential are highly significant (at the 1% level). To provide further insights into this long-run relation, we conduct long- run exclusion (LE), weak exogeneity (WE), and stationarity tests by imposing restrictions on α and β (see Johansen, 1995); the latter tests are conducted to provide further evidence with regard to the stochastic properties of the series. Even though it is often felt that the normalisation is innocuous, the significance of the LE test is informative of the likely appropriateness of a normalisation. According to Boswijk (1996), empirical identification generally requires satisfaction of further rank conditions. However, Burke and Hunter (2005; chapter 5) argue that any coherent strategy for identification ought to preclude normalisation on variables that are either long-run excluded or weakly exogenous. Cointegration is a property of two or more non-stationary series and thus normalisation is also inappropriate on stationary variables. The tests of LE, WE, and stationarity are asymptotically distributed chi-squared (Johansen, 1992) and in Table 5 we report our results on a variable by variable basis. The LE tests are conducted by imposing a zero restriction on the relevant elements of β. If a zero restriction on an element of β for a specific variable is not rejected, then the long-run relation cannot be normalised on this variable. The WE tests, by contrast, are carried out by imposing a zero restriction on elements of α in turn. If a zero restriction on an element of α for a particular variable is not rejected, then this variable can be considered weakly exogenous; it drives the system instead of adjusting to it. The stationarity tests are conducted, under the null hypothesis of stationarity, in the multivariate setting by fixing each element in turn in a single cointegrating vector to unity and the remaining elements to zero. As is evident from Table 5, the LE tests indicate that, except for the relative income and long-term interest rate differential, all the other variables can be excluded from the cointegrating relation. Hence, any long-run model based on the exchange rate may be ill defined, as the related parameter cannot be distinguished from zero. In the subsequent panel, the proposition that the exchange rate and short-term interest rate differential are weakly exogenous also cannot be rejected. This implies that, at best, the long-run relation ought to be conditioned on the exchange rate instead of being normalised on. Hunter (1992) among others presents similar findings for the exchange rate. In conclusion, despite the existence of a long-run relation among the variables of the RID model, such a relation cannot explain the behaviour of the exchange rate as this variable can be both excluded and viewed as weakly exogenous for the cointegrating vector. However, the tests of stationarity following from the restriction mentioned before on the VAR support the proposition that the series are all difference stationary (see Table 5), in line with the results of single unit root tests in Tables 1 and 2. Long-run analysis of the HM model ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The findings given above cast serious doubt on the conventional monetary model regarding the dollar–yen exchange rate. Therefore, we consider it of paramount interest to investigate the reasons for this failure. To this end, the VAR model is now based on the vector XHM that represents the hybrid version. Since the price of oil is a global factor and all other factors are differentials between the United States and Japanese variables, we treat the real oil price as exogenous to the system.10 Indeed, the test suggests a non- rejection of the null hypothesis of weak exogeneity with a p-value of 0.741. This finding is also consistent with the intuition of Amano and van Norden (1998) that oil prices in the decades preceding their study were governed by the major supply-side shocks resulting from political instability in the Middle East, and are thus external to the developed economies. With regard to the VAR specification, the SBIC, HQIC and AIC suggest a lag length p = 1, while diagnostic tests imply that p = 3 is required to improve the specification. The reported diagnostics in Table 6 suggest that, at the 5% level, the model does not suffer from serial correlation using the LM test up to order 8, and the same applies for ARCH effects up to order 8. However, the multivariate normality test is rejected, where the sources of such failure seem to result from excess kurtosis in the money supply and productivity differentials. Since Gonzalo (1994) demonstrated a lack of sensitivity of the cointegrating rank to excess kurtosis, we conclude that these findings are robust. Accordingly, Table 7 reports the trace test related to the HM model. It is evident that the null hypothesis of no cointegration is rejected, but evidence for more than one cointegrating vector cannot be rejected at the 5% level. Since the cointegrating rank does not change by the inclusion of the augmenting factors, this indicates that these factors follow stochastic trends common to the nominal exchange rate and its monetary fundamentals in the RID model. Long-run exclusion tests are likely to give more information regarding the nature of the contribution of the augmenting factors and also the variables on which the long-run relation may be normalised. Hence, Table 8 reports the LE, WE, and stationarity tests of the variables included in the HM model. The stationarity tests imply that none of the variables in the cointegrating relation is stationary, consistent with the results reported in Tables 1 and 2. The LE tests, by contrast, indicate that the real oil price is the primary candidate for exclusion in the long-run relation, while the money supply and real stock price differentials could be excluded on a single-variable basis, although this would be rejected at the 15% level. However, at this stage, we do not exclude any variable based on a single-variable test. In the next subsection, we use these results to obtain a more parsimonious long-run relation. Our key findings are that the long-run exclusion of the nominal exchange rate is rejected now, and that the exchange rate appears not to be weakly exogenous for the HM model (see panel B in Table 8). The change in WE status is a de facto indication of changes in long-run feedback and is of paramount interest (Juselius and Macdonald, 2004). Unlike the RID model, this finding indicates that the nominal exchange rate in the HM model adjusts to the long-run equilibrium. That is, it does not force the system when such a system accounts for the relative real stock prices, the productivity differential, relative government spending, and the real oil price. In addition to the real oil price on which the system is conditioned as stated earlier, the tests reported in Table 8 (panel B) also indicate that we cannot reject the findings that relative money supply, relative income, short-term interest rate differential, relative real stock prices, productivity differential, and relative government spending are weakly exogenous at the 5% level for the long-run relation, although the long-term interest rate differential is not. As shown from Eq. (13), the estimated coefficients on monetary fundamentals (relative money supply, relative income, and short-term and long-term interest rate differentials) are all significant and consistent with monetary theory. More specifically, the coefficient on the relative money supply is not materially different from 1, and is significant based on a one-sided test at the 5% level. All the other monetary variable coefficients (relative income and short-term and long-term interest rate differentials) have their hypothesised signs and are significant at the 1% level. Furthermore, as hypothesised by Frankel (1979), the parameter on the long-term interest rate differential is greater than that on the short-term interest rate differential in absolute value. Except for the real oil price, all the factors that have been used to augment the monetary model have significant parameters. This implies that the real oil price can be excluded from the long-run relation and, as with Johansen and Juselius (1992), treated as strictly exogenous. Consistent with Friedman (1988) and Caruso (2001), the coefficient on the relative real stock price is negative implying that the wealth effect dominates the substitution effect in the underlying money demand functions for the United States and Japan. The coefficients on the productivity differential across the industrial sectors and relative government spending are negative and significant. This suggests that higher domestic productivity or government spending compared to their foreign counterpart results in an exchange rate appreciation. The fact that the real oil price can be excluded from the long-run part of the VAR system suggests that it affects the long-run only indirectly by enhancing the econometric performance of the model. Hybrid model validation ~~~~~~~~~~~~~~~~~~~~~~~ The above results strongly indicate that the HM model dominates the RID model on theoretical and econometric grounds in explaining the dollar–yen exchange rate in the long-run. However, to check the robustness of our results, we conduct two further analyses. First, we use the results on LE and WE to obtain a more specific and robust formulation of the long-run relation based on the HM model. A similar approach has also been used by MacDonald and Nagayasu (1998), though they examine a simplified version of the RID model of the yen–dollar exchange rate excluding money demand functions. Next, we sequentially impose zero restrictions on the loading factors, α, of the standard monetary fundamentals related to relative money supply, relative income, and short-term interest rate differential. These weak exogeneity restrictions are empirically plausible given the size of the adjustment coefficients and also consistent with monetary theory. The tests, as displayed in Table 9, indicate that the imposed restrictions cannot be rejected. Moreover, the constrained final long-run relation normalised on the exchange rate suggests the significance of monetary variables with their hypothesised signs, as found in the previous subsection. The trace test also implies that there is still a single cointegrating vector among the variables (these results are unreported). Overall, this demonstrates the robustness of our results in terms of the long-run formulation and direct impact of the augmenting factors on the long-run exchange rate monetary model.11 Then, we subject our proposed HM model to an array of forward and backward recursive stability tests proposed by Hansen and Johansen (1999) to gain further insights into its adequacy as a long-run exchange rate model. The results reported here relate to the behaviour of the max tests of β and are displayed in Fig. 2. The forward and backward tests appear respectively in the Figure's left and right panels, with the corresponding 5% critical value represented by the solid line. Note that in providing these stability tests, the short-run effects (X(t)) compared to those of the long-run R1(t) are concentrated out. In a broad sense, the model shows a reasonable degree of stability of the parameters in the cointegrating vector. Hence, the model seems to be adequate and does not exhibit structural breaks in relation to the long-run for the period under observation.","In this paper, we re-examine the dollar–yen exchange rate using two versions of the monetary model. The first is the conventional real interest differential (RID) model of Frankel (1979) and the second is a hybrid monetary (HM) model, proposed herein, that incorporates on the one hand real stock prices to capture the stability of money demand and on the other, the productivity differential, relative government spending, and the real oil price to explain the persistence in the real exchange rate. Both models are estimated using the Johansen cointegration methodology and quarterly data from 1980 to 2009, a period characterised by high international capital mobility, as well as periodic volatility in the dollar–yen exchange rate. Although a single cointegrating vector exists for both models, the long-run exclusion and weak exogeneity tests inform us that the HM version gives an appropriate long-run explanation of the monetary model of the dollar–yen exchange rate. The enhanced performance of the HM model derives from the following considerations to the conventional monetary model. First, the stability of money demand relations is taken into account by the inclusion of key variables that impact on transactions (Friedman, 1988). A key feature of globalised financial markets is a highly active market in cross-border investments, mergers and acquisitions, and cross-listed stocks. In particular, the futures contract on the Nikkei is listed as an asset in the US stock market. Second, the persistence of the real exchange rate, which reflects primarily the impact of the non-traded goods, is taken into consideration by accounting for productivity and government expenditure differences. In essence, these differences may be due to the relatively insular nature of Japanese society limiting the effectiveness of arbitrage. The literature also suggests that the real oil price affects such persistence, but the empirical findings herein show an indirect impact of such a price via the dynamic specification of the VAR model. Contrary to the conventional monetary model, the results also suggest that the dollar–yen exchange rate in the hybrid model is driven by money, income, and short-term interest rate differentials, but not the reverse. This implies a substantial role for real economic and financial market variables in a well-formulated monetary model for the determination of the long-run exchange rate."],["Reservation price equilibria (RPE) do not accurately assess market power in consumer search markets. In most search markets, consumers do not know important elements of the environment in which they search (such as, for example, firms' cost). We argue that when consumers learn when searching, RPE suffer from theoretical issues, such as non-existence and critical dependence on specific out-of-equilibrium beliefs. We characterize equilibria where consumers rationally choose search strategies that are not characterized by a reservation price. Non-RPE always exist and do not depend on specific out-of-equilibrium beliefs. Non-RPE have active consumer search and are consistent with recent empirical findings. --------------------------------------------------------------------------------","In consumer search markets, firms have market power due to the fact that some consumers do not make price comparisons. Firms take this market power into account when deciding on price. This paper addresses the question of whether, by focusing on consumers following reservation price strategies, the existing consumer search literature accurately evaluates this market power due to search frictions. A reservation price strategy is a cut-off strategy: after observing a price at or below some critical value, consumers decide to buy, otherwise they continue to search. In markets where there is uncertainty about the underlying factors determining firms' pricing behavior, there are important theoretical reasons to consider other search strategies than reservation price strategies. Rothschild (1974) drew attention to the fact that when consumers do not know from which distribution of offers they obtain their information, the optimal consumer search rule may well be different from the typical reservation price rule.1 The main reason is that on the basis of past search observations, consumers learn about the environment in which they search. Depending on the environment, it may well be that, after observing a relatively good outcome, consumers infer that even better outcomes are likely to be observed in the next search round and rationally conclude to continue to search, whereas, after observing a relatively bad outcome, consumers infer that better outcomes are unlikely and thus stop searching. The consumer search literature has, by and large, neglected this observation. The celebrated models by Stahl (1989) and Wolinsky (1986), and much of the literature that takes these models as a starting point, study environments without underlying uncertainty and in such theoretical environments the optimal search rule is indeed a reservation price rule. In consumer search markets where consumers are uninformed about firms' underlying costs (and this probably comprises most markets where consumer search is important), learning is an important part of the search process. There are some papers on learning and consumer search that take consumer uncertainty about firms' costs into consideration (see, Benabou and Gertner, 1993; Dana, 1994; Fishman, 1996 and more recently, Yang and Ye, 2008; Tappata, 2009; Janssen et al., 2011 and Chandra and Tappata, 2011). The observations by Rothschild (1974) are of immediate concern to these environments, but the relevant economics literature has continued to focus on equilibria where the consumer search rule is characterized by a reservation price. Some of this literature is inspired by retail gasoline markets where the common wholesale price of crude oil is the most important determinant of the (variation in) costs of retailers, and consumers are uncertain about these costs due to the large fluctuations of this wholesale price on the world market. Although our focus in this paper is on consumer search in retail markets, the issues we address are also relevant for other markets. For example, Benabou and Gertner (1993) is motivated by macroeconomic concerns about inflationary uncertainty and the consequences for firms' mark-ups, while a recent paper by Duffie et al. (2017) considers over-the- counter (OTC) financial markets and the role of benchmarks in these markets. The current paper is also relevant for the labor search literature where workers search for a better wage. In labor markets, it is natural that the wage distribution depends on the business cycle and that firms are better informed about the business cycle than workers. In that case, workers learn about the wage distribution while searching for another job and their search behavior does not need to follow a reservation wage strategy. In all these markets, there is uncertainty and asymmetric information about a common component that determines the distribution of offers and one needs to understand how the uninformed side (consumers, workers) search and simultaneously learn in such an environment. Our paper is the first to systematically incorporate Rothschild's observations on non-reservation price strategies into an equilibrium search model with endogenous firm behavior.2,3 Benabou and Gertner (1993) also mention the fact that in their model reservation price equilibria (RPE) may not exist. They set up the equations that have to be satisfied in a non-RPE. They perform some numerical analysis for some parameter values, but they neither have an analysis characterizing these non-RPE, nor do they show the conditions under which these equilibria exist.4 The literature studying RPE in environments where consumers are uninformed about firms' cost is unsatisfactory for a number of reasons. First, RPE are known to exist only if the search cost is relatively large and/or the uncertainty about costs is relatively small (cf., Dana, 1994 and Janssen et al., 2011). It is unclear what type of equilibria do exist for small search cost or large uncertainty about common costs. Second, RPE implicitly assume certain out-of-equilibrium beliefs and it is unknown whether these out- of-equilibrium beliefs satisfy game theoretic refinement concepts commonly employed in asymmetric information games. Third, one would expect that when costs are uncertain consumers may engage in active costly search in equilibrium. When consumers observe a high price, they are uncertain about whether this is due to a relatively high (common) production costs or whether this particular firm is charging a high margin. RPE in these homogeneous goods markets have firms charging prices below the consumer reservation price, however, and therefore all consumers buy at the first firm they visit. This lack of consumer search gives firms substantial market power, but it may well be that RPE overestimate the true market power because they underestimate consumers' search intensity. In response to these points, this paper first sharpens existing results on RPE. We show (i) that independent of out-of-equilibrium beliefs, RPE do not exist when the uncertainty about production cost is relatively large and (ii) that an RPE, even if it exists, is sensitive to the specification of out-of-equilibrium beliefs and do not satisfy, for example, the logic of the D1 equilibrium refinement (hereafter the D1 logic, see Cho and Sobel, 1990). If the uncertainty about cost is relatively large, any equilibrium should have active search. We then continue to characterize non-RPE and show that they exist for all parameter values and that there are parameter values for which multiple non-RPE exist. Thus, non-RPE resolve the non-existence problem that RPE suffer from. Moreover, in any non-RPE, consumers actively search beyond the first firm. In particular, there is a region of “high” prices that are set with positive probability such that consumers are indifferent between buying and searching and consumers continue to search with strictly positive probability. When the cost uncertainty is large, market prices may be substantially below the market prices predicted by RPE due to active search by consumers. On the other hand, when cost uncertainty is small, expected market prices are larger in non-RPE. Thus, whether or not RPE overestimate the market power of firms depends on the uncertainty about cost. In a recent empirical paper, De los Santos et al. (2012) show that, when buying books online consumers do not follow reservation price strategies. These strategies predict that (i) consumers buy from the last store visited unless all stores have been visited and (ii) the decision whether to continue to search depends on the outcome of the previous search with consumers observing lower prices deciding to buy, and consumers observing larger prices deciding to continue searching. Their evidence contradicts these predictions.5 In this paper, we show that their empirical findings are consistent with equilibrium behavior under non-reservation price strategies, as follows. When the cost uncertainty is relatively large, non-RPE have a region of intermediate prices where the probability of a sale is lower when the price is low. At lower prices, consumers rationally expect to get lower prices on the next search round and this may induce them to search more. In particular, we show that consumers may accept higher prices in the first search round, while rejecting lower prices. In an extension to oligopolistic markets, we also show that the optimal sequential search behavior of consumers is consistent with consumers going back to previously sampled firms, before they have sampled all firms.6 There is also a relationship with the marketing oriented literature on reference price effects (see, e.g., Putler, 1992; Kalyanaram and Winer, 1995 and Mazumdar et al., 2005). This literature points to the fact that consumers have particular pricing points around which consumer demand is very sensitive to price changes. This may lead to situations where consumer demand drops significantly if firms price above this reference point, whereas at higher prices, consumers are willing to buy again. Such “reference point” demand behavior can occur in non-RPE when the cost uncertainty is large. After observing intermediate prices above the “reference price”, consumers rationally infer that these prices are not chosen by high cost firms. Knowing costs are low, consumers find these prices too high to buy, however, and they continue to search for sure. This inference creates a gap in the equilibrium price distribution of the low cost firms. The rest of the paper is organized as follows. Section 2 describes the model and the equilibrium concept. Section 3 discusses how RPE depend on assumptions regarding out-of- equilibrium beliefs and why (regardless of these out-of-equilibrium beliefs) they do not exist when cost uncertainty is relatively large. Section 4 describes our analytical results on non-RPE. Section 5 shows, by means of a numerical analysis, the effects of cost uncertainty on profits, expected prices and consumer welfare. Section 6 briefly discusses a generalization of our model to the case of imperfectly correlated production costs and oligopoly markets with N firms. We show that with three or more firms, the optimal search rule may imply that consumers first continue searching another firm, and then go back to a previously sampled firm before all firms are sampled. Section 7 concludes with a discussion, while proofs are given in two Appendices.","The timing of the model is as follows. First, Nature chooses c for both firms. After observing c, firms simultaneously decide on their prices. Finally, consumers search and make their purchase decisions. It is by now a standard argument in the search literature with symmetric information that due to the presence of shoppers and non-shoppers there do not exist equilibria with mass points in the price distributions. This argument continues to hold in our model with asymmetric information as far as pure pricing strategies are concerned: even if all non-shoppers continue to search, an undercutting firm will sell to all shoppers and non-shoppers that first visit that firm. In the present model, this argument does not extend, however, to ruling out pricing distribution with mass points. Given that equilibria have to be in mixed strategies and that we prove that equilibria without mass points always exist, we restrict our attention to mixed pricing strategies without atoms. As explained in the Introduction, the existing literature focuses on reservation price equilibria, which are defined as follows. When investigating non-RPE, we focus on equilibria satisfying the logic of the D1 criterion (Cho and Sobel, 1990).10 The D1 criterion was developed in the context of pure signaling games with one sender. Our model is a two-sender game, where the beliefs of the receivers (the non-shoppers) are only based on the single price they have observed. As firms are of the same type, the out-of- equilibrium belief of non-shoppers is simply a mapping from the observed price to the type distribution of cost, as in the one-sender game. Thus, from the set of equilibria that do not satisfy Definition 1, we focus on those equilibria that satisfy the D1 logic and that are sufficiently smooth. For easy reference, we refer to such equilibria in the rest of the paper as non-reservation price equilibria (non-RPE). There may exist other equilibria that do not satisfy the properties of a RPE and that do not satisfy the D1 requirement and are not smooth. We do not consider these equilibria in our paper as we show that even with the additional requirements of D1 and smoothness, we can guarantee existence. Moreover, the equilibria we consider are interesting in their own right and do not depend on arbitrary out-of-equilibrium beliefs or on more technical issues related to non- smoothness.","In this Section, we summarize some existing results on RPE, (i) prove that they do not exist if the cost uncertainty is large and (ii) prove that they do not satisfy the D1 logic even if they do exist. We next show that if an RPE does not exist, any equilibrium without mass points should have a region of prices where non-shoppers actively search (search with positive probability). There are two important corollaries, which follow immediately from Proposition 3. First, as in any RPE firms' pricing distributions are atomless (see, e.g. Stahl, 1989), and there is no active search (see Dana, 1994), we immediately have the following corollary. All reservation price equilibria do not satisfy the D1 logic. Second, by Definition 2 we have In any non-reservation price equilibrium non-shoppers search with positive probability.","The proof is constructive and consists of several Lemmas. It is formally developed in Appendix B. For a range of parameter values the equilibrium is not unique, while for other parameter values it is unique. We now extensively describe how for any set of parameter values we construct an equilibrium. We end this Section discussing how our analysis may shed some light on the empirical observations we mentioned at the end of the Introduction. Numerically, one can compare for a given cost realization (i) the expected first price observation conditional on the price being accepted and (ii) the expected first price observation conditional on it not being accepted. De los Santos et al. (2012) observe that in their sample the first conditional expected price is larger than the second, and they rightly claim that this is inconsistent with RPE. For the parameter values used in Figs. 3–5, one can compute and compare both conditional expected prices to conclude that for the high cost realization these non-RPE are consistent with the findings of De los Santos et al. (2012): for Fig. 4 the respective numbers are 42.94 and 42.83, for Fig. 3 the numbers are 50.81 and 49.83, while for Fig. 5 they are 45.19 and 45.11, respectively. Thus, we conclude that the observations of De los Santos et al. (2012) are not necessarily inconsistent with sequential search, although they are inconsistent with reservation price strategies. Finally, equilibria where the low cost price distribution has a non-convex support may be interpreted as a search theoretic foundation for the reference price principle that is discussed in marketing (see the references in the Introduction). In our model, reference prices endogenously arise from the fact that consumers rationally infer that a certain low price will only be set when cost is low, and if the common cost is really low, then the chances of finding low prices are so high that it is rational to continue searching for better deals. Thus, it is better for firms not to set prices just above these reference prices.","We are now in a position to compare the equilibrium outcomes of our model with two benchmark models, and to perform some numerical comparative statics analysis. On one hand, we use Stahl (1989) as a benchmark to show the implications of cost uncertainty. On the other hand, we use Dana (1994), or equivalently Janssen et al. (2011), as a benchmark for the outcome of RPE with cost uncertainty. As shown in Janssen et al. (2011), the expected price under RPE is larger than the weighted average of the expected price of the high and low cost equilibria as developed by Stahl (1989) and in that sense, consumers are worse off under cost uncertainty. In this Section we show that this result may well be reversed for non-RPE. The first two panels (8(a) and 8(b)) show the dependence on search cost. For small search cost, a large fraction of non-shoppers performs two searches and the expected price is close to the average marginal cost of 25. When the search cost increases from initially low levels, the expected price increases and the fraction of non-shoppers performing two searches decreases (giving firms more market power). At search cost levels close to 2, there are multiple gap equilibria, and it may be that the expected price is decreasing in search cost. When the search cost further increases, a no-gap equilibrium emerges and the probability of non-shoppers searching twice becomes very close to 0. Panel (8(b)) also shows that starting from an initially small search cost, non-shoppers will search less when the search cost increases. In this way, non-shoppers partially mitigate the increase in market power typically associated with higher search cost. The last two panels (8(e) and 8(f)) show the dependence on the probability that the cost is high. When this probability is high, there is a no-gap equilibrium and consumers search very little, since there is a low probability of obtaining a substantially lower price. In this region the higher the α, the higher the expected price. For lower values of α, there is a monopolistic gap equilibrium with qualitatively similar properties. When α is sufficiently low, there are multiple gap equilibria and the incentives to search can be high, pushing the prices down. The expected price can be both increasing and decreasing in α depending on which of the regular gap equilibria is chosen.","In this Section, we deal with two important extensions of our general analysis. The first relates to introducing more general forms of correlation between firms' costs, the second relates to a first analysis of oligopoly markets with sequential search. Introducing an idiosyncratic cost component ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In Appendix C, we describe the analysis for the case where there is no uncertainty concerning the common cost component (the “pure idiosyncratic cost shock” case). The main take away from that analysis is that (for the same common cost component) a firm with a low idiosyncratic cost state randomizes prices over a support that is below and does not overlap with the support of the price distribution of a firm with a high idiosyncratic cost state. We conclude that it is entirely possible to extend the analysis in the main body of the paper and deal with situations where firms' cost has a common and an idiosyncratic component. Fig. 9 shows that if we add idiosyncratic cost uncertainty, the expected market price can be both lower and higher than without this uncertainty. The impact of idiosyncratic cost uncertainty on expected prices is very different than the impact of common cost uncertainty, as there is no consumer learning. The main effect is through the fact that the low and high cost distributions are not overlapping and that the reservation price is based on a weighted average of the expected prices of these two distributions. Oligopoly markets ~~~~~~~~~~~~~~~~~ It is not too difficult to reformulate our analysis to an oligopoly model by replacing sequential search with “newspaper search” a la Salop and Stiglitz (1977) and Dana (1994). Under newspaper search, a consumer pays a search cost only once to see all remaining prices. Thus, in both the duopoly model with sequential search and the oligopoly model with newspaper search, a consumer effectively has to take the decision whether or not to continue to search only once.","In this paper we have considered consumer search markets where firms' underlying common cost is unknown to consumers. If consumers do not know the prices different firms charge, it is natural that they also do not know the underlying cost. We have argued that in this environment of cost uncertainty, the standard RPE considered in the consumer search literature suffer from severe limitations. It was already known that RPE do not always exist, but we add that RPE implicitly assume specific out-of-equilibrium beliefs that do not satisfy standard game theoretic refinements. We characterize non-RPE that do not depend on specific assumptions regarding out-of-equilibrium beliefs and show that these equilibria always exist. Non-RPE may provide a significantly different assessment of the market power firms derive from search frictions. In non-RPE, non-shoppers are indifferent between buying and continuing to search over a range of prices. As prices in this range are set with positive probability, these non-RPE have active search with positive probability in equilibrium. Thus, we extend the Rothschild (1974) finding by showing in a model with endogenous price setting that in equilibrium firms price in such a way that consumers do not choose reservation price strategies. The fact that consumers rationally search more with cost uncertainty in non-RPE explains why market power may be overestimated in RPE. The additional search has a quantitatively important pro-competitive effect on prices. Our results on non-RPE also have important consequences for the empirical literature on consumer search models that has recently taken off. Non-RPE may explain the observations of De los Santos et al. (2012) and Honka and Chintagunta (2017), as in these equilibria (i) consumers may rationally continue to search at lower prices, while they buy at higher prices and (ii) consumers may stop searching and buy at a previously visited store, before they have observed all prices in the market (see our oligopoly extension in Section 6). Moreover, the price distributions of non-RPE are quite different from the regular price distributions found in RPE. It would be interesting to see whether these price distributions provide a good fit with empirical data. As a first inquiry into non-RPE, we have analyzed a stylized model limiting the immediate applicability of this paper to real world markets.23 In extensions, we have shown that some of the equilibria extend to oligopoly markets, and, importantly, we have dealt with markets where firms' cost consists of an idiosyncratic and a common cost component. Obviously, in an oligopoly framework one may want to consider a continuum of possible cost states. Such an extension of the present paper would be important in environments where the firms' cost is determined by an upstream firm (who can choose a continuum of different price levels). In such an environment, Janssen and Shelegia (2015a) have characterized interesting properties of RPE, but they also show such equilibria do not always exist. Non-RPE would solve this non-existence issue and it is natural to inquire into the qualitative properties of such equilibria. Bagwell and Lee (2014) provide such an analysis for the case where cost has an idiosyncratic component only. An obvious next step is to see whether our analysis on learning about a common cost component can be combined with their analysis. One important issue that needs to be addressed in the generalizations to oligopoly markets is how consumer inferences after observing two (or more) prices interact with the consumer search decisions. In the oligopoly extension analyzed in this paper, we dealt with the easiest of different possible cases that can arise. In general, however, different possible search behaviors interact in a complicated way with the incentive of firms to choose different prices. This paper made a first step analyzing non-RPE. There are many theoretical and empirical challenges that lie ahead."],["In a choice model, we characterize the loss induced by misperceptions of payoff-relevant parameters across a distribution of decision problems. When the agent cannot avoid misperceptions but has some control over the distribution of errors, we show that strategies that minimize loss from misperception exhibit systematic biases, akin to some documented in the behavioral and psychological literatures. We include illusion of control, order effect, overprecision, and overweighting of small probabilities as illustrative examples. --------------------------------------------------------------------------------","Within economic discourse, the idea that the human mind is imperfect and that some information is necessarily lost during any decision process dates back at least to Simon (1955). Under this premise, an agent faces, apart from the standard action choice, a problem of error management. In economic decision making, some mistakes in perception of payoff-relevant parameters are costlier than others, and this asymmetry has an impact on the frequency of different types of perception errors. In this paper, we relate error management to known observed behavioral biases. In order to understand the direction of biases arising under second-best perception strategies, we need first to understand the associated costs of under- or over-estimating payoff-relevant parameters. Our agent first receives some statistical information on a state of nature which can be either high or low, and this information translates into an objective belief p that the state is high. In a choice stage, she recalls this information imperfectly and ends up with a subjective belief q that the state is high. She then chooses an action that is optimal under her subjective belief q within a fixed choice set. Due to belief distortion, her choice may differ from the optimal one. We ask how large is the payoff loss incurred from the misperception of p for q. Equipped with this loss characterization, we study a class of error-management problems in which the agent chooses a distribution of perception errors that performs well across all decision problems that she encounters in her environment. We follow principles from the ecological rationality literature, in that our agent is unable to reoptimize the perception strategy in each encountered decision problem separately but, instead, must choose a perception heuristic that fits her environment.1 One feasible perception strategy is to memorize the true probability value. Under such a strategy, the recalled probability is in expectation equal to the true one. It turns out, however, that this unbiased perception strategy is generically suboptimal, and the agent benefits from memorizing m distinct from the true probability p. A misperception of q instead of p causes a loss only when the optimal choices at p and q differ. This happens precisely when there exists a probability s between p and q such that the agent is indifferent at s between the two lotteries. The likelihood that such a tie arises at s depends on s in an intuitive way. Since the expected value of each lottery is a convex combination of the two standard normal draws, its variance is lower the closer s is to 1/2. Thus, the likelihood of a tie at probability s is a single-peaked function of s attaining its maximum at 1/2. This implies that misperceptions of p for a value q towards the direction of 1/2 are more likely to distort choices than misperceptions in the opposite direction: undervaluation of one's own ability to predict the binary state leads to suboptimal choice more often than the symmetric opposite error. Since the agent can shift the distribution of her perception error by controlling the memorized probability value, it is optimal for her to memorize a value further away from 1/2 than the truth, hence to exhibit an overprecision bias. Apart from the overprecision application, we illustrate our methodology in three additional examples: illusion of control, order effect, and probability weighting. In each application, we make natural assumptions about the distribution of utility functions encountered by the agent in her environment and derive the optimal bias that arises. We relate these biases to stylized facts from psychology and behavioral economics. In the illusion-of-control example, the agent overvalues her impact on her own well-being. In the order-effect application, she exaggerates the quality difference among two available objects observed in the first of the two periods. In the probability-weighting example we revisit the model of Steiner and Stewart (2016), who derive the overweighting of small probabilities as a second-best perception heuristic. Related literature ~~~~~~~~~~~~~~~~~~ Rate-distortion theory studies optimal communication via a noisy channel (Shannon, 1948, 1959). One of the primitives of this theory is an exogenous map that specifies a loss to each input and output of the communication. A popular loss function is the square error. In our paper, the loss function is derived from the agent's environment and is equal to the average welfare loss caused by the distortion.2 Once we establish the relevant loss function, we let the agent engage in optimal error management: she avoids the costlier types of the errors. Beyond information theory, error management has been studied in several scientific disciplines such as biology and psychology. See Johnson et al. (2013) for an interdisciplinary literature review and Alaoui and Penta (2016) for a recent axiomatization of the cost-benefit approach to error management within economics. We contribute to this literature a formal characterization of loss from misperception based on a statistical description of the agent's environment. Sims (1998, 2003) has introduced an exogenous information-theoretic capacity constraint to economics. The subsequent literature on rational inattention studies the information acquisition and processing of an agent who, as in our model, is unsure about a payoff-relevant parameter but who, unlike our agent, knows her payoff function. In that setting (assuming signal cost is nondecreasing in Blackwell informativeness), the agent acquires only an action recommendation and no additional information beyond that needed for choice. In our model, the agent does not know her payoff function when she processes information and thus forms beliefs beyond those needed for the mere action choice in any given problem. The resulting optimal information structure can then be naturally interpreted as a perception of the payoff parameter.3 Our assumption that the perception strategy is optimized across many decision problems has appeared in the literature on the evolution of the utility functions. This literature studies the performance of decision processes across a distribution of the fitness rewards to the available actions. Depending on the constraints assumed, the optimal decision criterion is either the expected utility maximization, as in Robson (2001), or its behavioral variants, as in Rayo and Becker (2007) and Netzer (2009). (See Robson and Samuelson, 2010 for a literature review.) Most of this literature focuses on choice under certainty. Two exceptions who, like us, focus on the probability perception are Herold and Netzer (2010), who propose that biases in probability perception serve as a correction of another behavioral distortion; and Compte and Postlewaite (2012), who conclude that an agent with limited memory benefits from ignoring weakly informative signals. We contribute to the branch of the bounded rationality literature that emphasizes limited memory. Mullainathan (2002) studies the behavioral implications of exogenous imperfect memory usage. Like us, Dow (1991), Hirshleifer and Welch (2002) and Wilson (2014) study the behavioral implications of an agent who optimizes her usage of limited memory, but they examine effects that are different from ours. In his survey, Lipman (1995) focuses on memory-related frictions. This paper generalizes the model of Steiner and Stewart (2016), who study optimal probability perception in a choice between a binary lottery with random rewards and a fixed outside option. The contribution of this paper beyond Steiner and Stewart is threefold: (1) We significantly enlarge the class of settings by allowing arbitrary payoff distributions and arbitrary action sets; (2) we characterize the perception loss function from the more fundamental value of information function, and (3) we deliver new behavioral insights. Two of our derived biases, the illusion of control and overprecision, are classified among so-called positive illusions (Taylor and Brown, 1988). Such positive belief biases are often rationalized by assuming that the agent derives felicity from holding favorable beliefs about herself; see for example Brunnermeier and Parker (2005), Caplin and Leahy (2001), and Köszegi (2006). In contrast, belief distortions in our model are purely instrumental—they guide choice—and arise in the absence of any felicity benefits. The paper is organized as follows. Section 2 introduces the model, and our loss characterization is presented in section 3. We develop general results on the direction of biases in section 4, and applications in section 5. We show an example of a sophisticated Bayesian agent in section 6. Section 7 concludes.","Our aim is to characterize how the distribution of payoff functions u translates into the loss function L. The role of v″ ~~~~~~~~~~~~~~ The second derivative of the value function plays a central role in our analysis. Here, we offer intuitive explanations of its role in our results. Illusion of control ~~~~~~~~~~~~~~~~~~~ The term Illusion of control, introduced by the psychologist Langer (1975), refers to the overestimation of one's ability to impact one's own well-being. As an extreme example, casino visitors may overestimate the relevance of their choices over payoff-equivalent lotteries. Order effects ~~~~~~~~~~~~~ Psychologist Baron (2000) defines an order effect as order-dependent weighting of the observed pieces of evidence when the order of the evidence presentation is normatively uninformative. Primacy effect arises when the agent overweights her first impression and the recency effect refers to overweighting the last impression. Page and Page (2010) document the order effect in a field study of talent judgment and they argue for the relevance of the effect in hiring practices. Here we offer a stylized model in which the order effect arises as a second-best perception strategy for an agent with imperfect memory. Our next result shows that the second-best perception avoids the relatively costly types of error by a systematic exaggeration of the first impression. Overprecision ~~~~~~~~~~~~~ Proposition 3 together with Lemma 4 imply the direction of the bias. Overweighting of small probabilities ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We revisit here Steiner and Stewart (2016), who provide a microfoundation for overweighting of small probabilities akin to the one in prospect theory (Kahneman and Tversky, 1979). The purpose of the revisit is twofold. First, we show how the problem from Steiner and Stewart can be solved by the general method from this paper. Second, we contrast their model with the setting from the previous subsection in which underestimation of small probabilities arises, and clarify the forces driving the opposite biases in these two frameworks. Proposition 3 implies that overvaluation of small probabilities arises in this setting.","The agent from section 4 is naive in that she does not fully utilize the information available to her at the choice stage. This section studies the sophisticated perception strategies of an agent who loses some information during the decision process, but who is then fully capable of utilizing the information retained until the choice. Sophisticated perception","We have studied the cost of misperception in a model of decision making under uncertainty. Our main technical result, Lemma 1, shows that the loss function can be entirely characterized from the value function of the underlying decision problem. In our applications, the decision problem is specified by a random payoff function, which makes our environment rich enough. Note however that our loss formula applies to all settings, whether or not the payoff function is random, as long as the value function admits a second derivative. We presented a series of applications of this loss formula to settings in which the agent has imperfect memory, and hence faces an error-management problem; these examples allowed us to microfound several well-known behavioral biases. Microfoundations for behavioral biases can contribute to normative discussions of debiasing. Our model suggests that a bias relative to the precision-maximizing perception may be second-best optimal and thus presence of a bias is not enough to justify an intervention without further arguments. As the standard maladaptation argument goes, an intervention into the decision process may be beneficial if the agent's environment has changed since the biases have evolved. Interestingly, the model suggests that an intervention may be justified even if the agent's environment has not changed since adaptation took place. This is the case if an intervening outsider knows more about the agent's decision problem than the perception designer—evolution—has known."],["The paper investigates the incentives of Salop-type oligopolistic firms to cooperate and the architecture of the resulting collaboration networks. We find that when spillovers are exogenous, firm profits are not affected by the network structure. On the contrary, with endogenous spillovers (absorptive capacity) firms tend to form less dense networks. We also seek out the architecture of socially efficient networks, showing that social welfare is maximised in the complete network. Also, given the network structure we conclude that a Salop industry could be characterized by a general tendency to under-connection. --------------------------------------------------------------------------------","Our key findings are as follows. First, in presence of exogenous spillovers, firm profits are not affected by the network structure. Thus, empty or partially connected networks are desirable as much as a complete network. Second, with endogenous spillovers (absorptive capacity), firm profits decrease with the number of cooperation links. Therefore, there is no private incentive to form dense networks (empty network). Third, independently of whether spillovers are exogenous or endogenous, social welfare is maximised in the complete network. The structure of the paper is as follows. Section 2 describes the model and its properties. Section 3 presents the main results. Section 4 concludes.","In this section we present the main blocks of the model. We first introduce the preferences (demand and supply side) then we account for the timing of the moves. Endogenous spillovers: absorptive capacity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In a Salop oligopoly with imperfect information and endogenous spillovers (absorptive capacity), firms’ profits decrease with the number of cooperation links.","Proposition 3 states that link formation is always welfare improving. Specifically, in the case of endogenous spillovers, there is a general tendency to under-connection with respect to social welfare. This result may be explained as follows. In the case of exogenous spillovers, forming pair-wise collaboration links keeps unchanged firm profits (Proposition 1) but both reduces market price and increases consumer surplus. As a result, social welfare increases. On the contrary, with endogenous spillovers firm profits decrease as m increases (see Proposition 2). Nonetheless, this change is more than compensated by consumer surplus increase, thus entailing a greater level of social welfare."],["A government designs anonymous income transfers between a continuum of citizens whose income valuation is privately known. When transfers are deterministic, the incentive constraints imply equal treatment independently of the government's taste for redistribution. We study whether random transfers may locally improve upon the egalitarian outcome. A suitable Taylor expansion offers an approximation of the utility function by a quasilinear function. The methodology developed by Myerson to deal with incentive constraints then yields a necessary and sufficient condition for the existence of a socially useful randomization. When this condition is met a large set of lotteries are locally improving. A special menu made of two lotteries only is of interest: all the agents with low risk aversion receive the same random transfer, financed by a deterministic tax paid by the high risk aversion agents. --------------------------------------------------------------------------------","It is known that in a second-best world a principal may find it valuable to propose random contracts to the agents. For instance, in the presence of asymmetric information, risk can be used to relax the incentive constraints (Laffont and Martimort, 2002). In Gauthier and Laroque (2014) we give a necessary and sufficient condition for useful/useless randomization near a deterministic optimum, but our previous analysis only applies to well-behaved problems where the constraints are qualified, i.e., the gradients of the binding constraints at the optimum are linearly independent. In this note we deal with a case where the constraints are not qualified. We consider a government that allocates a given sum of money deterministically between potential recipients with different income valuations. As observed by Lerner (1944) only equal sharing can be implemented if valuations are not observed by the government and recipients always prefer more income to less. The constraint set reduces to a single point and qualification is not met. In this setup a random allocation may allow the government to screen individuals according to their attitudes toward risk. Pestieau et al. (2002) provide a necessary condition for local randomized redistribution to improve upon the equal sharing outcome. However their proof uses a Taylor expansion where the variance of the lotteries is negligible and so it does not give tools to design the optimal differential risk exposure. One contribution of our note is to provide a class of expansions where the noise component is first-order non- negligible. An appealing feature of this approach is that it enables us to apply the standard quasilinear toolkit of contract theory developed by Myerson (1982). This yields a necessary and sufficient condition for the existence of locally improving stochastic allocations. The social weights put on the agents with the smallest risk aversions must be large enough that a transfer in their favour more than compensates for the extra randomness required to meet the incentive constraints. There are many ways to design the locally improving randomizations. Still it turns out that there is no loss in generality in limiting the attention to simple schemes that work as follows. The agents have to choose between two (small) deviations from the status quo. One is a certain tax, the other is a random transfer with positive expectation and variance. The deviations are built so that all the agents with a risk aversion larger than a threshold choose to pay the certain tax, while the agents with a smaller risk aversion take the other (risky) option. The paper is organized as follows. Section 2 lays down the deterministic framework, which is extended to a random environment in Section 3. Section 4 presents an example. Finally Section 5 states and proves the main propositions.","We extend the power of the government and allow for randomized redistribution. The government can now design a menu of lotteries, such that every individual must choose some lottery in this menu. Given the random draw from the lottery, there is commitment from both the government and the players to conform to the outcome. Suppose also that a law of large number holds, so that with independent draws the cost of the lottery is equal to its mathematical expectation. With this class of small randomizations the variances of income transfers explicitly enter the incentive constraints (7). In the absence of noise these constraints reduce to (4) and prevent any income redistribution. Randomized income transfers expand the set of possible redistribution schemes consistent with incentive compatibility. Although the set of allocations consistent with incentive compatibility is larger than what is seen in the deterministic Lerner world, incentive constraints restrict possible redistribution. These restrictions, which follow from standard arguments, are given in Lemma 1."],["This paper explores how the interaction of different framework conditions affect the way in which top corporate R&D investors organise their cross-border operations worldwide. The analysis uses location-specific aspects, socio-economic factors and other controls common in the economic geography literature to investigate the distribution of a company's international subsidiaries. The location drivers are estimated using a multilevel mixed-effects logistic regression, controlling for both country characteristics and company specific random effects. Our results confirm that framework conditions, as product market regulation (PMR) and labour market legislation (EPL), affect the location strategies of top R&D investors. Adding to the literature we also found that: (i) PMR and EPL exert a mutually reinforcing negative effect on the location of subsidiaries, (ii) the effect of EPL is not significant for low levels of PMR, and (iii) barriers to trade and investment is the PMR component with the largest negative effect. Policy implications are drawn accordingly. --------------------------------------------------------------------------------","The policy debate on the European economic recovery is currently focusing on one key question: will the European economy be able to generate a self-sustained and balanced expansion once the three temporary tailwinds that have sustained its recent performance — the sharp fall in oil prices, a supportive macro-economic policy and the depreciation of the euro's exchange rate — have died down? To answer this question, both the shortfall of investment over the past few years, which has reduced the EU economic growth potential, and the declining trend in EU productivity growth, which has not yet been reversed, must be carefully considered. This calls for a reboot of growth drivers — such as productivity and innovation — to ensure a sustainable recovery and to prevent reverting to a situation of consistent weak growth. The Investment Plan for Europe1 launched at the end of 2014 by the European Commission represents a first step in this direction. In addition to providing a financial commitment, the plan has a long-term relevance for innovation and productivity growth because of its ‘third pillar’, which aims to create an investment- friendly environment. The underlying rationale of this pillar is that providing greater regulatory predictability, removing barriers to investment across Europe and further reinforcing the Single Market (i.e. creating optimal framework conditions) will unlock the full potential of investment in Europe, including investment in research and development (R&D). Accordingly, ‘more efficient labour and product markets’ are recognised as key policy priorities in the Five Presidents’ Report (European Commission, 2015a; p.7). Finally, the European Commission communication “A renewed European Agenda for Research and Innovation — Europe’s chance to shape its future” stress the importance of stimulating private investment in R&D and of regulation that provides the right environment to promote innovation (European Commission, 2018). Against this background, and being aware that fostering private R&D investment in Europe is essential to revitalise EU innovation performance, this paper analyses the impact that third-pillar-relevant variables have in shaping the organisation of top R&D investors cross-border activities worldwide. Namely, we test whether changes in the level of regulation in product and labour markets — and their possible interplay — have an impact in attracting foreign long-term investments. In doing so, we investigate the role that structural policies play in shaping firms’ location decision, as Aghion et al. (2005, 2009) did for their effects on innovation and productivity. To this end, we use the 2014 EU Industrial and R&D Investment Scoreboard (Hernández et al., 2014) dataset, collecting information on the top 2500 R&D investors worldwide and their ownership structure. In the empirical application, we model the probability of a company international subsidiary to be located in a particular country upon location-specific regulatory framework conditions, socio-economic factors, and other controls commonly used in the economic geography literature. The location decision drivers are estimated using a multilevel mixed-effects logistic regression, controlling for both country characteristics and company specific random effects. Following previous empirical literature in the field (e.g. Aghion et al., 2005), the synthetic indicators EPL (employment protection legislation) and PMR (product market regulation), provided by the Organisation for Economic Co-operation and Development (OECD), are used to characterise rigidities in the labour and product markets, respectively. From an economic point of view, well-functioning labour and product markets (able to absorb shocks and allocate resources in favour of the most productive and innovative firms) and lower administrative costs should help foster innovation and attract (and/or increase) investment in knowledge- based capital (OECD, 2013a) through the creation of a friendlier business environment. In fact, the effectiveness of the wider innovation system often hinges on the quality of framework conditions and the capacity to ensure an innovation-friendly climate in both more and less R&D-intensive parts of the economy (European Commission, 2015b). The remainder of this paper is organised as follows. Section 2 briefly reviews the literature on the role of well-designed framework policies in promoting innovation through the creation of a business-friendly and attractive environment and in influencing the location decisions of R&D-intensive multinationals; this is done with a specific focus on the empirical studies using PMR and EPL. Section 3 discusses the dataset and methodology employed for the empirical analysis. Section 4 presents a short descriptive analysis of the geographical location of the companies in the sample and comments on the results of the econometric analysis. Section 5 summarises the main results and presents the major policy implications. Section 6 briefly concludes.","The capacity to attract high-value FDI from knowledge-intensive and innovative multinationals is a key element for sustained growth as, especially in advanced knowledge- based economies; frontier innovation has become the main source of growth (Aghion and Griffith, 2005). In this context, multinational corporations (MNCs) have become central economic actors because of their large international investments and numerous foreign subsidiaries, which enable them to shift activities within their multinational network according to changing conditions. The literature on countries’ attractiveness in terms of international investment and location factors is broad and diverse and has not yet resulted in many clear-cut policy implications. Stated simply, a single theory explaining the location decisions of MNCs is still lacking, and a variety of models attempting to explain such decisions emphasise different drivers, getting also to different results.2 Among the drivers affecting the location decisions of MNCs, we pay special attention to the role of competitive product and labour markets. These have emerged as particularly relevant factors taken into account by technological leaders (Belderbos et al., 2008). Accordingly, in what follows we briefly summarise the most relevant findings of the empirical literature regarding the expected effects of product market and labour regulations on firm strategic decisions and performance. In doing so, our theoretical reference will be the Schumpeterian innovation paradigm. In a Schumpeterian framework, growth is driven by innovative firms introducing quality-improving innovations within the economy that tend to displace previous technologies. In this context, growth involves a conflict between old and new technologies and old and new skills. Structural policies — such as product market liberalisation and/or labour market flexibility — may have a role in facilitating this process of creative destruction (Aghion et al., 2005, 2009) by favouring a business-friendly environment. In other words, the Schumpeterian creative destruction process seems to be favoured by a dynamic business environment.3 This can be reached by setting framework conditions that, for instance, allow young innovative firms to easily enter a market and grow — without facing impeding regulation barriers — and allow inefficient firms to exit the market, freeing resources for those that stay in business. Moreover, the typically long-term character and high volume of investment involved in setting up a subsidiary abroad makes the stability of the regulatory framework a crucial factor. In the presence of high administrative entry costs, stringent PMR and EPL, this reallocation of resources may be less efficient, damaging the most innovative and productive firms and sectors (European Commission, 2015b). Aghion et al. (2005) empirically tested the Schumpeterian theory on the relationship between product market competition and innovation rate. They showed the existence of an inverted U-shaped relationship between product market competition and innovation. This implies the presence of an ‘escape competition effect’ at lower levels of product market competition, while the aforementioned ‘Schumpeterian effect’ — pointed out in earlier endogenous growth models and in the industrial organisation literature4 — dominates at high (but not too high) initial levels of product market competition.5 Therefore, although competition may increase the incremental profit from innovating, it could also reduce innovation incentives for laggards.6 The increasing body of literature and empirical research analysing the effect that changes in the level of PMR have on economic performance through the process of entry and exit also focuses on the final effect that this has on productivity (Conway et al., 2006). Scarpetta et al. (2002) found that the overall PMR level (and, in particular, administrative barriers to start-ups) has a significant negative impact on firm entry (see also Brandt, 2004). Blind (2012) studied the impacts of six different regulatory framework conditions on the innovation performance of a set of OECD countries, finding mixed results. Using OECD data, Bourlès et al. (2013) showed that stringent PMR has a negative indirect effect on total factor productivity, especially for countries close to the productivity frontier. This negative effect is mainly transmitted through investments in R&D and ICT (Cette et al., 2017), and operate along with a negative direct effect on production prices and wages Cette et al. (2014). Accordingly, Franco et al. (2016) found that PMR has a negative impact on R&D efficiency and the magnitude of this impact changes along the distribution of PMR. Amable et al. (2009) study the effect of market liberalisation (proxied with PMR) on innovation and productivity, finding that this is positively related to productivity the closer we get to the technological frontier. Barbosa and Faria (2011) study the impact of PMR, EPL, intellectual property rights and credit access on innovation intensity at industry level using CIS data. They find a negative effect of PMR and EPL on sector innovativeness. Andrews and Criscuolo (2013) relate allocative efficiency to framework policies, such as the level of administrative burdens on start-ups and the cost of closing a business, and to the level of EMP. Looking at OECD countries, Barone and Cingano (2011) study the effect of high levels of PMR the service sector on the performance of manufacturing industries heavily relying on them finding a negative relation. Along this line, Canton et al. (2014) found that a less strict PMR that applies to professional services, which are increasingly and intensively used by knowledge-intensive manufacturing firms (Ciriaci and Palma, 2016), improves business dynamics and, as a result, firms’ ability to allocate resources efficiently. For economies to thrive on innovation, labour also needs to be reallocated within and across firms and sectors. In a constantly changing economic context and/or with the introduction of labour-saving technologies, labour market frictions may generate inefficient labour allocation and unemployment for specific skills. In these cases, a stringent EPL may hinder the redirection of resources towards the most productive uses and delay the match of labour supply and demand. In line with this expectation, high EPL may reduce R&D investment, hampering the growth of innovative firms that are in need of skilled personnel and complementary resources to implement and commercialise their innovations (Andrews and Criscuolo, 2013; Amoroso et al., 2015). Similarly, Bassanini et al. (2009) find a negative impact of layoff restriction on aggregate labour productivity growth in OECD countries. Cingano et al. (2010) analyse the impact of EPL and financial market imperfections on investment, capital-labour substitution, labour productivity and job reallocation at firm level, finding a negative relationship with differences depending on the sector considered. Calcagnini et al. (2014) use EU firm level data for the period 1994–2000 (from Amadeus) to study the impact of labour regulation on firm investment decisions, finding a negative relationship. Similarly, Griffith and Macartney (2014) study the impact of EPL on innovation activity, finding that MNCs locate innovation activity mainly in countries with low EPL. Some studies go beyond the usual approach of considering the different types of regulations and barriers as independent factors by introducing an interaction term in their empirical setting. The interaction term allows considering possible mutually reinforcing effects among the different types of regulations. Égert (2016) uses OECD country level data to study the drivers of multifactor productivity, looking in particular to market and employment regulations. He finds an inverse relationship between high regulations (both market and employment) and multifactor productivity. By interacting the market and labour regulation proxies he uncovers a differential impact of market regulation depending on the level of labour regulation. Amable et al. (2011, 2016) perform a country level analysis on the relationship between product market and labour regulations and unemployment, also using an interaction term between the two regulation proxies. They found EPL and PMR are substitute rather than complementary policies.7 Our paper aims to contribute to this literature by verifying the extent to which differences in the strictness of regulation (in both product and labour markets) affect the location strategy of R&D-intensive multinationals. In particular, we test whether environments characterised by lower levels of product and labour market regulation are indeed more attractive to R&D-intensive multinationals looking for dynamic innovation ecosystems and market conditions that favour returns to their innovative investments. Moreover, we explore possible joint effects of employment and PMR on firms’ location decision. Despite the scant literature on this mechanism, we may presume that high PMR and EPL have a mutual reinforcing effect in reducing country attractiveness for R&D intensive MNCs. However, we do not have sufficient elements to determine ex-ante whether this effect would be relevant; we let the empirical analysis disentangling the significance and strength of this mechanism. Apart from PMR and labour market regulation, many other socio-economic factors may influence the strategic location decision of MNCs. Besides firm-specific characteristics, such as firm size and corporate performance, sector and technological intensity, other important drivers in the host country have been identified as particularly relevant. These include: (i) administrative burden (i.e. red tape), (ii) market features, such as size, growth potential and purchasing power; (iii) the presence of high-quality scientific infrastructure and human resources; (iv) the presence of tax breaks and government support, and the strength of intellectual property protection; and (v) geographical, technological and cultural distance between the source and the destination country.8 When looking at the relationship between PMR, EPL and firms’ location decisions, we will also considering these factors. The database ~~~~~~~~~~~~ The empirical application is based on the EU Industrial R&D Investment Scoreboard 2014 data (Hernández et al., 2014, http://iri.jrc.ec.europa.eu/scoreboard.html), which contains economic and financial information for the world's top 2500 corporate R&D investors. It is based on company data taken directly from companies’ annual reports.9 Each of these top 2500 R&D companies invested at least €15.5 million in R&D in 2013. Information on the subsidiaries of these top R&D investors was obtained by matching the Scoreboard dataset with the ORBIS databank of Bureau van Dijk, and reconstructing the corporate structure of the Scoreboard companies in place at the end of 2014. We were able to identify 218,343 industrial subsidiaries10 belonging to 2,338 of the top corporate R&D investors worldwide. Unfortunately, ORBIS does not report location or main sector of activity for all these subsidiaries. Only those for which both location and main sector of activity were known were included in the analysis. The final number of subsidiary companies included in the analysis is 140,382, belonging to 2,288 out of the top 2,500 corporate R&D investors worldwide.11 The Scoreboard dataset has been merged with data on PMR and EPL from the OECD. PMR is constructed with a bottom-up approach, through rounds of aggregation of different indicators, in which each component of the regulatory framework is equally weighted. The components of the PMR summary index are three composite indicators: (1) state control; (2) barriers to trade and investment; and (3) barriers to entrepreneurship. ‘State control’ measures the extent of public ownership in the economy and the involvement of the state in business operations (e.g. via controlling prices). ‘Barriers to trade and investment’ is the intensity of explicit (like barriers to FDI) as well as other barriers to trade and investment (e.g. the existence of a differential treatment for foreign suppliers). Finally, the indicator ‘barriers to entrepreneurship’ captures the complexity of the regulatory procedures, the administrative burdens on start-ups and the regulatory protection of incumbents. Data used to calculate PMR come from a questionnaire sent to governments in OECD and non-OECD countries (see Koske et al., 2015, for a full description of the methodology used to construct the PMR). It should be noted that, although being the only comprehensive, rich and comparable across countries available indicator on product market regulation the PMR does not capture all regulatory barriers, but only a selection of them. Therefore, it provides a partial approximation of the regulatory framework in each country. The EPL indicator measures the procedures and costs involved in dismissing individuals or groups of workers and the procedures needed to hire workers on fixed-term or temporary work agency contracts. The EPL is constructed as a weighted average of 21 items covering different aspects of employment protection regulations (as they were in force on 1st January of each year), which can be classified in three main areas: (i) protection of regular workers against individual dismissal; (ii) regulation of temporary forms of employment; and (iii) additional, specific requirements for collective dismissals (see OECD, 2013c for details and the specific weights applied to each item). PMR and EPL indicators measure product and market regulation on a 0–6 scale respectively; higher values represent a stricter regulation. We retrieved data on the administrative cost of starting a business and on the level of taxation on profits from the World Bank's Doing Business dataset, which allows cross-country comparison in a field where available (and comparable across time and countries) data are very limited. The cost of starting a business records the procedures officially required, or commonly carried out in practice, for an entrepreneur to start up and formally operate an industrial or commercial business, as well as the time and cost to complete them and the paid-in minimum capital requirements.12 Differently from ‘barriers to entrepreneurship’, which measure the complexity of the regulatory framework, the cost of starting a business tries to quantify the associated costs to a new venture, providing a slightly different information. We also include the Park–Ginarte index to control for the strength of intellectual property rights protection, which in turn is related to the R&D and innovation activities of a country (Ginarte and Park, 1997) and to the possibility for innovative MNCs to protect from imitation. From the World Development Indicators we extract data on gross domestic product (GDP), trade and unemployment to control for the macro-economic condition of a country. Finally, as is common in the relevant empirical literature, we also used a series of geographical and cultural distance measures provided by the GeoDist database, which was developed and described by Mayer and Zignago (2005, 2011).13 The econometric strategy ~~~~~~~~~~~~~~~~~~~~~~~~ This setting allows a more appropriate modelling of a firm decision to locate in a given country by directly dealing with the clustered structure of the data, where each cluster has its own choice behaviour. Indeed, in the present setting, instead of considering all observations at once, these are organised as a series of N independent clusters (companies) nested into I different clusters (industries). A similar setting has been used by Basile et al. (2008) to analyse the subsidiary location of multinational firms in 50 European regions, and by Griffith et al. (2014) to ascertain the importance of corporate income taxes in determining where firms choose to legally own their intellectual property rights. Among the drivers of a company location decision, xci, we included (i) the third- pillar relevant variables discussed above, PMR and EPL; (ii) country controls: cost of starting a business, tax rate on profits GDP level, differences in per capita GDP between source (headquarter location) and destination country, trade as a share of GDP, unemployment rate; and (iii) geographical controls: distance in kilometres between the capital of the hosting country and original country location of the Scoreboard company, and sharing a common language (binary 1/0). It is worth noting that, as PMR emerged as the main determinant of such a probability, we decided to analyse also the impact of its three components, i.e. state control, barriers to trade and investment, and barriers to entrepreneurship. Among the country controls, we had initially also included the share of people with a tertiary education, but, as results were confirmed and its insertion caused a large drop in the number of observations (especially because of missing data for China), we excluded it from the final estimations (results are available upon request). PMR values are available every 5 years (the last available year is 2013). In our analysis we used the average value between 2008 and 2013. The same holds for the high-level indicators (state control, barriers to entrepreneurship and barriers to trade and investment), whereby the simple average constitutes the value of PMR. However, EPL values are available yearly, and we used the average for the period 2008–2012. For the cost of starting a business (expressed as share of income per capita) we used the average from 2008 to 2012.14 Apart from EPL, PMR (and its three components) and common language, all the variables are expressed in logarithms. Table A1 (in the Appendix) shows the descriptive statistics for the explanatory variables. The identification strategy ~~~~~~~~~~~~~~~~~~~~~~~~~~~ Before discussing the estimation results, we deem important discussing the limitation of our approach and the mitigation strategy we take to try identifying the correct parameters of the model. Indeed, the identification of the parameters associated to PMR and EPL requires that governments’ decisions with respect to employment and market regulation are independent from firms’ location choices. In part, we rely to the fact that the decision of locating in a new country may take years (from the conception to its operationalisation) and that therefore the linkage from environment to location is the one more likely to have the higher signal to noise ratio. However, we also try to control for the possible influence a multinational may have on governments’ decisions. Industry level studies using PMR or EPL (e.g. Bassanini et al., 2009; Barone and Cingano, 2011) address the possible reverse causality issue using a ‘double difference’ approach mutated from the financial literature (Rajan and Zingales, 1998). This approach relies on the assumption that industrial sectors in different countries have similar characteristics. This homogeneity allows taking the differences of some characteristic of interest (e.g. how much each sector relies on external financing) between sectors in a benchmark country (typically the US) and use them as reference for the other countries. The exogenous differences so observed in the benchmark country are used to characterised the industrial sectors in the other countries (hence the ‘double difference’ approach), and then the analysis on the variable of interest is run. For example, Bassanini et al. (2009) use the different propensity to layoff among the US industries as reference for the other economies when analysing sectoral productivity growth. Our focus on firm data (rather than industry data) makes this approach not particularly suitable for the analysis. Indeed, the assumption of an industrial homogeneity across countries is excessively restrictive in a framework dealing with large multinationals presenting a persistent heterogeneity in their investment strategies (Montresor and Vezzani, 2015b; Coad, 2019). We therefore try to replicate the country/sector arguments in line with our theoretical framework. In particular, we control for the number top R&D investors operating in a specific sector and headquartered in a given country. Local companies — and in particular large multinationals — might have stronger linkages (and influence) on government decisions than foreign companies operating in the same sector and that may consider locating there. However, this variable may also represent a measure of the innovative capabilities of the industry/country. For this reason we also control for the GDP of the potential hosting economy. Again, while this measure may capture market determinants — larger markets may be considered more valuable by firms —, it also controls for the possible influence a foreign company might have on government decisions: larger economies are less susceptible to foreign influences. Furthermore, we run a series of robustness checks using specific exclusion restrictions (results are reported in Table A3 of the appendix). In particular, we first include the number of scoreboard companies headquartered in a given country and operating in a different sector from the foreign company that might potentially locate there (column 1). We then exclude from the estimations the 5% most internationalised companies (in term of number of countries in which they operate): being particularly globalised, these companies are more likely to move subsidiaries across countries and therefore might have an actual leverage on countries’ regulation decisions (column 2). We finally exclude countries not hosting a Scoreboard company headquarter (column 3). A part from this last estimation, which is particularly restrictive because basically keeps only the most advanced economies, the results with respect to EPL and PMR are not statistical different across all the estimations. All-in-all, these results suggest that our parameters are correctly identified. Where do top R&D intensive multinationals locate? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Out of the 2,500 companies present in the R&D Scoreboard, we managed to find information on subsidiaries for 2,288 of them. For these companies we considered only those subsidiaries that are not branches (as Castellani et al., 2017) and for which we could retrieve information on the sector of activity. Therefore, the final number of subsidiary companies included in the econometric analysis is 140,382, of which 89,940 (64.1%) are ‘international’, that is subsidiaries located in a country different from the one where the Scoreboard parent company headquarter is located (see Table 1). International subsidiaries located in the EU represent 58.3% of all subsidiaries, 10.3% are in the US, 1.5% in Japan and 2.7% in China (the remaining 27.2% are located in the rest of the world). Looking at the top countries where the Scoreboard companies locate their international subsidiaries, 56.7% are found in only 10 countries (see Table 2), a level of concentration that is lower than that for Scoreboard headquarters. Some interesting facts emerge if we compare the number of Scoreboard headquarters in a country and the number of international subsidiaries of other Scoreboard companies located in the same country. If we take the first 10 countries by number of international subsidiaries, only seven of them are also among the first 10 countries by number of Scoreboard headquarters. Spain, Canada and Mexico account for very few headquarters but many subsidiaries. For Canada, and especially for Mexico, this can depend on how close ties are to the US, as these three countries are part of a free trade agreement. As regard Spain, this observation can be explained by the general high level of attractiveness of this country for FDI during the 1990s and the 2000s — at least until the financial crisis (see, e.g., Barrios et al., 2004; Rodriguez and Pallas, 2008; Guimón, 2009). In contrast to the above, Switzerland, Sweden and Japan are in the top 10 for number of Scoreboard headquarters but are in a lower position on the chart for Scoreboard subsidiaries. The position of Sweden does not come as a surprise, given the size of its internal market. In the case of Switzerland — host to many pharmaceutical and financial multinationals — the explanation may reside in the way we count subsidiaries: considering only industrial ones and not branches, which is the kind of subsidiary that is typical of financial MNCs. Finally, the reason why Japan is at the bottom of the top 20 destination countries chart may be the historical closed nature of the Japanese market and its onerous institutional rules (see Nakamura and Oyama, 1998). So it comes as no surprise that Japan is the country with by far the highest number of national subsidiaries (11,525), accounting for 22.8% of all national subsidiaries of the Scoreboard companies included in the analysis. This is in line with the so-called ‘home bias’ of major Japanese R&D intensive firms (Belderbos et al., 2013). A similar picture emerges when the number of Scoreboard companies headquartered in a particular country is compared with the number of Scoreboard companies present in the same country through one or more of their subsidiaries. Some countries, such as Spain and Canada, register a much higher presence of Scoreboard companies than might be expected if only the number of headquarters is considered. Before moving to the empirical application, we look for descriptive evidence supporting our hypothesis of the role played by regulatory framework in the location decision of companies. In Table 3, we report the pairwise correlation coefficients between the number of international subsidiaries located in country i, and our main variables of interest, PMR and EPL and the other framework conditions considered in this study. Table 3 shows a statistically significant negative correlation between the level PMR and the presence of international subsidiaries of R&D-intensive Scoreboard companies. Also, the level of labour market regulation is negatively correlated with the presence of international subsidiaries and positively with PMR, but in both cases the coefficients do not appear to be significant at the standard statistical level. The patent protection index is instead positively correlated with the presence of international subsidiaries and negatively with PMR. Econometric analysis ~~~~~~~~~~~~~~~~~~~~ The estimation results for the different specifications used are reported in Table 4 (columns 1–4). First we include PMR and EPL, the variables discussed in the identification strategy and geographical controls (column 1). Then we introduce the interaction between EPL and PMR (column 2) and later all the other control variables, including cost of starting a business, corporate income tax and patent protection index (column 3). Finally, we report the results on an estimation excluding unemployment and trade (column 4). This accounts for the possible impact of PMR and EPL on unemployment and trade; given their effect on the probability of location may be underestimated if part of it goes through these variables.15 The specification reported in column 4 is our reference model and we use it also to identify which component of the PMR indicator — which emerged as much more relevant than EPL — is the most influential (Table 6). From a policy point of view, this analysis would help not only to detect which of the two aforementioned third pillar relevant variables affects the most the kind of investment considered here (the presence of a subsidiary in the country), but also to rank the different possible reforms in terms of their potential impact. Before discussing the general estimation results, a caveat should be pointed out. As no information was available on the year when the subsidiary was opened, a lack of data common to many studies in the field, identifying causal relations is not straightforward. However, our econometric analysis aims not to explain the decision to open a subsidiary in a particular year, but to analyse the correlation between a multinational's decision on how to organise its worldwide activities and a series of country characteristics and geographical variables. These variables are those likely to be considered when such a decision is taken and/or when the mother company decides whether or not to stay in a particular host country. The majority of these characteristics are structural and display a low degree of variation over time. Nonetheless, we use average values over a 5- or 6-year period for our main variables of interest, and a 1-year lag for the controls. Finally, it should be recalled that, as we use a particular class of logistic estimator, the coefficients attached to our regressors could not be interpreted as elasticities. In order to quantify the effects of our main variable of interest (PMR and EPL) on the location conditional probabilities we will compute and comment their marginal effects. The effects on location probability of top R&D multinational for third-pillar relevant variables Our results confirm that product and labour market regulations significantly affect the location probability of top R&D multinationals (Table 4). However, the coefficient attached to EPL is not statistical significant when its interaction with PMR is not taken into account. This is an interesting result pointing to the importance of considering the joint effect of the two framework conditions addressed in this paper. In the baseline specification (column 1 of Table 4), the PMR indicator is negative and significant, while the EPL is not. In the following specifications (columns 2–4) we include the interaction term between PMR and EPL. This makes meaningless the simple reading of the individual coefficients of the two indexes. To ascertain their overall impact on location decision we will compute the marginal effects later on. What is important noticing here is the coefficient capturing the interaction effect between PMR and EPL, which is negative and significant also when additional controls are included. This means that PMR and EPL do not have independent effects on the location choices of our firms. This result is in line with Aghion et al. (2009) that, exploiting a macro- panel data for OECD countries, found that the interaction between PMR and labour market rigidity explained a major part of the total factor productivity growth of countries close to the technological frontier, the effect of EPL being significantly lower if considered on its own. Moreover, our results suggest that there can be a certain value of the PMR above which the EPL becomes a significant negative explanatory variable of the location probability, an aspect that will be dealt with in more detail later. All the country controls included are significant and enter the equation with the expected signs. In addition, in line with previous empirical evidence (e.g. Barrios et al., 2012), the cost of starting a business16 and the tax rate on profits show a significant negative correlation with the probability of locating in a particular host country.17 As expected (Hall, 2011), the patent protection index enters with a positive and significant coefficient. Also the two variables discussed in the identification strategy, GDP and # of SB headquartered (in the same sector) enters with a positive sign: the market potential of a country as well as the presence of national top R&D investors may attract further knowledge investments. Our companies also tend to favour countries with levels of GDP per capita lower than the source country. As far as the included geographical controls are concerned, we found that distance from headquarters negatively affects the probability of locating subsidiary activities in a given foreign country. In fact, the relative estimated coefficients are negative and strongly significant for all the specifications. On the contrary, sharing a common spoken language affects it positively. In fact sharing a language increases the efficiency in transferring and aggregating knowledge (Grant, 1996). The importance of this factor makes particular sense in the case of knowledge- intensive companies, such those in this analysis. As we said, because of the interaction term between the PMR and EPL indicators, the signs of the individual coefficients reported in columns 2–4 cannot be directly interpreted. What counts are their overall marginal effects, which considers also the interaction term. Therefore, we have computed the overall marginal effects for EPL, PMR and the other variables proxying the framework conditions of a country. Results are reported in Table 5. The marginal effect of PMR is much larger than that of EPL, indicating that PMR have a higher influence on the location decisions of R&D-intensive multinationals than employment protection. More specifically, a 1-point increase/decrease in the PMR indicator decreases/increases the probability of choosing a particular location by 16.4%. The same increase in the EPL indicator decreases said probability by 1.8%. This effect is similar in magnitude and sign to the one of corporate income tax, while the cost of starting a business plays a minor role. The marginal effect of the patent protection index is instead positive and particularly relevant (18.4%). The last two results are not surprising given that we are dealing with R&D intensive MNCs, which may be more concerned about the creation and exploitation of knowledge assets rather than the administrative cost related to starting a new business. Because of the non-linearity of the logit model, and to better exploit the interaction between EPL and PMR, we also computed the marginal effects of each of these two variables for different values of the other one (Fig. 1). As anticipated, the non-significance of EPL in the baseline specification (where EPL and of PMR are considered separately) allows the identification of a threshold value of PMR above which EPL becomes significant in determining the location probability. Fig. 1 reports the marginal effects of PMR and EPL (lines) alongside the distribution of EPL and PMR (bars), respectively. The negative impact of PMR on the location probability increases as the level of EPL rises (left panel). This means that the higher the level of EMP in a given country, the greater the negative marginal effect of PMR on location probability. This reinforcing mechanism is similar to the one found by Fiori et al. (2012) studying the effect of PMR and employment protection on the employment rate of OECD countries. More interestingly, for low levels of PMR, the effect of EPL on the same probability is null and eventually positive (right panel). In sum, the higher the level of PMR, the higher the negative effect of employment protection, and vice versa. Fig. 1 also shows that some countries — USA and New Zealand — have levels of EPL notably lower than the rest of the sample, while others — China and India — have levels of PMR notably higher. As a further robustness check, we have run our main specification excluding these countries. Our main results are not influenced by the inclusion or exclusion of these countries (see Table A4). The effects of the different components of PMR Given the prominent role of the level of PMR emerged in the analysis, being the PMR index used a composite indicator (see Section 3.1), we also explore which one of its three main components — State Control, Barriers to trade and investment, and Barriers to Entrepreneurship — has the greater influence on the location probability. Table 6 reports the estimations using the full specification used in Table 4 (column 4), including separately the three components of PMR as explanatory variables and their interactions with the EPL composite indicator. The three columns of Table 6 report the coefficients of the independent variables, each including a different component of PMR in the set of explanatory variables. Since signs and significance of all the control variables do no change compared to the main estimation (Table 4), we will comment only the results for the components of PMR. As before, the presence of the interaction term between EPL and the PMR components (a different one for every specification) makes the simple reading of the individual coefficients’ signs and magnitude misleading. Therefore, Fig. 2 reports the overall marginal effects at the mean values of the three PMR components. Barriers to trade and investment is the component with the greatest negative impact on location probability (−13.6%), followed by state control (−10.4%) and by barriers to entrepreneurship (−4.2%). Barriers to trade (such as differential treatment of foreign suppliers) and investments, and excessive state control (which may lead to unfair competition from state-owned enterprises) have a much greater effect on the probability of an MNC's subsidiary to be located in a particular country than barriers to entrepreneurship. This latter component may be more hampering for local entrepreneurship, eventually limiting competition and thus favouring local and foreign MNCs. However, the negative effect we found suggests that foreign MNCs still valuate negatively this type of obstacles.","Our results confirm previous empirical evidence showing that R&D-intensive MNCs tend to locate their international activities in countries with greater market potential, head- start opportunities and the right environment to create and exploit their knowledge assets. Companies prefer to move into countries close rather than far away from their headquarters. They also choose the location for their international activities on the basis of cultural factors, such as a common language. This could make the task of sharing and integrating technical knowledge (often tacit) easier, which would lead to lower costs. The empirical evidence confirms that product and labour market regulations significantly affect the location decision of top R&D investors’ subsidiaries. However, when taken separately, EPL does not appear to play a significant role in such choices. When considering instead the interaction between EPL and PMR, the former becomes significant, while the latter has the greatest negative effect on companies’ location decisions. These results show that these two regulations exert a mutually reinforcing negative effect on the decision of top R&D investors about where to locate their subsidiaries. In other words, the lower the level of PMR, the lower the negative effect of employment protection and vice versa. Translated into policy terms, these results suggest the importance of bearing in mind both markets when considering possible reforms of one of the two. The interplay between PMR and EPL calls for integrated/coordinated policies in the two realms. While the effect of PMR on location decisions is negative for each level of EPL, the latter does not exert a significant effect for low levels of PMR. Our setting allows us to identify the threshold level of the PMR index (1.5) below which EMP has no relevant effect on the location decision of top R&D investing MNCs. In Fig. 3 we rank the EU-28 countries according to their average PMR level in 2013, from the highest to the lowest, and report the aforementioned threshold value. The sample of countries is divided into two groups, above (left) and below (right) the threshold. A reduction of their PMR in countries above the threshold (Lithuania, Sweden, Malta, Bulgaria, Latvia, Poland, Cyprus, Romania, Slovenia, Greece, and Croatia) would mitigate and eventually neutralise the direct negative effect of EPL on location decision. Clearly, this does not imply that those countries below the threshold would not benefit from a reduction of PMR, as the coefficient attached to the PMR indicator is always negative and significant. As already said, our finding does not rule out the possible beneficial effect of reducing EPL. A reduction of EPL would decrease the negative conditional effect of PMR on location decisions (Fig. 1). Among the different components of the PMR index, barriers to trade and investment have the greatest impact on location decisions, followed by state control and barriers to entrepreneurship. By lowering barriers to trade and fostering investment, EU policy-makers may facilitate the market uptake of new products and have the greatest impact in attracting foreign investments. The cost of starting a business and the corporate income tax rate also play a negative role on companies’ location decisions. However, their effect is lower than that of the other two framework conditions discussed above. The decisions of the companies considered in our study seem to be driven more by a desire to improve efficiency than by cost reduction considerations. This seems supported also by the importance of patent protection for the location decision.","This paper aimed at investigating the extent to which product and labour market regulations affect, among other socio-economic factors, the probability that top corporate research and development (R&D) investors locate in a particular country. Our results show that PMR and EPL exert a mutually reinforcing negative effect on the location of top R&D investors’ subsidiaries. Moreover, of the three components of the PMR indicator — barriers to trade and investment, state control and barriers to entrepreneurship — the first is the one with the highest negative marginal effect. Finally, a caveat should be kept in mind when reading our results. In our analysis we consider a particular set of firms (top R&D investors worldwide) and their current subsidiaries, for which we don’t have the year of establishment. This focuses our analysis on the observation of conditional ex-post probabilities, and not on the genesis of the multinational foreign subsidiaries network (Egger et al., 2014). Second, the aforementioned PMR/EPL mutually reinforcing effect does not automatically imply that markets should be deregulated as such because, as shown in Aghion et al. (2005) for innovation, an excessive reduction of market regulation could even negatively influence country attractiveness. Therefore, further research on this area may provide additional insights on the necessity to carefully consider how PMR and employment protection interact in different socio-economic environments taking into account the multiple effects they have on the society."],["This paper assesses the performance of two recently developed tariff aggregators in reducing tariff aggregation bias by analysing Swiss beef market liberalisation scenarios. Specific relevant sources of bias are addressed: substitution effects on import demand, Tariff Rate Quotas and overprotection in tariffs. The aggregators are linked to a global large-scale partial equilibrium model and benchmarked against a standard aggregator. The choice of the aggregation method shows considerable effects on simulated economic impacts, specifically if the dispersion in tariffs or tariff cuts is large. A large bias is revealed in simulated gains from trade liberalisation using the standard aggregator. The impacts on traded quantities are found to be overestimated, while price and welfare effects can be higher or lower by switching to alternative aggregation methods. By reducing aggregation bias and depicting negotiated tariff schedules more directly, the proposed aggregators enhance the contribution of trade modelling to evidence-based policy making. --------------------------------------------------------------------------------","Market access policies, and import tariffs in particular, are typically defined at the detailed tariff line level.3 The “tariff schedule” of a country normally includes thousands of tariff lines, and trade statistics are also often recorded at this fine level of disaggregation.4 As trade negotiation offers and requests made in trade talks (tariff cuts, exceptions to tariff cuts, sensitive products etc.) usually refer to tariff lines, economists need to analyse and process tariff distributions and detailed trade data when they support negotiators by providing simulated potential economic impacts of policy options. However, most applied models of international trade operate by using more aggregated commodities, each of them often covering a wide array of tariff lines. As a result, detailed tariff line level data first needs to be aggregated, which in turn leads to aggregation bias in simulated model outcomes as well as to partial loss of information on detailed trade and trade policy patterns. In this paper we assess the performance in reducing tariff aggregation bias of two recently developed tariff aggregators. The focus is on three specific sources of bias which are relevant to a wide range of agri-food markets and empirical applications: substitution effects on import demand, Tariff Rate Quotas (TRQ) and existing overprotective buffers of tariff protection. The proposed approaches depart from state-of-the-art consistent tariff aggregation techniques. The first one, the trade expenditure aggregator (TE), is an equivalence measure of tariff protection, defined as the uniform tariff rate that is equivalent to a set of individual tariffs in terms of its impact on trade expenditures. The second one, the Tariff Reduction Impact Model for Agriculture (TRIMAG) aggregator, adjusts aggregation weights based on a CES function defined at the tariff line level, and provides an explicit representation of the overprotective part of the duties. The proposed aggregators are linked to a global large scale partial equilibrium model as standalone modules, thereby demonstrating an impact assessment approach that can be replicated to a wide array of trade modelling applications. This paper provides a concrete application of the proposed impact assessment approach by examining a set of Swiss beef market liberalisation scenarios. Simulation results are also benchmarked against a standard aggregator for systematic comparison. Our assessment shows that the choice of the aggregation method has considerable effects on simulated economic impacts, specifically in the case of heterogeneous tariff cuts and large tariff dispersion. Simulation results also suggest that the proposed aggregators reduce aggregation bias. A further advantage of the proposed tariff aggregators is that they allow for a more accurate implementation of negotiated tariff schedules in applied equilibrium models, by exploiting tariff line level details of trade- and trade policy data. Against this background, we argue that the proposed aggregators enhance the potential contributions of trade modelling to evidence based policy making. The paper is structured as follows. In Section, 2, we discuss possible sources of tariff aggregation bias and its measurement (Section 2.1), and then formally introduce our strategy for consistent tariff aggregation (Section 2.2), presenting in detail the TE and TRIMAG aggregators (Section 2.3 and 2.4, respectively). Then, in Section 3 we briefly describe the CAPRI model. After that, in Section 4 we define the application to the Swiss beef market as well as the test tariff dismantling scenarios. Data and simulation results are presented and discussed in Section 5. Finally, concluding remarks are reported in Section 6.","This section presents a comprehensive yet concise review of the most relevant sources and measurement issues related to tariff aggregation bias (Section 2.1), followed by the presentation of our approach (Section 2.2) and of the proposed tariff aggregators (Sections 2.3 and 2.4, respectively). Sources and measurement of aggregation bias ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The straightforward way to improve exploitation of existing datasets at the tariff line level is to extend trade models to the tariff line. Indeed, there have been several attempts in this direction in the literature. Grant et al. (2007) iteratively link a Computable General Equilibrium (CGE) model with a satellite partial equilibrium (PE) model for dairy products at the HS 6 level. Narayan et al. (2010) opt for a nested general equilibrium (GE)/PE modelling approach in order to use tariff line level data to assess the impact of tariff liberalisation on the Indian auto industry. Britz and van der Mensbrugghe (2016) advocate advanced database filtering methods and algorithmic improvements to increase regional and commodity-wise disaggregation in trade models and therefore limit the need for aggregation at the same time. However, these approaches still only offer partial solutions because they only cover selected sectors or they still do not disaggregate models to the tariff lines. Therefore, computational and data issues still seem to force applied economists to keep to more aggregated commodity and regional groups in applied trade modelling. Global statistics on both demand and supply are still lacking at the tariff line level, and when extending import demand systems to the tariff line, the number of possible bilateral trade flows quickly becomes computationally unmanageable. Practitioners are consequently confronted with the choice of a tariff aggregation method that fits both the technical implementation and structure of their modelling tools as well as the focus of the specific policy question to be addressed. Aggregation over tariff lines inevitably leads to biases in simulated impacts of trade liberalisation, at least for some model outcomes. Measuring this bias involves comparing the simulated results derived using both a disaggregated (tariff line level) and an aggregated version of the same model. However, impacts simulated using the disaggregated model cannot be considered to be ‘true’ impacts because they are determined by the model parameterization (Pelikan and Brockmeier, 2008b). Therefore, the lack of a clear reference point does not allow for an absolute measurement of the bias, but instead all measures of the aggregation bias are necessarily relative to a selected disaggregated model structure and its parameterization. Empirical evidence still suggests that the aggregation bias is substantial. Laborde et al. (2017) compare simulated impacts of global trade liberalisation using different tariff aggregation methods in the LINKAGE model (Van der Mensbrugghe, 2005) and find that the optimal tariff combination of Anderson (2009), which we also discuss below, almost double welfare impacts compared to traditional trade weighted aggregators. The simulated terms- of-trade and real income effects largely vary between regions, and is generally largest for developing countries. They also find a crucial impact of model parameterization on the aggregation bias: increasing substitution elasticities in their modelling framework leads to larger differences in simulated welfare gains. In a modelling exercise with using trade restrictiveness indices (Anderson and Neary, 2005) in the GTAP model (Hertel, 2007), Pelikan and Brockmeier (2008a) also find that standard trade weighted tariff aggregation leads to largely underestimated welfare gains from global trade liberalisation. However, the trade modelling literature in general (including gravity models) is inconclusive about the relative empirical size and even about the direction of the aggregation bias. While some such as French (2016), who find a downward bias in a flexible model allowing for comparative advantage across products for every country, Bektasoglu et al., 2017 calculate an upward bias in estimating non-tariff measures using traditional gravity modelling techniques. Regarding the main sources of the aggregation bias, the trade protection measurement literature attributes them either to (i) the conversion of different types of tariffs (ad valorem, specific, and compound) to a common metric; or (ii) to the aggregation process from the tariff line to a more aggregated commodity level (Pelikan and Brockmeier, 2008a).5 With regard to the aggregation process, standard methods only use trade data for defining aggregation weights. In principle, other supply and demand related data (e.g. production and consumption) are also relevant for tariff aggregation, but are often neglected due to data availability. For example, the estimation of product- substitution in the consumption bundle requires observed consumption patterns at the tariff line level. Many ways of averaging tariffs for trade modelling can be observed in the literature. Tariffs might receive the same weight when calculating a simple average; medians can be used to take skewed tariff distributions into account; or weighted averages can be formed by applying, for example, import shares valued at border prices as aggregation weights. The latter approach is specifically prone to endogeneity bias since highly taxed goods tend not to be imported and so might even receive zero weights. Therefore, trade weighted average tariffs typically underestimate trade restrictiveness. They not only tend to smooth out tariff peaks but they also tend to underestimate simulated welfare gains from trade liberalisation (see Martin et al., 2003). Specific weighting schemes are readily available to reduce the endogeneity bias, e.g., the “reference group” method (Bouet et al., 2008; Guimbard et al., 2012) which clusters countries into groups of similar GDP per capita and trade openness indices. Bach and Martin (2001) proposed and Anderson (2009) fine-tuned an approach for eliminating all aggregation bias in the GE framework for a selected model outcome. Their welfare consistent aggregation corrects the entire aggregation bias for simulated welfare impacts by introducing an optimal combination of tariff aggregators in a modified balance of trade condition into GE models. Restrictions for the consistency include imposing separability conditions for consumer utility maximisation and assuming zero price impact on the domestically produced good. While the first assumption is commonly present in most CGE models through the Armington assumption, the latter is quite restrictive, and it has not been possible to significantly relax it even in recent extensions of the approach for large-scale CGEs (Laborde et al., 2017). The approach also requires the parameterization of complete supply and demand systems against a large number of behavioral parameters, such as substitution elasticities. This paper does not aim at a comprehensive classification of the possible sources of tariff aggregation bias. Instead, the focus is on three specific sources: (i) demand-side substitution effect observed at the tariff line level; (ii) the endogenous determination of tariffs under Tariff Rate Quotas (TRQ); and (iii) the difference between initial unit quota rents and applied out-of-quota tariffs under binding TRQs, potentially increased with “water” in tariffs. The latter source of bias is treated in the following discussion as an over-protective part of the duty. The above sources of aggregation bias are not only relevant to our empirical investigation of the Swiss beef market liberalisation, but they are also relevant to a wider range of policy applications. Firstly, the substitution effect-related aggregation bias is closely linked to the heterogeneity of commodity groups. Trade liberalisation can change the composition of the consumed commodity group significantly. Assuming separable homothetic preferences, as is often the case in trade models, this demand side adjustment is attributed to relative price changes within a larger product group. When facing large tariff dispersion (either due to differences in initial tariff rates or due to different tariff cuts), significant changes in relative import shares can be expected and so the fix trade shares assumption underlying standard tariff aggregation becomes too restrictive. Consistent aggregation (Anderson, 2009; Himics and Britz, 2016) offers an approach for correcting the demand side adjustments by setting up an explicit import demand system at the tariff line. This approach is also followed in this paper and Constant Elasticity of Substitution (CES) type utility aggregators at the tariff line are set up to either (i) adjust aggregation weights according to shifts in import shares; (ii) or to calculate an equivalent average tariff in terms of shifts in trade expenditure shares. The point of departure from the literature is that the authors of this paper do not aim at full consistency in terms of a single selected model outcome (e.g. welfare), but rather the applicability of the proposed aggregators in reducing the substitution effect-related bias in a range of simulated outcomes is tested. Secondly, covering TRQ related aggregation bias is crucial for the liberalisation of global agri-food markets, which are often regulated by complex TRQ systems (de Gorter and Kliauga, 2006; WTO, 2012). Opening TRQ- regulated market access (e.g. in the form of expanding quota limits and lowering in- or out-of-quota tariff rates) has been included in many recent trade negotiations, including the EU-Mercosur (Ramos et al., 2010) talks and also the current debate on post-Brexit trade relations of the EU, and the UK with third countries (Revell, 2017). TRQs are two- level tariffs with a limited volume of imports permitted at the lower “in-quota” tariff and all subsequent imports charged the higher “out-of-quota” tariff. In other words, the tariff rate applied on imports depends on the quota fill rate. Standard aggregation techniques do not always take the possibility of a TRQ regime shift into account (e.g., from an in-quota to an out-of-quota situation), and neither the induced adjustment in the applied tariff rates, or at least not at the tariff line level where TRQ thresholds are usually defined.6 Neglecting TRQ regime shifts in standard tariff aggregation (e.g., by converting TRQs into an equivalent tariff rate, additional trade cost or price gap), can lead to both over and underestimation of the applied tariff rate. The literature is inconclusive regarding the size of the potential aggregation bias from TRQ regime shifts. Using aggregated TRQ data in the MIRAGE CGE model, Decreaux and Ramos (2007) find that the aggregation bias itself is relatively small compared to the bias induced by not modelling the TRQ regime shifts explicitly in the aggregated (CGE) model. On the other hand, Himics and Britz (2016) found that the aggregation bias can be significantly large when using fix aggregation weights, even for scenarios involving moderate TRQ expansion. In order to capture the interplay of tariff rates and imported volumes under a TRQ regime, we follow different approaches in the two aggregators we propose. One of the aggregators replicates the Mixed Complementarity Problem (MCP) formulation of Himics and Britz (2016) for modelling TRQ regime shifts at the tariff line, while the other one aggregates in-quota and out-of-quota rates separately, and leaves the explicit TRQ modelling for the aggregated equilibrium model at hand, similar to Decreaux and Ramos (2007). A significant impact of tariff “water” or “binding overhang”, which is defined as the gap between bound and applied rates (Bchir et al., 2006), on simulated gains from trade liberalisation has long been identified in the literature (e.g. Beshkar et al., 2015). The gap between applied and bound rates also adds another source of complication to TRQ modelling, further motivating the research presented in this paper in this direction. The “water” in tariffs provides trade negotiators with a leverage to offer tariff cuts that have no significant impact on domestic prices (and therefore domestic producers) in the short term (Bchir et al., 2006). Not only the “water” in tariffs, but also the difference between unit quota rent and the out-of-quota rate provides negotiators with an additional buffer to accept a decrease in applied rates as the decrease will not necessarily affect import prices. To illustrate this point, consider a binding TRQ (Fig. 1) in the small country case. The equilibrium price is determined as the intersection (A) of the net import demand and excess supply curves, the latter being fully determined by the quota threshold and the TRQ tariff rates (see Skully, 2001). In case the TRQ is binding, net import demand is strong enough to induce imports up to the TRQ limit but fails to trigger imports beyond the threshold. Consequently, an economic (quota) rent is generated. Tariff cut scenarios would only have an impact on imports, ceteris paribus, if the out of quota rate is reduced beyond the per unit quota rent. In other words, as long as the applied rate after tariff cuts stays above the unit quota rent, import demand will not exceed the quota level and thus no price impact will be triggered. In order to highlight this additional buffer for trade negotiators to decrease tariffs without an impact on imports, the sum of the “water” or “binding overhang” in tariffs and the gap between unit quota rent and applied rates is denoted as the overprotective part of the tariff. Finally, the overprotective part of the tariff (partly defined by the size of the quota rent) is the third source of aggregation bias specifically examined in this paper. Lips and Rieder (2005) identified that standard tariff aggregation yields biased aggregate quota rents, and proposed using aggregation weights based on quota rents and tariff revenues. They find significantly smaller price impacts in a TRQ expansion scenario using the Global Trade Analysis Project (GTAP) model when quota rents are taken into account although the initial unit quota rents in their study are based on simplified assumptions rather than on real data. Himics and Britz (2016) also highlight the impact of quota rent assumptions on simulated results when showing that setting quota rents uniformly at half of the difference between in- and out- of-quota rates largely overestimates effective tariff reduction. A data driven approach is also implemented in this paper to specifically define the unit quota rent and therefore decrease the potential aggregation bias from arbitrarily setting initial quota rents. Quota rents for estimating trade liberalisation impacts are not only relevant for negotiators, but farmers might also perceive potentially decreasing economic rents as a policy risk (Nogueira et al., 2012). The proposed tariff aggregators ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Two strategies to go beyond standard tariff aggregation and address the above sources of aggregation bias are proposed. Firstly, the method presented in this paper departs from state-of-the-art consistent tariff aggregation techniques and sacrifices full consistency for the sake of improving a multitude of simulated model outcomes. The trade expenditure aggregator (TE), evaluated here is an equivalence measure of tariff protection, defined as the uniform tariff rate that is equivalent to a set of individual tariffs in terms of its impact on trade expenditures. Formally, the TE aggregator is part of the optimal tariff combination of Himics and Britz (2016), originally developed for flexible welfare consistent aggregation over exporter countries. The TE aggregator includes an import demand system at the tariff line level, driving the simulated changes in the composition of commodity groups and the related trade expenditures. As proposed here, the TE aggregator also includes explicit TRQ functions at the tariff line level to take into account the adjustments in applied tariff rates under TRQ. The TE aggregator therefore directly addresses two of the investigated sources of aggregation bias: (i) changes in the composition of the imported consumption bundle; and (ii) TRQ regime shifts. With regard to the over protective part of the duty, the TE aggregator takes the tariff “water” into account but sets the initial quota rent on the basis of ad-hoc rules (arbitrarily). The strategy in the second approach adopted in this paper is to include some of the methodological elements of the consistent aggregation into weighted average tariff calculations. Aggregation weights are adjusted to the changes in consumption shares within a commodity group. The proposed Tariff Reduction Impact Model for Agriculture (TRIMAG) aggregator adjusts aggregation weights based on a CES aggregator function, defined at the tariff line level. TRIMAG also explicitly takes into account the overprotective part of the duty by calculating the price impacts of tariff cuts at the tariff line level. A price transmission function in TRIMAG calculates import price changes, taking into account the initial overprotective part of the duties estimated from a Switzerland-specific database of wholesale domestic prices and CIF prices at the HS 8 level (Listorti et al., 2013). Consequently, TRIMAG addresses the aggregation bias stemming from both substitution effects and from TRQ regime switches. Furthermore, in the case of binding TRQs, the above domestic price database enables TRIMAG implementing a data driven approach to estimate the overprotective part of the duty and take this into account in the price transmission from tariff cuts to domestic import prices. Both the TE and the TRIMAG aggregators neglect the income effect of tariff revenues that might accrue to consumers. The assumption of ignoring tariff revenues does not seem to be too restrictive in the PE framework of this paper, and, more generally, if tariff revenues has little impact on consumer income.7 In addition, both aggregators are built for the small country case and assume completely elastic excess supply from the rest of the world, at the tariff line level (see section 3.2). Following a standard assumption of the consistent aggregation literature, prices of domestically produced and imported goods are independent, simplifying the consumers’ utility maximisation (expenditure minimisation) problem for finding an optimal consumption bundle after trade liberalisation. The features of the proposed aggregators, including a traditional fixed weight aggregator for comparison, are summarised in Table 1. The two aggregators proposed were tested by analysing alternative Swiss tariff dismantling scenarios for beef imports, following a typical two-stage approach (Francois et al., 2005; Philippidis and Sanjuán, 2007; Egger et al., 2015 to name a few). First aggregated ad valorem equivalent (AVE) tariff rates were estimated using disaggregated data at the tariff line level. Second, the overall impact on agricultural markets was analysed by plugging in the estimated AVEs in a PE model working at a more aggregated level (e.g. commodity level). More precisely, the TE and TRIMAG aggregators, and additionally a traditional fixed weight aggregator for comparison, provide a set of aggregated AVE tariff rates for the Common Agricultural Policy Regionalised Impact (CAPRI) model (e.g. Britz and Witzke, 2014) before and after trade liberalisation, making a systematic and direct comparison of the approaches possible. In the first stage, aggregated tariffs are calculated using the different methods, which then enter the calibration process of CAPRI in the first step. The outcome of the first stage is the set of three reference scenarios (one corresponding to each aggregation method). The assumed tariff cuts are applied in the second stage, where the liberalised tariff schedule is used to calculate aggregate tariffs again post-liberalisation. The aggregated tariffs are transferred into CAPRI to simulate trade (market balances and prices) and welfare impacts for the trade dismantling scenarios. The trade expenditure (TE) tariff aggregator and the TRIMAG tariff aggregator are discussed in detail in the following Sections 2.3 and 2.4, respectively. The trade expenditure (TE) tariff aggregator ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ With respect to the link to the CAPRI model, the aggregate (importer-specific) tariff rates of equation (6) are implemented in CAPRI as AVE tariff rates. As a result, the endogenous modelling of the TRQ system in the TE aggregator is shifted from CAPRI to a standalone aggregation module. Shifting policy representations out of the core (aggregated) model has various numerical advantages as the number of equations decreases. On the other hand, removing the (beef) TRQ function from CAPRI breaks the link between aggregated quota thresholds and tariff rates, possibly altering simulated trade impacts. The TRIMAG tariff aggregator ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The TRIMAG model, developed by the Swiss Federal Office for Agriculture (FOAG) (Listorti et al., 2013), calculates aggregate tariffs based on current tariff data (reference mode) as well as considers possible tariff cut scenarios (simulation mode). A particularity of the Swiss tariff schedule is that in- and out-of-quota tariff rates of the same TRQ are registered under different tariff lines.8 Following this separation of tariff rates, TRIMAG calculates aggregate in- and out-of-quota rates separately. Tariff rates are also differentiated by the geographical region of origin (EU and rest of the world RW9). The following sections provide more methodological detail on the reference and simulation modes. The reference mode TRIMAG combines three weighting methods in the reference mode to reduce the endogeneity bias related to small traded volumes of highly protected tariff lines (see discussion in Section 1): (i) bilateral import weighted average, which gives a large weight to the most important trading partners; (ii) total imports weighted average, which puts the largest weights on tariff lines with the largest import demand within the aggregated commodity groups; (iii) a simple arithmetic average, which is not affected by traded volumes and is therefore free from the endogeneity bias. The simulation mode In simulation mode, TRIMAG updates the aggregation weights according to the assumed change in the current tariff schedule. Unlike the linear aggregation in equation (17), the updated weights enter a CES-type tariff aggregator, which calculates the modified aggregate tariffs. Once the updated import prices at the 8-digits level have been determined, they enter a CES utility aggregator, which drives substitution within the consumption bundle, corresponding to a standard consumer expenditure minimisation framework under fix utility assumption. Consequently, the adjustment in the consumption mix is triggered by relative price changes at the tariff line level. The adjustment in the consumption mix (i.e. substitution effect) is therefore driven by the equation system (19)–(21). Note that a demand system has not been formally defined at the tariff line, as it is in consistent aggregation, but only adjust aggregation weights depending on a CES- type aggregator.15 Instead of using CES share equations and the associated composite price indexes as is typical e.g., in the CGE modelling literature, a theoretically equivalent formulation based on a numeraire tariff line was used in this paper. The price transmission function in equation (18) follows a standard price gap approach based on fix world (CIF) prices, but takes a possible overprotective part in the duties under binding TRQs into account. Tariff cuts are only transmitted to the import prices if the overprotective part under TRQs has been eroded. However, possible pricing behaviour of exporting firms is not considered in the analysis. Exporting firms might pass through tariff rate changes to the import markets to varying extent, adjusting their mark-ups and pricing their exported goods depending on the tariff protection of the export destination (e.g., see Mallick and Marques, 2008, 2012). The aggregated in- and out-of-quota tariff rates in the empirical application below are transferred directly16 to the CAPRI model, and there they are used to parameterize the beef TRQ equations.","The following section provides a short description of the CAPRI modelling system used for the scenario analysis of the beef tariff dismantling scenarios. The focus is only on those aspects of CAPRI that are relevant for the simulation exercise presented in this paper. A more detailed description can be found e.g. in Britz and Witzke (2014). CAPRI is a global comparative-static deterministic partial equilibrium model with a focus on European agriculture. Nevertheless, CAPRI includes a global market module covering the main agricultural and food commodities. The market module covers 77 countries or country aggregates in 40 trade blocks and about 50 products. The model follows the Armington approach for simulating bilateral trade flows, taking into account the price impacts of bilateral and multilateral trade policy instruments (including ad valorem and specific tariffs and tariff-rate quotas). The market model consists of structurally identical template equations for all regions and commodities. Regional and commodity-wise specificities are expressed by the differences in parameterization. Supply and demand equations are consistent with microeconomic theory by imposing homogeneity and other curvature conditions during calibration. The supply of agricultural and feed compound sector are derived from a Normalized Quadratic profit function while final demand is based on Generalized Leontief demand systems (Diewert, 1971). For Switzerland, in order to improve the empirical foundation of the supply response, the supply elasticities are based on estimates derived through sensitivity analyses carried out using the SWISSland agent based model (Möhring A. et al., 2016). The CAPRI model outputs comprises several economic indicators which include welfare figures based on a partial equilibrium setting for ceteris paribus conditions for the sectors outside agriculture. The CAPRI model has been applied for trade-related policy impact analysis in the literature (e.g. Burrell et al., 2011; Burrell et al., 2014; Pelikan et al., 2015). However, for the sake of this study, CAPRI was extended with tariff aggregation modules17 calculating aggregate tariffs for CAPRI in both calibration and simulation. In the case of the TE and fixed weight aggregators the model-endogenous TRQ instrument also had to be shifted from CAPRI to the standalone aggregation modules.","Beef production represents slightly more than a quarter of the total Swiss agricultural production in value terms (BLW, 2017). A self-sufficiency rate of around 80% renders Switzerland a net importer. Meat markets are regulated by a multilateral TRQ system with a large number of sub-quotas for specific commodities (Loi et al., 2016). The multilateral TRQ for red meats includes 22 in-quota and 23 out-of-quota sub-quota tariff lines18 for beef products. The TRQ covers a very heterogeneous product group ranging from live animals to fresh or frozen carcasses, fresh or frozen meat with bones in or boneless, and offal, with a total notified volume of 22,500 tonnes. Out-of-quota quota tariffs are often prohibitively high, and therefore imports mostly occur within the quota limit.19 See also Loi et al. (2016) for more details. The proposed tariff aggregation methodologies were implemented at the 8-digits level, therefore taking advantage of all of the detail in the Swiss tariff schedule, which is also defined at that level of detail. Although Swiss beef TRQs are normally totally filled, assumptions on the initial unit quota rents still need to be made. Theoretically, the unit quota rent can be set to anything between the in- and out-of-quota rates in a binding TRQ regime (Skully, 2001). The TRQ functions of the TE aggregator are built on complementarity slackness conditions mimicking potential TRQ regime shifts at the tariff line level. The calibration of the system still requires an assumption on the initial quota rents, as initial import prices faced by consumers include the quota rent. Initial unit quota rents for the TE aggregator are uniformly set to the maximum for all tariff lines, reflecting a generally strong Swiss import demand for beef commodities. On the other hand, TRIMAG follows a data-driven approach and calculates initial unit quota rents as the difference between wholesale domestic and the corresponding CIF prices (plus in-quota tariff if applicable). Traditional fixed weight aggregate tariffs are also calculated for comparison. More precisely, an import weighted aggregator that is quite standard in the literature was selected in this paper. This follows the approach of Bouet et al. (2008) in calculating applied tariff rates under TRQ: (i) if the fill rate is between 95% and 99%, the applied rate is an arithmetic average of the in-quota and out-of-quota rates; (ii) if the fill rate is below 95%, the in-quota rate is taken; and (iii) above a 99% fill rate, the out-of-quota rate is calculated. To avoid mixed impacts which would complicate a systematic comparison of simulated results, the tariff dismantling scenarios presented in this paper are designed to be simple. Two liberalisation scenarios are implemented. All bound rates for Swiss beef imports from the EU in the homogeneous cuts scenario are cut by 50% at the 8-digits HS level. More heterogeneity regarding the tariff cuts is introduced in the second (heterogeneous cuts) scenario by making two tariff lines (0201.3099, fresh boneless beef meat, and 0202.3099, frozen boneless beef meat) exempt from the general 50% tariff cuts. Although the two exempted tariff lines are characterized by similar tariff protection, they have significantly different aggregation weights in the reference scenario, and the overprotective part in the duties varies considerably. The heterogeneity in tariff rates and tariff cuts as well as the presence of significant overprotection in tariffs makes these scenarios particularly useful in evaluating the different tariff aggregators. In fact, it is expected that the three aggregation biases discussed in this paper will be relevant to the scenarios by design. Firstly, initial aggregation weights for a product group are not uniformly distributed, and so the cuts on the tariff lines have different impact on the aggregate tariff after liberalisation. The heterogeneity in initial import prices induces significant substitution effects, including in the case of uniform tariff cuts, as relative price changes after liberalisation might differ substantially. Making the tariff cuts heterogeneous may further increase this substitution effect. Secondly, explicitly considering TRQ regime shifts at the tariff line is expected to provide more precise aggregate tariff cuts. The increased precision was assessed by comparing the calculated aggregated tariff cuts across the aggregation approaches. As an additional methodological difference, TRIMAG calculates aggregated in- and out-of-quota rates, and does not convert the whole TRQ mechanism into an equivalent tariff rate as the other two aggregators do. Therefore, using TRIMAG aggregated tariffs means the aggregated (beef) TRQ function in CAPRI can be taken advantage of. Comparing TRIMAG simulated results to the other two aggregators therefore provides an estimate of the bias that might be introduced by substituting the TRQ mechanism with a simple equivalent tariff rate. Finally, comparing TE and TRIMAG results sheds light on the importance of estimating the overprotective part in the duties using domestic price data.","In this section both the input data used for the modelling exercise are presented and the simulation results derived with the three tariff aggregators are systematically compared. Data base construction ~~~~~~~~~~~~~~~~~~~~~~ Import values and quantities from the Swiss-Impex database (Swiss-Impex, 2015) were used for the TE tariff aggregator. The data were extracted at 8-digits level and averaged over the years 2009–2014. Exporter countries are mapped and potentially aggregated to the CAPRI regions before setting up the demand system of the TE aggregator, so the exporter-specific aggregate tariffs of equation (6) can be plugged into CAPRI directly. In-quota and out-of- quota tariffs in the Swiss tariff schedule are registered under different tariff lines. In order to calibrate the TRQ functions of the equation system (9)–(12), all out-of-quota tariff lines had to be paired with their corresponding in-quota tariff lines. Therefore, the TE aggregator covers 22 (merged) TRQs for Swiss beef imports. The standard (fixed weight) aggregated tariffs are calculated based on the same Swiss-Impex data and the same merged TRQ system. The TRIMAG aggregator has its own trade and tariff database at the 8-digit level, with a base year defined as an average of the 2004-09 in the case presented in this paper. The TRIMAG database includes bound and applied tariffs, imports values and quantities, as well as CIF prices for the main trading partners20 and domestic wholesale Swiss prices.21 Beef related TRIMAG data used in this analysis are illustrated in Table 2. Tariff cut calculations (stage one) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the first stage of the modelling exercise, the aggregate beef tariff cuts are calculated according to the three tariff aggregation methods, each of them implemented as pre-model aggregation module to CAPRI. Differences in calculated aggregate tariffs are already present in the corresponding reference scenarios (Table 3). TRIMAG calculates an aggregated in-quota rate of 15% and an aggregated out-of-quota rate of 145%. The 142% AVE in Table 3 is the marginal tariff rate that the aggregated beef commodity in CAPRI is calibrated to. The large AVE, close to the out-of-quota rate, implies a strong import demand and a high unit quota rent respectively in the reference scenario. The TE aggregator calculates a significantly lower (83%) aggregate AVE rate for imports from the EU. The AVE for beef calculated with the fixed weight aggregator is between the other two aggregation approaches (99%). Most TRQs at the tariff line level are close to the 100% fill rate, and so the fixed weight aggregator sets the unit quota rent under these TRQs at exactly the midpoint between in- and out-of-quota rates during the tariff aggregation. Using fixed aggregation weights in the homogeneous cuts scenario, a −50% uniform tariff cut on all tariff lines simply translates to a −50% aggregated tariff cut. By allowing aggregation weights to change, the aggregate tariff cut on EU imports departs from the −50% middle value, and respectively becomes −48% and −51% for the TE and TRIMAG aggregators in the homogenous scenario. The similar estimates suggest only marginal relative price changes and so only small adjustments in the consumption mix due to substitution effects in the case of uniform tariff cuts. However, introducing heterogeneity by making two important tariff lines exempt from tariff cuts (heterogeneous cut scenario) opens the gap between estimated aggregate tariff cuts. Here, larger relative price changes induce more adjustments in the consumption mix toward commodities with higher initial tariffs. The aggregate tariff cut therefore becomes relatively smaller when calculated with tariff aggregators that take the substitution at the tariff line into account. The explicit TRQ modelling in the TE aggregator adjust applied tariff rates under TRQ upwards as imports reach the quota threshold. This upward adjustment increases the aggregate tariff rate in the scenarios, further decreasing the aggregate tariff cut compared to standard approaches. Accordingly, aggregate cuts in the heterogeneous cuts scenario are significantly smaller using the TE and TRIMAG approaches. In the TE aggregator, the average tariff cut is only −18% for the EU. Even though the consumption shares of the exempted tariff lines decrease, their initially high relative shares and large initial tariff rates keep the average tariff at a higher level than in the homogeneous cuts scenario. The relative average tariff cut calculated by TRIMAG (−39%) is also lower compared to the fixed weight aggregator, again mainly due to the large initial tariffs and trade shares of the exempted tariff lines. The marginal tariff rates reported for TRIMAG in Table 3 coincide with the out-of-quota tariff rates in both scenarios as the increasing imports push the TRQ regime to the out-of-quota situation. Regarding the exempted tariff lines, one of them has a large aggregation weight and therefore reduces the average tariff cut substantially, while the other tariff line has a much lower weight and so has almost no impact on the aggregated cuts (Fig. 2). It is important to note here that such a difference in tariff cuts between the homogeneous and heterogeneous tariff cut scenarios is not found for the fixed weight aggregator. Aggregation weights do not change for the two exempted tariff lines, limiting the impact on the aggregated tariff cut compared to aggregation techniques explicitly addressing the substitution effect (−46%). CAPRI simulation results (stage two) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the second stage of the modelling exercise, the average beef tariff cuts are introduced in CAPRI, and the respective comparative static simulations are performed. The smaller aggregate tariff cuts of the first stage induce smaller simulated trade impacts in CAPRI in the second stage for the TE and TRIMAG aggregators, specifically in the heterogeneous cuts scenario. According to the TE aggregator the EU (dominating Swiss beef imports) countries increase their exports by 65% in the homogeneous and by 22% in the heterogeneous cuts scenario results whereas using TRIMAG an increase of 64% and 45% respectively was simulated (Table 4). In contrast, the fixed weight aggregator only delivers marginal differences in simulated trade impacts between the two scenarios (70% and 63%). The larger import price effects of TRIMAG in the second stage are largely due to the explicit beef TRQ functions in CAPRI, which allow for a more flexible price adjustment compared to the ad valorem representation of TRQs used by the TE and standard aggregators. Firstly, traded quantities are too small to have significant impact on the beef market prices at the source of origin, therefore import price effects are mostly defined by changes in the applied tariff rates. Applied tariff rates, on the other hand, tend to drop more in trade liberalisation scenarios using explicit TRQ functions, because the applied rate may settle in the new equilibrium between the in- and out-of-quota tariff rates. Secondly, non-EU import prices only decrease for TRIMAG, with an additional negative impact on the average beef import price. The impact on import prices from non-EU countries is due to the explicit TRQ functions on these bilateral trade links (Brazil, Argentina and USA all take advantage of a multilateral beef TRQ). As demand for non-EU beef in Switzerland is decreasing, the unit quota rents (and therefore the shadow prices) for imports from non-EU countries decrease as well. Regarding the different import price impacts calculated with the fixed weight and TE aggregators (ad valorem representation of TRQs in both cases), the effects closely follow the relative differences in tariff cuts of stage one (−23% and −19% decrease in the average import price in the homogeneous scenario, respectively). As observed for the tariff cuts above, the heterogeneous cuts scenario results in smaller impacts also on import prices, for all three aggregation methods. Average import prices decrease by −20% for the fixed weight aggregator, −7% for the TE aggregator, and −23% for TRIMAG. The above import price changes are transmitted to the producer prices in CAPRI, taking into account both the substitution of domestically produced goods with imported ones as well as the substitution of imported goods from different origin in the consumption bundle. Consistently to the effects on average import prices, simulated impacts on producer prices are the largest for the TRIMAG approach (Fig. 3). In the homogenous tariff cuts scenario a −29% decrease in the average beef import price induces a −16% reduction in the domestic producer price with the TRIMAG approach. The price transmission is similar in relative terms for the other two aggregators as well: from a −19% change in the import price into a −9% decrease in the producer price for the TE aggregator; and from a −23% to a −12% decrease for the standard aggregator. The adjustment in Swiss beef supply in CAPRI is driven by the respective supply elasticities, empirically estimated using the SWISSland agent based model (Möhring A. et al., 2016). The supply response for most agricultural products is inelastic in the medium run (including beef), causing beef supply to decrease less than proportionally to simulated producer price changes. Producer price and production impacts are, in general, smaller in the heterogeneous scenario. At the same time, more heterogeneity in tariff cuts widens the relative differences in simulated impacts between the TE and standard aggregators. The equivalent variation measure of consumer welfare increases in all scenarios and under all aggregation methodologies (see consumer surplus in Fig. 4). This is a standard result for tariff reduction in Armington models, and is due to decreasing consumer prices for beef. Producer surplus, on the other hand, decreases due to the reductions in beef producer price and domestic supply. Consumer and producer surplus changes point to the same direction using all tariff aggregation approaches. However, we also note that there are key differences in the magnitude of net welfare increases. In fact, TRIMAG systematically delivers the highest impacts regarding all items of the welfare calculation. For the homogenous cuts scenario, the net welfare impact of TRIMAG (90 million EUR) is almost twice as high as the net welfare impact computed relying on the fixed weights approach (50 million EUR). Simulated net welfare gains are the lowest using the TE approach while results using the fixed weight approach are between the other two aggregators. The order of simulated impacts is consistent with the impacts on import prices, domestic prices and production as presented above. Introducing more heterogeneity in the tariff schedule enlarges the gap between simulated impacts across the aggregators. For example, in the homogeneous cuts scenario the increases in simulated consumer surplus with the TE aggregator are only marginally below the simulated changes using the fixed weight aggregator. In contrast, the impact in the heterogeneous cuts scenario is more than three times larger using the fixed weight aggregator than using the TE approach. When comparing simulated welfare changes in the two scenarios, we note that making two tariff lines exempt from the cuts obviously decreases welfare gains. Returning to the previous example, simulated increase in consumer surplus for the TE aggregator in the heterogeneous cut scenario is only about one third of the impact in the uniform cut scenario. The impacts on consumer surplus, producer surplus and net welfare consequently mirror the different impacts on the import prices between the tariff aggregators. Again, simulations with the TRIMAG tariff cuts provide the largest impacts, followed by the fixed weight and TE approaches. Finally, it should be noted that whether an ad valorem equivalent approach (fixed weight and TE aggregators) or an explicit TRQ function approach (TRIMAG) is used for modelling TRQs strongly affects the simulated tariff revenue impacts. Using the ad valorem equivalent approach scenario, impacts on tariff revenues are negative because increasing beef imports (quantity effect) do not compensate for the decrease in tariff rates (price effect). With the explicit TRQ function approach of TRIMAG, the increase in imports pushes the TRQ regime into an overfill position, and therefore beef imports above the quota limit face reduced out-of-quota duties. That generates an increase in tariff revenues in both scenarios. On the other hand, the decrease in the out-of-quota duties has a significant impact on the quota rents. Quota rents decrease in both scenarios, indicating that the overprotective part of the duty is fully eroded after tariff cuts and consequently the unit quota rent decreases22. Two considerations concerning quota rents and tariff revenues are relevant here. Firstly, the magnitude of the tariff revenue and quota rent changes is significantly smaller than the changes in producer and consumer surplus (for all aggregators and scenarios). Secondly, in TRIMAG, the increase in tariff revenues and the decrease in quota rents almost neutralize each other (the combined effect remains below 2 million euros in both scenarios). Although the impacts on tariff revenues and quota rents may be significant for government revenues, they are not decisive for the direction of net welfare changes. Net welfare change is mostly defined by the impacts on import and domestic prices through their effect on consumer and producer surplus.","Market access policies, and import tariffs in particular, are defined at a finer product detail than the one at which applied equilibrium models operate. This requires aggregating data on trade and market access, which leads both to an aggregation bias in the simulated results and to a departure from exact implementation of negotiated tariff schedules. This paper develops a two-stage modelling approach, based on the TE and TRIMAG tariff aggregators, to reduce aggregation bias in applied trade modelling. Three sources of aggregation bias are focussed on: substitution effects at the tariff line, TRQ regime shifts, and overprotective part of the duty. The tariff aggregators are tested against tariff dismantling scenarios for the Swiss beef market, including a standard (fixed weight) aggregator for comparison purposes. The three above-mentioned sources of aggregation bias characterize many agri-food markets and particularly the Swiss beef market since: a) the product heterogeneity of the beef tariff lines is sufficiently large to induce large substitution effects; b) the market is regulated by a complex system of TRQs; and c) the overprotective part of the duties varies to a great extent between the beef tariff lines. The simulation results indicate that standard (fixed weight) aggregators do not appropriately take into account either substitution effects at the tariff line level, or TRQ regime shifts, or possible overprotection in duties, and so they tend to provide biased estimates of the gains from trade liberalisation in terms of simulated trade and welfare implications. Firstly, we find that the standard aggregator overestimates the trade liberalisation impacts on traded quantities, specifically when tariff cuts are heterogeneous. This is partly due to ignoring the substitution effect on the import demand, which leads to larger tariff cuts on the bilateral trade links. Ignoring TRQ regime shifts further contributes to the overestimated tariff cuts derived using the standard method, confirming previous findings in the literature (Himics and Britz, 2016). Secondly, simulated impacts with the standard aggregator on import prices, producer prices, production and welfare are found to be in between the impacts derived using the TE and TRIMAG approaches. This result shows a considerable bias in simulated welfare gains using standard methods, and drives the attention to the importance of TRQ modelling. On the one hand, net welfare impacts are reduced to one-third by using the TE aggregator, which shifts TRQ modelling from the core PE model to the standalone tariff aggregation module at the tariff line level. On the other hand, net welfare impacts are even increased by using the TRIMAG approach, which keeps the explicit beef TRQ functions (e.g. commodity level) of the core PE model. Thirdly, by comparing result of the homogenous and heterogeneous tariff cut scenarios, it was demonstrated that the aggregation bias increases with greater dispersion in tariffs or tariff cuts. Real-life tariff schedules typically show large tariff dispersion, which is often increased by trade negotiations conducted at the tariff line level. This finding supports the relevance of the proposed aggregators for direct policy support. Fourthly, the different simulated impacts between the TE and TRIMAG aggregator suggest that information on the overprotective part of the duties further reduces aggregation bias. Consequently, using data at detailed product level to estimate the height of the unit quota rent under a TRQ regime is well warranted, and further motivates statistical data collection and data processing directly on observable unit quota rents and on domestic prices. A final consideration concerns the technical implementation of the TE and TRIMAG aggregators as standalone modules linked to a large-scale PE model (CAPRI). Linking the aggregation modules to CAPRI does not require significant changes in the original model. In fact, the trade policy representation simplifies in CAPRI for the TE approach by shifting TRQ modelling into the standalone tariff aggregation module. Given that the proposed tariff aggregation approaches only simulate demand side adjustments, the respective aggregation modules are relatively easy to set up and solve numerically, and detailed trade statistics and policy data are available to parameterize them. Against this background we can state that the proposed tariff aggregation techniques can be implemented in existing modelling systems without the heavy burden of modifying core model structures and without jeopardizing the numerical feasibility of the overall modelling system. Consequently, more widespread use of advanced tariff aggregation techniques for applied trade modelling seems to be within reach. The above results support the findings of Countryman and Narayanan (2017), which highlight that the simplified ad valorem representation of trade policy measures, such as specific tariffs in their case or TRQs in this paper, lead to biased simulated price effects on the domestic markets. The biased welfare results, identified above, have been also highlighted in the literature. Gouel et al. (2011), for example, simulate smaller welfare gains after extending their CGE modelling framework to the product level and considering tariff line level exceptions (sensitive products). Our strategy to better exploit tariff line level data by using standalone tariff aggregation modules is in line with Britz and van der Mensbrugghe (2016) who advocates avoiding aggregation in model databases the utmost to reduce the induced bias in ex ante policy evaluation. However, several caveats to the proposed approaches must be emphasised. Both the TRIMAG and the TE aggregators only account for the consumption gains from trade and not for any production or specialization gains. Whether this is an acceptable restriction remains case specific. In the empirical application presented in this paper, the potential consumption gains from liberalising the beef trade largely outweigh the losses linked to domestic production, and therefore the assumption is viable. Although the test scenarios intentionally focus on a specific market and country, the approach also has the potential to correct significant bias stemming from heterogeneous tariff cuts in more complex and broader (global) applied trade liberalisation scenarios. With the proposed tariff aggregators, negotiated tariff schedules can be introduced to applied equilibrium models directly, without first aggregating tariff cuts over tariff lines. Our simulation results suggest that they also reduce aggregation bias, consequently enhancing the potential contributions of trade modelling to evidence based policy making.","The views expressed in this article are purely those of the authors and may not in any circumstances be regarded as stating an official position of the European Commission or the Federal Office for Agriculture (FOAG). The work was carried out while Giulia Listorti was Scientific Advisor at the Swiss Federal Office for Agriculture."],["We address the curse of dimensionality in dynamic covariance estimation by modeling the underlying co-volatility dynamics of a time series vector through latent time-varying stochastic factors. The use of a global–local shrinkage prior for the elements of the factor loadings matrix pulls loadings on superfluous factors towards zero. To demonstrate the merits of the proposed framework, the model is applied to simulated data as well as to daily log-returns of 300 S&P 500 members. Our approach yields precise correlation estimates, strong implied minimum variance portfolio performance and superior forecasting accuracy in terms of log predictive scores when compared to typical benchmarks. --------------------------------------------------------------------------------","The joint analysis of hundreds or even thousands of time series exhibiting a potentially time-varying variance–covariance structure has been on numerous research agendas for well over a decade. In the present paper we aim to strike the indispensable balance between the necessary flexibility and parameter parsimony by using a factor stochastic volatility (SV) model in combination with a global–localshrinkage prior. Our contribution is threefold. First, the proposed approach offers a hybrid cure to the curse of dimensionality by combining parsimony (through imposing a factor structure) with sparsity (through employing computationally efficient absolutely continuous shrinkage priors on the factor loadings). Second, the efficient construction of posterior simulators allows for conducting Bayesian inference and prediction in very high dimensions via carefully crafted Markov chain Monte Carlo (MCMC) methods made available to end-users through the R (R Core Team, 2017) package factorstochvol (Kastner, 2017). Third, we show that the proposed method is capable of accurately predicting covariance and precision matrices which we assess via statistical and economic forecast evaluation in several simulation studies and an extensive real-world example. Concerning factor SV modeling, early key references include Harvey et al. (1994), Pitt and Shephard (1999) and Aguilar and West (2000) which were later picked up and extended by e.g. Philipov and Glickman (2006), Chib et al. (2006), Han (2006), Lopes and Carvalho (2007), Nakajima and West (2013), Zhou et al. (2014) and Ishihara and Omori (2017). While reducing the dimensionality of the problem at hand, models with many factors are still rather rich in parameters. Thus, we further shrink unimportant elements of the factor loadings matrix to zero in an automatic way within a Bayesian framework. This approach is inspired by high-dimensional regression problems where the number of parameters frequently exceeds the size of the data. In particular, we adopt the approach brought forward by Caron and Doucet (2008) and Griffin and Brown (2010) who suggest to use a special continuous prior structure – the Normal-Gamma prior – on the regression parameters (in our case the factor loadings matrix). This shrinkage prior is a generalization of the Bayesian Lasso (Park and Casella, 2008) and has recently received attention in the econometrics literature (Bitto and Frühwirth-Schnatter, 2018; Huber and Feldkircher, 2017). Another major issue for such high-dimensional problems is the computational burden that goes along with statistical inference, in particular when joint modeling is attempted instead of multi-step approaches or rolling-window-like estimates. Suggested solutions include Engle and Kelly (2012) who propose an estimator assuming that pairwise correlations are equal at every point in time, Pakel et al. (2017) who consider composite likelihood estimation, Gruber and West (2016) who use a decoupling–recoupling strategy to parallelize estimation (executed on graphical processors), Lopes et al. (2018) who treat the Cholesky-decomposed covariance matrix within the framework of Bayesian time- varying parameter models, and Oh and Patton (2017) who choose a copula-based approach to link separately estimated univariate models. We propose to use a Gibbs-type sampler which allows to jointly take into account both parameter as well as sampling uncertainty in a finite-sample setup through fully Bayesian inference, thereby enabling inherent uncertainty quantification. Additionally, this approach allows for fully probabilistic in- and out-of-sample density predictions. For related work on sparse Bayesian prior distributions in high dimensions, see e.g. Kaufmann and Schumacher (2018) who use a point mass prior specification for factor loadings in dynamic factor models or Ahelegbey et al. (2016) who use a graphical representation of vector autoregressive models to select sparse graphs. From a mathematical point of view, Pati et al. (2014) investigate posterior contraction rates for a related class of continuous shrinkage priors for static factor models and show excellent performance in terms of posterior rates of convergence with respect to the minimax rate. All of these works, however, assume homoskedasticity and are thus potentially misspecified when applied to financial or economic data. For related methods that take into account heteroskedasticity, see e.g. Nakajima and West (2013, 2017) who employ a latent thresholding process to enforce time-varying sparsity. Moreover, Zhao et al. (2016) approach this issue via dependence networks, Loddo et al. (2011) use stochastic search for model selection, and Baştürk et al. (2018) use time-varying combinations of dynamic models and equity momentum strategies. These methods are typically very flexible in terms of the dynamics they can capture but are applied to moderate dimensional data only. We illustrate the merits of our approach through extensive simulation studies and an in-depth financial application using 300 S&P 500 members. In simulations, we find considerable evidence that the Normal-Gamma shrinkage prior leads to substantially sparser factor loadings matrices which in turn translate into more precise correlation estimates when compared to the usual Gaussian prior on the loadings.1 In the real-world application, we evaluate our model against a wide range of alternative specifications via log predictive scores and minimum variance portfolio returns. Factor SV models with sufficiently many factors turn out to imply extremely competitive portfolios in relation to well-established methods which typically have been specifically tailored for such applications. Concerning density forecasts, we find that our approach outperforms all included competitors by a large margin. The remainder of this paper is structured as follows. In Section 2, the factor SV model is specified and the choice of prior distributions is discussed. Section 3 treats statistical inference via MCMC methods and sheds light on computational aspects concerning out-of-sample density predictions for this model class. Extensive simulation studies are presented in Section 4, where the effect of the Normal-Gamma prior on correlation estimates is investigated in detail. In Section 5, the model is applied to 300 S&P 500 members. Section 6 wraps up and points out possible directions for future research. Factor SV model ~~~~~~~~~~~~~~~ Third, note that even though the joint distribution of the data is conditionally Gaussian, its stationary distribution has thicker tails. Nevertheless, generalizations of the univariate SV model to cater for even more leptokurtic distributions (e.g. Liesenfeld and Jung, 2000) or asymmetry (e.g. Yu, 2005) can straightforwardly be incorporated in the current framework. All of these extensions, however, tend to increase both sampling inefficiency as well as running time considerably and could thus preclude inference in very high dimensions.","There are a number of methods to estimate factor SV models such as quasi-maximum likelihood (e.g. Harvey et al., 1994), simulated maximum likelihood (e.g. Liesenfeld and Richard, 2006; Jungbacker and Koopman, 2006), and Bayesian MCMC simulation (e.g. Pitt and Shephard, 1999; Aguilar and West, 2000; Chib et al., 2006; Han, 2006). For high dimensional problems of this kind, Bayesian MCMC estimation proves to be a very efficient estimation method because it allows simulating from the high dimensional joint posterior by drawing from lower dimensional conditional posteriors. MCMC estimation ~~~~~~~~~~~~~~~ The above sampling steps are implemented in an efficient way within the R package factorstochvol. Table 1 displays the empirical run time in milliseconds per MCMC iteration. Note that using more efficient linear algebra routines such as Intel MKL leads to substantial speed gains only for models with many factors. To a certain extent, computation can further be sped up by computing the individual steps of the posterior sampler in parallel. In practice, however, doing so is only useful in shared memory environments (e.g. through multithreading/multiprocessing) as the increased communication overhead in distributed memory environments easily outweighs the speed gains. The shrinkage prior effect: An illustration ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Fig. 1 shows smoothed kernel density estimates of posterior loadings under the different prior assumptions. The signs of the loadings have not been identified so that a multimodal posterior distribution hints at a “significant” loading whereas a unimodal posterior hints at a zero loading, see also Frühwirth-Schnatter and Wagner (2010). It stands out that only very little shrinkage is induced by the standard Gaussian prior. The other priors, however, impose considerably tighter posteriors. For the nonzero loadings on factor one, e.g., the row-wise Bayesian Lasso exhibits the strongest degree of shrinkage. Little difference between the various shrinkage priors can be spotted for the nonzero loadings on factor two. Turning towards the zero loadings, the strongest shrinkage is introduced by both variants of the Normal-Gamma prior, followed by the different variants of the Bayesian Lasso and the standard Gaussian prior. This is particularly striking for the loadings on the superfluous third factor. The difference between row- and column-wise shrinkage for the Lasso variants can most clearly be seen in row 9 and column 3, respectively. The row-wise Lasso captures the “zero-row” 9 better, while the column-wise Lasso captures the “zero-column” 3 better. Because of the increased element-wise shrinkage of the Normal-Gamma prior, the difference between the row-wise and the column-wise variant are minimal. In the context of covariance modeling, however, factor loadings can be viewed upon as a mere means to parsimony, not the actual quantity of interest. Thus, Fig. 2 displays selected time-varying correlations. The top panel shows a posterior interval estimate (mean plus/minus two standard deviations) for the correlation of series 1 and series 2 (which is nonzero) under all five prior settings; the bottom panel depicts the interval estimate for the correlation of series 9 and 10 (which is zero). While the relative differences between the settings in the nonzero correlation case are relatively small, the zero correlation case is picked up substantially better when shrinkage priors are used. Posterior means are closer to zero and the posterior credible intervals are tighter. Medium dimensional Monte Carlo study ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ An overall comparison of the errors under different priors is provided in Table 4 which lists RMSEs and MAEs for all prior settings, averaged over the non-trivial correlation matrix entries as well as time. Note again that results under the Lasso prior are sensitive to the particular choices of the global shrinkage hyperparameters as well as the choice of row- or column-wise shrinkage, which is hardly the case for the Norma-Gamma prior. Interestingly, the performance gains achieved through shrinkage prior usage are higher when absolute errors are considered. This is coherent with the extremely high kurtosis of Normal-Gamma-type priors which, while placing most mass around zero, allow for large values.","The presentation consists of two parts. First, we exemplify inference using a multivariate stochastic volatility model and discuss the outcome. Second, we perform out-of-sample predictive evaluation and compare different models. To facilitate interpretation of the results discussed in this section, we consider the GICS5 classification into 10 sectors listed in Table 6. A four-factor model for 300 S&P 500 members ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To keep graphical representation feasible, we only focus on the latest 2000 returns of our data set, i.e. 5/3/2006 to 12/31/2013. This time frame is chosen to include both the 2008 financial crisis as well as the period before and thereafter. Furthermore, we restrict our discussion to a four-factor model. This choice is somewhat arbitrary but allows for a direct comparison to a popular model based on four observed (Fama–French plus Momentum) factors. A comparison of predictive performance for varying number of factors is discussed in Section 5.2; the Fama–French plus Momentum model is introduced in Section 5.3. Median posterior factor loadings are visualized in Fig. 5. In the top panel it can be seen that all series significantly load on the first factor which consequently could be interpreted to represent the joint dynamics of the wider US equity market. Investigating the second factor, it stands out that due to the use of the Normal-Gamma prior a considerable amount of loadings are shrunk towards zero. Main drivers are all in the sector Utilities. Also, companies in sectors Consumer Staples, Health Care and (to a certain extent) Financials load positively here. Both the loadings on factor 3 as well as the loadings on factor 4 are substantially shrunk towards zero. Exceptions are Energy and Materials companies for factor 3 and Financials for factor 4. The corresponding factor log variances are displayed in Fig. 6. Apart from featuring similar low- to medium frequency properties, each process exhibits specific characteristics. First, notice the sharp increase of volatility in early 2010 which is mainly visible for the “overall” factor 1. The second factor (Utilities) displays a pre-crisis volatility peak during early 2008. The third factor, driven by Energy and Materials, shows relatively smooth volatility behavior while the fourth factor, governed by the Financials, exhibits a comparably turbulent volatility evolution. Predictive likelihoods for model selection ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Even for univariate volatility models evaluating in- or out-of-sample fit is not straightforward because the quantity of interest (the conditional standard deviation) is not directly observable. While in lower dimensions this issue can be circumvented to a certain extent using intraday data and computing realized measures of volatility, the difficulty becomes more striking when the dimension increases. Thus, we focus on iteratively predicting the observation density out-of-sample which is then evaluated at the actually observed values. Because this approach involves re-estimating the model for each point in time, it is computationally costly but can be parallelized in a trivial fashion on multi-core computers.","Further research could be directed towards incorporating prior knowledge into building the hierarchical structure of the Normal-Gamma prior, e.g. by choosing the global shrinkage parameters according to industry sectors. Alternatively, Villani et al. (2009) propose a mixture of experts model to cater for smoothly changing regression densities. It might be fruitful to adopt this idea in the context of covariance matrix estimation by including either observed (Fama–French) or latent factors as predictors and allowing for other mixture types than the ones discussed there. While not being the focus of this work, it is easy to extend the proposed method by exploiting the modular nature of Markov chain Monte Carlo methods. In particular, it is straightforward to combine it with mean models such as (sparse) vector autoregressions (e.g., Bańbura et al., 2010; Kastner and Huber, 2018), dynamic regressions (e.g., Korobilis, 2013), or time-varying parameter models (e.g., Koop and Korobilis, 2013; Huber et al., 2019)."],["Many river sharing agreements in transboundary river basins are inherently unstable. Due to stochastic river flow, agreements may be broken in case of drought. The objective of this paper is to analyze whether river sharing agreements can be self-enforcing, or sustainable. We do so using an infinitely-repeated sequential game that we apply to several classes of agreements. To derive our main results we apply the equilibrium concepts of subgame-perfect equilibrium and renegotiation-proof equilibrium to the river sharing problem. We show that, given the upstream-downstream asymmetry, sustainable agreements allow downstream agents to reap the larger share of the benefits of cooperation. --------------------------------------------------------------------------------","We extend and apply the theory of repeated games to the river sharing problem. Our main contribution is the design of agreements that are sustainable to stochastic river flow in a dynamic setting. Doing so, we add to the rapidly growing literature on the analysis of solutions to the river sharing problem (cf. Béal et al., 2013; van den Brink et al., 2012; Ambec et al., 2013), which has largely ignored dynamics and stochasticity. In an international river basin, when water is scarce, countries may exchange water for side payments (Dinar, 2006; Carraro et al., 2007). This type of exchange is generally formalized in a river sharing agreement. The aim of river sharing agreements is to increase the overall efficiency of water use. This increase in efficiency can be obstructed by the stochastic nature of river flow, because countries may find it profitable to break the agreement in case of drought (Dinar et al., 2015; Ward, 2013). A recent example is Mexico's failure to meet its required average water deliveries under the 1944 US-Mexico Water Treaty in the years 1992–1997 (Gastélum et al., 2009). Additional case study evidence on agreement breakdowns because of droughts can be found, for instance, in Barrett (1994) and Beach et al. (2000). Only a minority of current international agreements take into account the variability of river flow (de Stefano et al., 2012). Most agreements do not; they either allocate fixed or proportional shares, or they are ambiguous in their schedule for water allocation. Both the efficiency and stability (Bennett and Howe, 1998; Bennett et al., 2000; Ansink and Ruijs, 2008; Ambec et al., 2013) of such agreements may be hampered. These effects could be worsened by the impacts of climate change on river flow. In order to accommodate for stochastic river flow, Kilgour and Dinar (2001) developed a flexible river sharing agreement that provides an efficient allocation for every possible level of river flow. This agreement maximizes the overall benefits of water use, after which side payments are made such that each country benefits from cooperation. This flexible agreement assures efficiency, but not stability because it ignores the repeated interaction of countries over time. Countries have an incentive to defect from the agreement when the benefits of defecting outweigh the benefits of compliance. Note that there is no supra-national authority that can enforce this type of international agreements. This implies that a stable agreement has to be self-enforcing or sustainable, in the sense that each agent should have an incentive to comply with the agreement. In such a setting, application of repeated-game theory to the setting of river sharing seems natural, but to the best of our knowledge, this has not been done yet.1 Given the asymmetry imposed by the geography of the river, we adopt an infinitely-repeated sequential game, in which upstream agents move before downstream agents.2 The Folk Theorem for infinitely-repeated sequential games is a limit result on the discount factor. For practical purposes, the issue is firstly, given some discount factor, how to construct sustainable agreements, and secondly whether certain classes of agreements have properties that may be appealing for implementation by policy makers. For instance, we will assess the effects of restrictions on per-period payoffs and we will look at some disadvantages of fixed-payment agreements, which are common in practice. To derive our main results we apply the Folk Theorem to the river sharing problem using the equilibrium concepts of subgame-perfect equilibrium and renegotiation-proof equilibrium. We will see that, given the upstream–downstream asymmetry, sustainable agreements allow downstream agents to reap the larger share of the benefits of cooperation. This distribution of gains is the opposite of some papers that assess agreements on river sharing in a static setting (e.g. Ambec et al., 2013). Our results provide economic intuition for an empirical result (downstream states managing to negotiate a substantial share of upstream river water) that has, up till now, mostly been explained by political factors (Dinar, 2009; Katz and Moore, 2011). In the next section we introduce our model. We derive equilibrium conditions in Section 3, which we use in an example in Section 4. After an intermezzo in Section 5 on the trade-off between efficiency and stability, we continue in Sections 6 and 7 with a detailed analysis of actual agreements. Finally, in Section 8, we provide some concluding remarks. Two appendices contain proofs as well as detailed information on context and generalization of our analysis. Preliminaries ~~~~~~~~~~~~~ There are several focal allocations of river flow that will be used extensively in the remainder of the paper: the Nash allocation,4 the efficient allocation and the minmax allocation. Definitions of the Nash allocation and efficient allocation are given below, while the minmax allocation is the subject of Proposition 1. Nash allocation Because of increasing benefits of water use, the Nash allocation is evident in the absence of water trade or other types of agreements on water use. Efficient allocation With a strictly concave quasi-linear utility function as in (1), the efficient allocation is unique and equal to utilitarian welfare maximization as in e.g. Kilgour and Dinar (2001) and Houba et al. (2014). The interesting and non-trivial case occurs when water is scarce: Water scarcity Realizations without incentives to cooperate can be ignored because any efficient agreement will specify no cooperation for each of these realizations, which is trivially sustainable. Our model setup is consistent with much of the river sharing literature. The case of two agents (Ansink and Ruijs, 2008; Houba, 2008) provides no limitation for reasons similar to infinitely-repeated games with n players and full dimensionality. Full dimensionality is assured in river sharing problems because with n players the monetary transfers allow redistribution of transferable utility in all n utility dimensions. By focusing on two agents instead of the general case with more agents and more realistic river geographies (Ansink and Houba, 2012; van den Brink et al., 2012), we are able to avoid some complexity and excessive notation. It also provides a better understanding of the issues involved in repeated interaction of the agents over time in the presence of stochastic river flow. In Appendix A, we elaborate on the general case. River sharing agreements ~~~~~~~~~~~~~~~~~~~~~~~~ We proceed to describe the possibility of river sharing agreements between the two agents. Our focus is on what we call ‘implicit’ agreements. Such an agreement coincides with a set of agents’ strategies in the infinitely-repeated sequential game that can be sustained in equilibrium, and which improves upon the Nash allocation. It does not require outside enforcement. An agreement (and thereby the agents’ strategies) specifies the following three elements: An allocation rule for river flow; A payment rule for monetary transfers; and Punishment strategies in case one of the agents deviates from the agreement (which may also involve monetary transfers). In the infinitely-repeated sequential game, that we introduce below, both the allocation of water and the payment may be contingent on realized river flow as in Kilgour and Dinar (2001). In reality, however, we also observe lump-sum payments and simpler allocation rules, including fixed and proportional allocations (Ansink and Ruijs, 2008; Drieschova et al., 2008). Punishment strategies are essential for the sustainability of the agreement because, in absence of a supra-national authority, agreements are non-binding. Punishment strategies determine what happens upon deviation and range from simple trigger strategies to more advanced strategies that assure ex post credibility of the punishment. In actual agreements on river flow allocations, punishment strategies are often lacking, although many do contain clauses on conflict resolution (Beach et al., 2000; Ward, 2013). Minmax values ~~~~~~~~~~~~~ The Folk Theorem for infinitely-repeated sequential games states that any utility vector that yields each agent more than his minmax value can be supported as an SPE (subgame- perfect equilibrium) utility vector for sufficiently large discount factors (Wen, 2002). Following this reference, we start our analysis by deriving the minmax values in the river sharing problem. We do this in Appendix A, where we also discuss the general case (i.e. more than two agents) as well as the implications of assuming increasing benefit functions for the minmax values. We state the following result. For every realization (e1, e2), agent i's minmax value is bi(ei). The agents’ minmax values coincide with the unique Nash equilibrium utilities and can be supported as a stationary SPE outcome path for all discount factors δ ∈ [0, 1) by the stationary strategies ‘always play the Nash equilibrium’. Because no SPE strategy is worse than the minmax value, this stationary SPE not only supports the worst possible outcome, but it also punishes at least as harsh as more sophisticated stick-and-carrot strategies as in e.g. Abreu (1988). This result has an important implication in that the Folk Theorem can be derived within the class of trigger strategies, which simplifies the analysis (details are in Appendices A and B). Trigger strategies Both agents use trigger strategies in which deviation from cooperative play under the agreement is punished by non-cooperative play forever (i.e. the Nash allocation).","Any agreement that satisfies both (4) and (5) for each realization (e1, e2) is able to sustain cooperation. This requires a choice of allocation and payment rules that is sufficiently flexible such that the bounds are not violated for any possible realization of river flow. Note the asymmetry in these bounds with respect to the payments; the lower bound always depends upon the realization (e1, e2) whereas the upper bound is independent of this realization and stated in terms of expectation. This asymmetry is caused by the extensive form of the game, which requires agent 2 to provide a minimum compensation to agent 1 for passed water in the current period, which enters the current realization in the lower bound. In addition, both bounds contain terms that reflect the expected benefits of cooperation in future periods. In the next section we use these bounds to illustrate the choice of a payment rule given the unique allocation rule that maximizes utilitarian welfare.","In this section we show how bounds (4) and (5) can be used to construct a payment rule for monetary transfers given a fixed parameter value for the discount factor, using a simple illustrative example that will recur throughout the paper. In order to do so, we make two additional assumptions. Two realizations of river flow Efficient agreement Fig. 1 shows the polygon for selected parameter values, illustrating the range of possible combinations of payments that provide sustained cooperation under Assumptions 2–4. Any point in the graph represents a payment rule, but only those in the shaded area sustain cooperation. Somewhat counter-intuitively, Fig. 1 illustrates the possibility of a negative payment under one of the two possible realizations of river flow. This gives the striking possibility that agent 1 delivers water and a payment to agent 2. Obviously, a negative payment under one realization is accompanied by a relatively large positive payment under the alternative realization of river flow. This possibility of negative payments makes clear that, while theoretically sound, some sustainable payment rules may be inapplicable in real life situations. Fig. 2 shows a comparison of results for various values of the discount parameter δ. The figure illustrates that there is no feasible payment rule for low levels of the discount factor (for the parameter values in the figure, the threshold is δ = (5/7) ≈ 0.7). At the threshold, the bounds for the realization of low river flow first converge and below this threshold, they switch place and diverge such that there are no combinations of payments that provide sustained cooperation. The intuition for this result is standard in that a lower δ reduces the expected present value of the benefits of cooperation in all future periods, which are compared with the benefits of non-cooperation in the current period. Summarizing, given a fixed parameter value for the discount factor and given trigger strategies, the illustration in Fig. 1 shows how to construct sustainable agreements. Such agreements are not possible for a sufficiently low discount factor and some sustainable payment rules may be unrealistic if they include negative payments. Because it is not clear how such agreements could ever be put into practice, we assess several subsets of agreements that do not allow negative payments in Sections 6 and 7, in which we will also drop Assumptions 3 and 4.","In this section we consider two subsets of agreements and assess their sustainability: fixed-payment agreements and individually-rational agreements. The reasoning for assessing these particular subsets is as follows. First, most existing agreements with payments explicitly coupled to water deliveries employ fixed payments. Second – and related – fixed-payment agreements reflect the lack of flexibility with respect to variability in river flow (de Stefano et al., 2012; Giordano et al., 2014). Third, individually-rational agreements offer more flexibility and they exclude the unrealistic and objectionable option of payoffs lower than minmax payoffs (including negative payments as discussed in Section 4). Both subsets serve as a realistic benchmark for the design of more sustainable agreements in Section 7. Note that we will not need Assumptions 3 and 4. Fixed-payment agreements ~~~~~~~~~~~~~~~~~~~~~~~~ By Assumption 1 on water scarcity, the left-hand side is positive for xc(e1, e2) = x*(e1, e2) and these belong to a well-defined compact set of agreements xc(e1, e2) that admit a non-negative left-hand side. For the subset of agreements xc(e1, e2) that admit a positive left-hand side, as δ goes to 1, the left-hand side converges to some positive number while the right-hand side converges to 0, so that the inequality holds. Therefore, we know that for this particular subset of agreements xc(e1, e2) there exists a threshold discount factor above which the range of sustainable payment rules is non-empty. Obviously, agreements xc(e1, e2) for which the left-hand side is either zero or negative cannot be sustained for any δ ∈ [0, 1). One may interpret fixed-payment agreements as insurance contracts where downstream pays a fixed amount for a flexible scheme of water deliveries. In this interpretation, there is nothing against such agreements. One problem of fixed- payment rules, however, is that they may cause payoffs lower than minmax payoffs for some realizations of river flow. This property makes such rules unattractive for application in practice. We will see in Section 6.2 that, for the example of Section 4, any fixed-payment agreement violates this condition. Individually-rational agreements ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Whether these restrictions hold in practice is an open question. The concept of individual rationality directly suggests that countries would never sign an agreement that yields them payoffs lower than their minmax payoffs. Yet, this situation could occur for fixed payment agreements as is illustrated by Fig. 4. Even proportional allocation rules, often advocated as offering flexibility to adapt to variability of river flow (cf. McCaffrey, 2003; Drieschova et al., 2008) fall within this class if they rely on fixed payments and may therefore not be individually rational (Ansink and Ruijs, 2008). When payments are flexible rather than fixed – and individually rational agreements allow for this – this barrier is resolved. The combination of a flexible allocation rule and flexible payments was suggested by Kilgour and Dinar (2001) to achieve an “efficient schedule” of river water allocation. Kilgour and Dinar (2001) consider efficient and individually rational agreements and therefore our Proposition 3 can be compared to their result. Specifically, this proposition indicates that their efficiency result may hold, but these “efficient schedules” are not stable for sufficiently low δ. This difference illustrates the importance of considering repeated interaction between agents in a river basin. In addition, the consideration of repeated interaction may cause two additional effects. One is that it may affect the distribution of the benefits of cooperation (see below). The other consequence is that stability may come at the cost of reduced efficiency (see Section 5). Both effects are absent in the (static) model by Kilgour and Dinar (2001). Note that the subset of individually-rational agreements contains various types of agreements. One example is a price-dependent agreement, in which the allocation rule is the efficient allocation and the payment implements the efficient water price, such that marginal benefits of water use are equal to both agents. This type of agreement mimics an international water market (cf. Ansink and Houba, 2012). An alternative example is the (asymmetric) Nash-bargaining solution, which we analyze in Section 7.1. Recall the interpretation of fixed-payment agreements as insurance contracts. Individually-rational agreements can also be interpreted as insurance contracts with a stochastic price sc(e1, e2) that depends upon the realization of river flow. This means that next to the risk over the allocation of water there is also risk with respect to this price. Because monetary payments enter the agents’ quasi-linear utility functions in (1) as the linear term, agents are risk neutral with respect to this category of risk. Consequently, they are indifferent between individually-rational agreements and fixed-payment contracts with the same allocation and payment E{sc(e1, e2)}.","In this section, we focus on agreements that are the outcome of a negotiation procedure and assess under which conditions such agreements can be sustained in equilibrium. Compared with the agreements assessed in Section 6, the current agreements potentially offer more flexibility. Here, our aim is not to assess an exogenously given subset of agreements, but rather to endogenize the agreement given exogenous model parameters (e.g. the discount factor δ and a weight α, introduced below). Nevertheless, to guide our assessment we do need some structure on the types of agreements to assess and, through our selection of equilibrium concepts, we focus on efficient agreements only. We will approach such agreements using the equilibrium concepts of subgame-perfect equilibrium and renegotiation-proof equilibrium. Nash-bargaining agreements are assessed to illustrate how negotiations may lead to efficient and stable agreements. Renegotiation-proof agreements show how this additional stability requirement affects the agreement design. As before, we will not need Assumptions 3 and 4. Nash-bargaining agreements ~~~~~~~~~~~~~~~~~~~~~~~~~~ Note that application of the ANBS implies that we assume, to some extent, an ‘ideal’ setting where bargaining strengths are readily available and there are no restrictions imposed by sustainability considerations or other factors at play (see Fig. 5 and its discussion below). In general, the bargaining strength parameter can be based on factors related to economic, political, or military dominance in the river basin. For specific case studies, an empirical estimate of bargaining strength can be obtained by evaluation of an existing agreement. Assuming that such an agreement is a Nash bargaining solution obtained in a setting not too different from the ‘ideal’ setting, its outcome can be used to derive the corresponding α parameter. Houba et al. (2014) provide a framework for making such a derivation. For Nash-bargaining agreements, given weight α, we show in our next result that any efficient agreement that improves upon the utilities under the minmax allocation xN(e1, e2) = (e1, e2) is an SPE for sufficiently large δ < 1. The difference with Proposition 3 is subtle. By construction, the ANBS as applied in (13) satisfies Conditions (10) and (11) that describe the set of individually-rational agreements and, therefore, Nash-bargaining agreements are a subset of this set. The Pareto efficiency of Nash-bargaining agreements implies that the threshold discount factor for which an agreement can be sustained will coincide with the one for the associated Pareto efficient individually-rational agreement. As is standard in the literature on repeated games, we present our result in terms of a threshold discount factor, above which subsets of agreements can be sustained in equilibrium. For the practical purpose of this paper, however, the real issue is how to construct sustainable agreements when δ is given. In our case this means how to construct the range of sustainable Nash bargaining agreements. This is why our next result consists of two parts: one in terms of a threshold lower bound on δ and one in terms of a threshold upper bound on α. Before discussing these results, we stress that Proposition 4 depends on the specific payment rule adopted in (13) which limits its scope, see Footnote 12. The result by Ambec et al. (2013) is particularly noteworthy, since they find that the downstream incremental distribution is actually the least sustainable agreement (Ambec et al., 2013, Proposition 3). This result is the converse of our result. The reason for this opposite result is the setting of the analysis; Ambec et al. (2013) present a static analysis and thereby ignore the importance of repeated interaction. This repetition turns out to completely reverse the relative advantage of upstream–downstream location in the river basin. In a static setting, the upstream agent can act like the hegemon to exploit his geographic advantage. In a dynamic setting, however, agents depend on each other and the downstream agent has an advantage in timing as explained in Footnote 2. The downstream incremental distribution has been criticized for being an extreme solution. In the next subsection, however, we will see that the requirement of renegotiation-proofness adds even more credibility to this ‘extreme’ solution. Renegotiation-proof equilibria ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In this section we assess the implications of requiring agreements to satisfy renegotiation-proofness. Up till here, we assumed trigger strategies, but these strategies have an important disadvantage: the agent who carries out the punishment by switching to non-cooperative play is also punishing himself. This gives the punisher an incentive to abolish his punishment, and re-negotiate with the defector in order to revert to cooperative play, which Pareto-dominates the non-cooperative path. As a result, punishments by trigger strategies lack credibility. In response to this lack of credibility, the key idea of renegotiation-proof equilibria (RPE) is to construct strategy profiles with punishments that include a non-negative reward to the punisher. Based on the concept of weakly renegotiation-proof equilibrium proposed by Farrell and Maskin (1989), our following result provides additional support for solutions in which the downstream agent gains most. Appendix B contains additional background information on RPE as well as an extensive proof of Proposition 5. We will first summarize the results of Appendix B before stating our main results. Since we are interested in Pareto efficient solutions, we focus, without loss of generality, on sustaining Nash bargaining agreements in RPE. In order to sustain the Nash bargaining agreement α0 ∈ [0, 1), we introduce two punishment paths, one for each agent. The punishment path for agent 1 corresponds to the Nash bargaining agreement α1 = 0 in every period, a path on which agent 1 has no incentive to deviate. The punishment path for agent 2 is somewhat more involved. In the first period of this path, agent 1 does not deliver any water and agent 2 pays a penalty that is equal to the present value of the entire expected net surplus of all future periods from the second period onwards, which is independent of the current realization. From the second period onward, we switch to the Nash bargaining agreement α2 = 0 in every period, which equals agent 1's punishment path (we postpone an explanation). For the moment, agent 2's punishment path coincides with agent 1's best RPE path. Therefore, agent 2's punishment path is Pareto efficient from the second period onwards with unavoidable efficiency losses only in the first period. In our next results, we state the theoretically largest set of Pareto efficient RPE Nash bargaining agreements that are derived in Appendix B, using the thresholds in (14) and (15). Before discussing this result, we stress that Proposition 5 is linked to Proposition 4 and thereby depends on the specific payment rule adopted in (13) which limits its scope, see Footnote 12. The positive news is that weakly renegotiation-proof equilibria exist and that the underlying strategies can replace trigger strategies. Surprisingly, imposing a more restrictive equilibrium concept does not result in a higher threshold level for the discount factor or a reduced upper bound on agents 1's bargaining weight. In Appendix B we show that agent 2 obtains exactly his minmax payoff on his punishment path. Since this is the same payoff as under the trigger strategies, both strategies employed as punishment strategies are payoff-equivalent to agent 2. As before, only agent 2 has an incentive to deviate. Given the payoff-equivalent punishments to agent 2, this must result in the same thresholds as derived under trigger strategies. We close this section with several remarks on the implications of the penalty in player 2's punishment path. Note that agent 2's worst RPE payoff consists of his expected payoff E{b2(e2)} minus a penalty in the first period, which is below his minmax value, followed by his expected maximum RPE payoff in all future periods, which is above his minmax value. Although we objected against payoffs below the minmax value in Remark 2, this argument does not apply here because, by paying the penalty, agent 2 invests in restoring the cooperation, which might be interpreted as either to repent and show remorse, or to regain agent 1's trust. Furthermore, the penalty equals the present value of the entire expected net surplus of all future periods from the second period onwards, which can be quite substantial. Even though agent 1 receives his minmax payoff in all future periods from the second period onwards, this agent receives almost the entire present value of the overall net surplus as the penalty in the first period (he only misses out on the net surplus of the first period, in which he only receives E{b1(e1)}). Theoretically, such a substantial penalty is fine, but in practice it might not be realistic. For practical purposes, one might resort to weights α1 < α0 < α2 that are closer together and construct paths similar as described above with one exception: the punishment path for agent 2 specifies (e1, e2) and some small penalty in the first period, followed by the Nash bargaining agreement α2 for several periods before it continues forever with Nash bargaining agreement α0 (instead of α1).14 Then, as before, agent 1 has no incentive to deviate and agent 2 should be given enough incentives to undergo any of these three paths. We leave this option for future research. The main message of this section is that Pareto efficient weakly renegotiation-proof equilibria can be derived and that we characterized its theoretically largest set, which happens to coincide with the set of sustainable Nash bargaining agreements in SPE.","This paper is the first to systematically assess the implications of repeated interaction for the sustainability of river sharing agreements between riparian neighbors. We obtain an interesting combination of results. First, our Folk Theorem for river sharing problems in Proposition 3, further refined in Propositions 4 and 5, provides clear conditions for sustainable agreements in terms of the distribution of the gains from cooperation. Remarkably, the complete set of sustainable Nash bargaining agreements under subgame- perfect equilibrium can also be sustained by the more restrictive weakly renegotiation- proof equilibrium. Second, in Proposition 2 we establish a trade-off between the efficiency and stability of agreements, which illustrates the importance of considering repeated interaction. Repeated interaction tends to favor the downstream agent, which may seem counterintuitive at first, but may explain empirical observations on downstream states managing to negotiate a substantial share of upstream river water. Our results provide non-cooperative support for solutions that assign larger shares of the pie to downstream agents. At the lowest possible threshold on the discount factor, only the downstream incremental distribution, proposed by Ambec and Sprumont (2002), that assigns all gains from cooperation to downstream agents can be sustained and this distribution remains sustainable for higher discount factors. Finally, the model developed in this paper offers ample scope for extensions and applications. One obvious extension is to allow for more general river geographies as in Khmelnitskaya (2010), Ansink and Houba (2012), or van den Brink et al. (2012). One obvious application is to repeat the analysis of the Bishkek Treaty in the Aral Sea basin by Ambec et al. (2013) in the dynamic setting of this paper. Their static analysis showed that actual payments under this agreement approximate the payment rule induced by the downstream incremental distribution. In a static setting, this result implies instability. When considering repeated interaction, however, the results in our paper suggest that this payment rule may actually be well- chosen."],["This paper investigates the extent to which cross-country differences in aggregate participation rates can be explained by differences in tax-benefit systems. We take the example of two countries, the Czech Republic and Hungary, which – despite a lot of similarities – differ markedly in labour force participation rates. Using comparable individual-level labour supply estimates, we simulate how the aggregate participation rate would change in one country if the other country's tax and social welfare system were adopted. The estimation results for the two countries are quite similar, suggesting that individual preferences are essentially identical in the two countries. The simulation results show that about one-third of the difference in the participation rates of the 15–74 year-old population and more than two-thirds of the participation of the prime-age population can be explained by differences in the tax-benefit systems. --------------------------------------------------------------------------------","A cross-country comparison of labour force participation rates is the most straightforward way to identify labour supply problems in a particular country or set of countries. International organisations, central banks and economic think-tanks often motivate their policy recommendations by comparing a country’s participation rate – usually broken down by selected sub-populations – against an international standard or against other similar countries’ statistics. The resulting recommendations frequently involve, but are not limited to, reforms of the country’s tax and welfare system to restore or improve work incentives.1 That is, the financial incentive to work is implicitly the primary candidate for explaining cross-country differences in the participation rates of apparently homogeneous groups of individuals. Is this really the most important factor in explaining cross-country differences, or are other aspects such as differences in preferences or other observed and unobserved characteristics (e.g. average health status) even more important? To what extent can differences in financial incentives explain the observed differences in labour force participation between countries? Although the impact of labour taxation and the welfare benefit system on labour force participation has been widely investigated in the empirical literature, the extent to which cross-country differences in participation rates can be explained by the diversity of tax-benefit systems has never been addressed directly. To fill the gap, we take the example of two countries, the Czech Republic and Hungary. We use, for both countries, a very detailed, complex microsimulation model and comparable micro estimates of labour supply at the extensive margin (i.e. the participation decision) to quantify the part of the difference in the two countries’ participation rates accounted for by different tax and welfare benefit systems. More precisely, we first estimate the same labour supply models for Czech and Hungarian individual-level data. Two models are tested for both countries: a structural, random utility labour supply model originally proposed by Van Soest (1995) and a reduced-form model relating the participation probability to the financial gains from accepting a job.2 The entirely comparable estimated equations for the Czech Republic and Hungary are then used to simulate how labour supply would change in one country if it adopted the other country’s tax and social welfare system. Assuming that the economies are supply-driven in the long run, the initial labour supply shock fully translates into participation and the results of the simulations can be interpreted as the effect of differences in the tax- benefit systems on the differences in the two countries’ participation rates. Overall, our empirical strategy follows a standard procedure in assessing the impact of tax and transfer system reforms on labour supply at the extensive margin. However, instead of assessing the effects of a specific past, planned or ad hoc hypothetical reform, we analyse the labour supply effects of adopting another country’s tax-benefit system. In doing so, this paper offers the first tentative explanation of cross-country differences in labour force participation rates by differences in taxation and welfare benefits. The closest strand of literature uses cross-country individual-level or disaggregated data either to investigate the labour supply effects of a specific policy change or to explain developments in participation rates from a cross-country comparative perspective. Some recent empirical studies analyse the differences in micro- and macro-level elasticities using cross-country individual-level data (Jäntti et al., 2015); others study the effect of hypothetical welfare reforms (e.g. Immervoll et al., 2007) or tax reforms (see, for example, Aaberge et al., 2000, for a cross-country analysis of flat-tax reforms); cross- country microsimulation models are also used to estimate governments’ redistributive preferences (Blundell et al., 2009; Bargain et al., 2011). Some other papers use individual-level or disaggregated data to describe the evolution of labour supply in some selected countries, either for the whole population (Balleer et al., 2014, or Blundell et al., 2013) or for a selected sub-population (e.g. Cipollone et al., 2014, for women). However, none of these or any other existing empirical papers have so far provided direct evidence to explain differences in participation rates between countries. The Czech Republic and Hungary provide an interesting case for comparison. The two countries exhibit a lot of similarities: their economies are geographically close, similar in size and level of economic development, and partly share a common history. In particular, both economies experienced full employment until the regime change at the end of the 1980s. Despite all their common factors, their participation rates differ markedly. In 2008 – the reference year we use for our simulation – with only 54.3% of the 15–74 year-old population working or actively seeking employment, Hungary recorded the second lowest participation rate among EU member states (behind Malta), while the Czech participation rate was close to the EU15 average (63.1% in the Czech Republic and 64.1% in the EU15, see Section 2.1). The low Hungarian participation rate is often considered to be the result of a combination of high taxes on labour and a relatively “generous” transfer system. After describing some stylised facts about Czech and Hungarian labour force participation and the key differences between the two countries’ tax and social welfare systems (Section 2), we present our methodology in Section 3. Our estimation results (presented in Section 4) show that the labour supply elasticities obtained from the structural and reduced-form models are quite similar in both countries. Moreover, the estimated labour supply elasticities for the Czech Republic are very close to the results for Hungary. Overall, the results suggest that, at least in this dimension, individual preferences are very close in the two countries. This even holds true for sub-populations depending on age group, level of education, gender and marital status. In both countries, lower educated people and married women are the most responsive to tax and transfer changes. Given the large dispersion of the labour supply elasticities found in the empirical literature for different countries and time periods, the similarities of the estimated elasticities for the two countries we analyse using two distinct models are in line with the evidence in Bargain et al. (2014), suggesting that a considerable part of the cross-country variation in the elasticities found in the literature is driven by methodological differences. The simulation results suggest that about one-third of the difference in the participation rates of the 15–74 year-old population and more than two-thirds of the participation of the prime-age population can be explained by differences in the tax-benefit systems. The results are quasi-symmetric, meaning that if the Czech system was adopted, the Hungarian participation rate would increase by about the same number of percentage points as the Czech participation would decrease if the Hungarian system were implemented. The highest responses are obtained for individuals with a low level of education and for married women. These differences are related to the higher personal income taxes for low earners and the more generous welfare system in general and maternity allowances in particular that are in place in Hungary as compared to the Czech Republic. Labour force participation in the Czech Republic and Hungary ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The Czech Republic and Hungary share many similarities. Together with Slovakia and Poland, the four countries of the Visegrád group3 have long been considered frontrunners of the fundamental systemic transformation from central planning to a market economy following the collapse of communism in 1989. After an initial recession, widespread changes to their economic systems shifted these countries to a relatively fast growing path. However, there are some marked differences in the transition paths they chose to pursue. These particular differences, together with all the common factors they share (such as their largely similar history, geographical closeness and the similarities in their initial conditions before the regime change), make the Visegrád countries a natural subject for comparative research. Comparing the two countries under investigation in this study, the Czech economy outperforms the Hungarian economy by a significant margin. By 2008 (the reference year for our comparison; more on this later), the Czech Republic had reached 76% of the EU15 economic average in terms of GDP per capita (measured in purchasing power standards; see Table A.1 in Appendix A). The Hungarian economy lags behind with only 56% of the EU15 average. However, the largest part of the difference in relative economic development between the two countries can be explained by the lower employment rate in Hungary: the difference between labour productivities, measured as GDP per employed person, is 71% of the EU15 average in the Czech Republic and 64% in Hungary. The total hourly compensation of employees is also similar in the two countries, reaching 54% of the EU15 average in the Czech Republic and 46% in Hungary. The two countries are also close in terms of openness and economic structure. With export shares of 63% in the Czech Republic and 79% in Hungary, both economies are remarkably open. The largest share of exports goes to EU countries, in particular to Germany. The Czech and the Hungarian economies are both largely industry-oriented, also in comparison to EU15 standards. On the other hand, the share of public administration in GDP is rather close to the EU15 minimum in both countries. Despite all the similarities, the participation rates differ markedly in the two countries. While Hungary lies near the low end among EU countries, the Czech Republic is around the middle of the ranking (see Fig. 1). This observation holds for both the working-age and prime-age populations. The significant differences in labour market performance are typically explained by the different policies adopted during the first few years of the transition process and kept – at least partly – unchanged since then. Both countries implemented fast-track reforms in the early 1990s, with two main differences. First, Hungary opted for case-by-case privatisation as opposed to the voucher privatisation method adopted by the former Czechoslovakia.4 Second, the Hungarian government introduced strict bankruptcy regulations in 1992. The differences in privatisation methods and the draconian bankruptcy regulations implemented in Hungary provoked a much larger drop in employment in Hungary as compared to Czechoslovakia. Between 1990 and 1993, total employment declined by 9% in Czechoslovakia and by 22% in Hungary. Moreover, in Hungary, a number of policy measures – such as alleviated conditions for entering the old-age or disability pension systems – contributed to pushing people out of the labour force rather than into unemployment. As a result, the participation rate declined continuously until 1997, when it was only 50.6% of the 15–74 year-old population, about 10 percentage points lower than the EU15 average and more than 13 percentage points lower than in the Czech Republic.5 Table 1 provides additional information by breaking down the participation rates by selected sub-populations. The participation rates (column 1) and population shares (columns 2–4) are expressed as a percentage of the corresponding sub-population. The numbers presented in the table are based on the database we use in this paper, namely the Survey of Income and Living Conditions (SILC) data for the Czech Republic and the Household Budget Survey (HBS) database for Hungary (see more on the datasets in Section 4). The total participation rates for the full sample of 15–74 year- old individuals (row 1, column 1) are somewhat closer in the two countries in our datasets than in the official data based on the Labour Force Survey (in Fig. 1). This is due to the fact that our datasets are representative at the household level, while the Labour Force Surveys are representative at the individual level. The difference of 7 percentage points is still considerable. The largest difference in participation rates in favour of the Czech Republic is observed among the youth population (15–24 year-olds, second row). While the large share of inactive youths can be explained by school attendance, the low share of youth participation in Hungary is still alarming, all the more so since the share of youths attending full-time education is also lower in Hungary than in the Czech Republic (column 2). Among the prime-age group, participation rates are significantly higher in the Czech Republic for lower-educated individuals, especially for those with elementary school as the highest level of education, while participation rates are close to each other for persons with tertiary education. Since lower educated individuals are usually found to be more responsive to tax and transfer changes (more on this in Section 4), the participation rates by educational attainment in the two countries are consistent with the higher taxes and more generous transfer system in Hungary as compared to the Czech Republic (see the next sub-section). Finally, the elderly population (aged 56 years or older, last row of Table 1) is considerably less active in the labour market in Hungary than in the Czech Republic. This divergence is likely to be related to the higher share of old-age and disability pensioners in Hungary within this population group. While the Hungarian old-age pension reform in 1997 and the successive strengthening policies adopted in order to reduce abuse of the disability pension scheme have gradually increased the effective retirement age and reduced the number of new disability claims granted, the effect of the policy responses to the transition shock is still visible among the older generation. Key elements of the Czech and Hungarian tax and transfer systems ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The Czech and Hungarian tax-benefit systems are both characterised by individual personal income taxation and a broad range of welfare benefits. In what follows, we briefly describe and compare the two systems effective in 2008, the reference year we used for our simulation.6 In this section, we only look at the main characteristics of the systems that apply to regular wage earners; the self-employed and other specific cases are not considered. In our empirical work, more details are taken into consideration. Additional details on the transfer systems of the two countries are presented in the Appendix. For a more comprehensive description of the two systems, see, for example, the country chapters of the OECD’s Benefits and Wages publication.7 In both countries, labour incomes are subject to personal income tax and employer and employee contributions. In Hungary, the tax schedule comprises three brackets, with tax rates of 18%, 36% and 40%.8 Employees with annual earnings under a given threshold are eligible for an earning income tax credit (EITC) of 18% of their wage income, which is phased out at a rate of 9%. A child tax credit (fixed amount) is available for families with three or more children. Social security contributions related to sickness, unemployment and pensions are applied. In the Czech Republic, family taxation for married couples with children was introduced in 2005. The joint taxation, however, became irrelevant when a flat tax rate of 15% was introduced in 2008. It is applied to the so-called super-gross income, which includes employer social security contributions. Tax credits are available per person, per spouse under a given income and per dependent child. If the tax credit per child is negative, it is provided to the household as a tax bonus. Comparing the two personal income tax (PIT) systems, Fig. 2 reveals that, first, the Hungarian system is more progressive and, second, average taxes are higher. Moreover, the Czech average tax rate can be negative for families with children if the calculated income tax is lower than the amount of family benefit the taxpayer is eligible for. The higher average PIT rates in Hungary are also reflected in higher tax revenues: in 2008 the Hungarian government collected 7.5% of GDP as PIT tax revenue, compared to 3.5% in the Czech Republic (see Table 2). Both countries provide a wide range of social benefits in order to – at least partly – compensate for temporary loss of labour income, to reduce and prevent poverty or to achieve other goals. Overall, if we look at entitlement periods and the share of GDP spent on various transfers in 2008, the Hungarian benefit system seems more generous than the Czech one (Table 2). In 2008, Hungary spent almost twice as much (expressed as a percentage of GDP) on unemployment benefits and about 60% more on old-age and disability pensions compared to the Czech Republic. The difference is even greater in the case of maternity allowances, especially in terms of entitlement period. Although the net amount of the maternity benefit is usually higher in the Czech Republic during the first 6 months after childbirth, the entitlement period is much longer in Hungary: conditional on past employment, Hungarian mothers can receive up to 70% of their previous income until the second birthday of the youngest child and parents are eligible for a fixed amount until the third birthday of the youngest child. Even longer entitlement periods apply for the third child, twins and disabled children. As a general rule, in Hungary, maternity benefits and unemployment benefits are either fixed amounts or depend on previous income and are subject to income taxation. On the other hand, other social benefits and pensions are exempt from taxation. In the Czech Republic, all benefits are non-taxable (except for very high pensions). In the case of maternity benefits, the amount is calculated based on previous gross income, whereas unemployment benefits are calculated based on previous net income. These differences make a direct comparison of the two systems rather challenging. Some details of the two welfare benefit systems are listed in Table 2. For a more detailed description, see the Appendix or the references presented in the first paragraph of this section.","We first estimate, separately for Czech and Hungarian individual-level data, a structural and a reduced-form labour supply model where both taxes and transfers are treated in a unified framework. In what follows, we refer to this step as the “estimation part”. In the second step, the “simulation part”, we use these estimated equations to simulate how each individual’s probability of being active – and hence the aggregate labour supply – would change in one country if it adopted the other country’s tax-benefit system. More precisely, we perform a microsimulation exercise: we replace the net income variables (wages, transfers and non-labour incomes) in the estimated equations by the values that would result from adopting the other country’s tax and transfer system. As a consequence of the benefits gained or losses suffered due to changes in the effective tax rates and the amount of (potential) transfers, individuals are expected to adjust their labour supply according to the estimated labour supply equations. The weighted average of the individual changes in the participation probabilities corresponds to the aggregate labour supply shock induced by the hypothetical reform. Under certain assumptions detailed below (see Section 3.4), the result of this exercise reveals the extent to which the difference between the two countries’ participation rates can be explained by the differences in the taxation and welfare benefit systems. In this framework, the part of the cross-country differences in participation rates that can be attributed to the differences in tax and welfare systems is assessed using counterfactual simulation scenarios. That is, the effects of the cross-country differences in tax-benefit systems are only indirectly identified. Ideally, to address this question directly, one would need to observe a country that adopts the entire tax-benefit system of another country. Such a quasi- experiment would allow us to use a more transparent identification strategy and to directly estimate the impact of such a large-scale reform.9 In the absence of such an experiment, the approach adopted in this paper provides a viable alternative strategy. Modelling labour supply at the extensive margin ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The labour-supply decision is modelled as a discrete choice problem. Originally proposed by Van Soest (1995), this approach has become increasingly popular and quite standard in recent years. In this framework, utility-maximising individuals are supposed to choose between a few alternative discrete sets of hours of work, such as inactivity (zero hours worked), part-time or full-time. Reducing the maximisation problem to choosing among a discrete set of possibilities yielding different utilities considerably simplifies the problem and provides a simple yet rather general way of representing labour supply decisions in the presence of nonlinear and non-convex budget constraints. The model is in principle very flexible. The only restriction on the model is the imposition of increasing monotonicity in consumption. In our empirical work, we model individual labour supply decisions. Married couples or individuals living with a partner take into account the labour market status of their husbands, wives or partners, but household members do not optimise collectively. Furthermore, we assume that there are only two labour-market states, active and inactive. Indeed, in both the Czech Republic and Hungary, part-time work is relatively rare: before the outbreak of the crisis, the share of part-time employees among all workers was less than 5% in both countries. By 2014, this share had increased to 6.4% in the two countries. For comparison, the average share of part-time employees in the EU28 reached 18.2% in 2008 and increased further to 20.4% in 2014.10 Limitations and interpretation of the simulation results ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Whether the initial labour supply shock induced by the adoption of the other country’s tax-benefit system fully translates into increased participation (and employment) in the long run depends on the general equilibrium interactions of the entire economy. Our approach implicitly assumes a textbook (fully supply-driven) neoclassical general equilibrium long-run model of a small open economy with perfectly elastic capital supply. Our results are thus equivalent to the output from the general-equilibrium behavioural microsimulation model of Benczúr et al. (2018) when perfectly flexible capital supply is assumed. The perfect flexibility of capital supply is equivalent to assuming the non- existence of demand-side constraints on the labour market.21 In this framework, the expected rate of return on capital is pinned down by international benchmarks, which implies that the capital-to-labour ratio and the ratio of factor prices stay constant in the long run. It follows that, for example, after a positive labour supply shock induced by a tax cut, capital accumulation will follow until the new equilibrium is reached with an unchanged capital-labour ratio. After the adjustment period, gross wages will also reach their pre-reform level, so factor prices remain unchanged in the long run. The initial labour supply shock therefore fully translates into participation in the long run. Consequently, under these assumptions, the results of our simulation exercises can be interpreted as the effect of differences in the tax-benefit systems in the two countries on the differences in the two countries’ participation rates.22 While these assumptions are quite standard in the literature, several factors can realistically impede the adjustment process. To start with, possible demand-side constraints can prevent the full adjustment of gross wages, which would limit the labour supply response (if the original shock comes from the demand side, for example following a decrease in employer contributions) or increase unemployment rather than employment (for example following a decrease in PIT or employee contributions, or following a strengthening of the benefit system). In the latter case, labour force participation is still impacted. However, the additional mass of unemployed people might give up searching after some time and withdraw from the labour market because of failed searches (discouraged worker effect). Second, search-and-matching frictions on the labour market might also impact the results. For example, although a shortened eligibility for unemployment benefits increases labour supply and strengthens job-search incentives, it might also affect the successful matching rate negatively. The overall effect of the shift in the aggregate successful matching rate is similar to the above-discussed demand-side constraints. Nevertheless, Horvath et al. (2018) adopt a similar modelling strategy to assess the impact of tax-benefit reforms in Slovakia and extend it with a search-and-matching block. Their results suggest that search-and-matching frictions have a limited impact in the long run. On the other hand, the impact of the tax and transfer reforms may also be higher than predicted by our simulations. We simulate the effects of adopting the other country’s tax and transfer system in a static framework. That is, we assume that the tax and transfer system is not correlated with other observable or unobservable characteristics. In the long run, people might change all their life decisions in response to a drastic change in the tax and welfare system. For example, reducing the rate of the labour tax or transfers might make individuals more inclined to pursue higher education. As a consequence, the composition of the population might also change following a tax and transfer reform, which might in turn affect the aggregate participation rate. Finally, the “unexplained” part of the difference between the two countries’ participation rates is possibly also affected by some elements of the transfer system that we do not explicitly model, most importantly the differences in the old-age and disability pension schemes. Although both types of pensions are included in the model, neither individuals’ retirement decision, nor their access to the disability pension scheme is taken into account in our microsimulation exercise. That is, we assume that individuals’ retirement decisions, their access to the disability pension scheme and the amount of these pensions remain unchanged when the other country’s welfare benefit system is adopted.","We restrict our analysis to the working-age population of individuals aged 15–74 years, but use information on the full population to generate the relevant variables, compute net incomes and generate (hypothetical) transfers.23 The labour supply estimation part is carried out using Survey of Income and Living Conditions (SILC) data for the Czech Republic for the years 2005–2010. For the simulation part, explained in detail in Section 3.3, we restrict the sample to the 2008 wave. The SILC is a yearly cross-section survey of households that provides detailed information on demography and incomes. For each year, the survey contains broadly 10,000 households and 23,000 individuals. All income variables (labour and non-labour incomes, transfers) are expressed in total annual terms. To compute net incomes and potential transfers that individuals are entitled to or would be entitled to in the absence of any income from work, we used a modified version of the tax-benefit model developed by Galuščák and Pavel (2012). We used the same codes to compute net labour incomes, but extended the original version of the model in order to take into account more complex features of the Czech tax and transfer system (e.g. taxes on capital revenues and maternity benefit). As for the transfers, we simulated the total annual amount of benefits the individual gets if he is not working or could get if he decided to stop working.24 To compute net incomes and hypothetical transfers for Hungary, we use the microsimulation model developed by Benczúr et al. (2018).25 The model features a very detailed static fiscal microsimulation model, combined with behavioural labour supply responses at both the intensive and the extensive margin and a small neoclassical macro model.26 The fiscal microsimulation model is parameterised for each year since 1998. The estimation part is carried out using all waves between 1998 and 2008, while the simulation part (see Section 3.3) uses only the 2008 wave. The Hungarian model runs on the Household Budget Survey (HBS) database, which is more detailed but otherwise comparable to the CZ-SILC database. The difference with respect to the CZ-SILC is twofold: first, several income variables in the Czech database are broken down into several categories in its Hungarian counterpart; second, some variables – such as income from abroad, income of minors, life annuity payments, scholar grants and severance indemnities – are reported in the Hungarian database but no information is available from the Czech database. Finally, we carefully linked all the variables to their counterparts in the other database. If one-to-one correspondence could not be achieved, the closest possible several-to-one match was established (see Tables C.1 and C.2 in the Appendix). Estimation results ~~~~~~~~~~~~~~~~~~ Table 3 reports the estimated conditional marginal effects evaluated at the population averages and by selected subgroups following the methodology presented in Sections 3.1 and 3.2. As expected, a higher net wage increases the probability of participation in the labour market, while means-tested transfers have the opposite effect. Overall, the structural models provide very similar conditional marginal effects to the reduced-form models and the Czech results are remarkably close to the Hungarian ones. For the whole population, the lowest wage elasticity is obtained for the Czech Republic with the structural model and the highest elasticity is obtained with the structural model for Hungary. On average, a 10% increase in the net wage leads to a 2.4–2.8 percentage point increase in the probability of being active in the Czech Republic and a 2.8–3.4 percentage point increase in Hungary. The estimated marginal effects of the transfers are even closer to each other across countries and model specifications. The effect of a 10% increase in transfers on the average individual’s probability of being active ranges from -0.9 percentage points in Hungary according to the reduced-form model to -1.2 percentage points in the Czech Republic with the structural model. The similarity of the results suggests that, on average, individual preferences are essentially identical in the two countries. This is also true for selected subgroups depending on age group, level of education, gender and marital status. Again, the structural models basically imply the same net wage and transfer elasticities as the reduced-form models and the Czech results are very close to the Hungarian ones. The conditional marginal effects calculated on the prime-age population averages (25–54 year-olds) are lower (in absolute value) than the results for the whole population, indicating that prime-age individuals are less responsive to tax and transfer changes than the rest of the population. Within the prime-age population, however, the education level strongly influences individuals’ responsiveness: the lower educated (“elementary school or less”), probably highly transfer-dependent group, is highly responsive, while the estimated elasticities for secondary and tertiary educated people are much smaller. The remaining rows in Table 3 suggest that women’s participation is more responsive than that of men, particularly for married women. The estimation results and the implied wage elasticities are qualitatively in line with existing results in the empirical literature. However, a direct comparison of the elasticities is not straightforward. There is a large consensus in the literature that women (and especially married women) respond more to changes in their net market wage than men. Lower educated individuals are also usually found to be more responsive than the average. The variation in the magnitude of the estimated labour supply elasticities found in the literature is nonetheless considerable and the reported elasticities are rarely directly comparable for several reasons. First, with some exceptions, early labour supply models using the continuous Hausman approach usually do not distinguish between the decision to participate (the extensive margin) and the decision regarding hours worked (the intensive margin) and only report uncompensated and/or compensated (Hicksian) hours elasticities, not the participation elasticity.27 Second, although the discrete choice approach by construction always includes participation elasticity, estimates from various studies are not easily comparable due to differences in specifications and definitions. A comprehensive comparison of the international evidence on labour supply elasticities by Bargain and Peichl (2013) shows that the large variance in the reported elasticities is – at least partly – explained by differences in specifications (most importantly, whether wages are predicted for all individuals or only for non-workers) and the time period considered. A meta-analysis by Chetty et al. (2013) provides a “consensus” elasticity: according to the authors, existing micro and macro evidence points towards an aggregate steady-state elasticity of labour supply at the extensive margin of 0.25. Overall, the marginal effects obtained for the Czech Republic and Hungary lie in the (rather large) range obtained in previous studies, especially studies using the discrete choice approach.28 Given the large dispersion of the labour supply elasticities estimated for different countries and time periods and using various methodologies, the similarities of the results we obtain for two different countries is in itself interesting and we tentatively suggest that methodological differences might be responsible for a considerable part of the cross- country variation in the elasticities found in the literature. The findings of Bargain et al. (2014), who estimate a comparable labour supply equation for a large set of countries and show that cross-country differences in genuine work preferences are rather small, corroborate this view. Simulation results ~~~~~~~~~~~~~~~~~~ The simulations were carried out as explained in Section 3.3. The results are shown in Table 4. The reported figures indicate the participation rates for 2008 (the reference year chosen for our simulations) for both countries on the overall sample of individuals aged 15–74 (row (a)) and for the prime-age subsample (rows (b) to (i)). The participation rates are considerably higher in the Czech Republic than in Hungary for almost all subgroups (see columns (1) and (6)). The difference in participation rates for the whole population is 7.0 percentage points according to our datasets. The simulation results using each country’s own tax and benefit systems are presented in columns (2), (3), (7) and (8). The reported numbers are weighted averages of the participation probabilities predicted by the estimated structural equations (columns (2) and (7)) and the reduced-form equations (columns (3) and (8)). As mentioned in Section 3.3, the constant terms in both the structural and the reduced-form equations are set so as to match the aggregate statistics. The simulated full-sample aggregate participations (cells (a2), (a3), (a7) and (a8)) are therefore equal to the observed participation rates (cells (a1) and (a6)). The differences in the participation rates for the selected subgroups come from differences in preferences (captured by the labour supply controls and taste shifters in the logit equations), differences in productivity and thus in gross wages, differences in average tax rates, which influence net wages, and/or differences in potential transfers. For both countries and both models, the simulated participation rates for the selected subgroups are close to the observed ones. This suggests that, overall, the estimated equations capture individuals’ heterogeneity quite well. In columns (4), (5), (9) and (10) we show how Czech and Hungarian labour force participation would change (in percentage points) if one country adopted the other country’s tax-benefit system. The results of our simulation exercise show that, overall, the Czech participation rate would decrease by 2.2 or 3.0 percentage points (depending on the model used) if the country adopted the Hungarian system. The Hungarian participation rate would increase by 2.4–2.5 percentage points if the Czech system were adopted. These results are close to symmetric and indicate that the differences in the two tax-benefit systems explain about one-third of the total difference in the two countries’ participation rates. On average, the difference is even greater for prime-age individuals, reaching (-2.6)–(-2.8) percentage points for the Czech and 3.0–3.3 percentage points for the Hungarian data (see row (b)). That is, according to our results, more than two-thirds of the total difference in participation rates for the prime-age population between the two countries can be explained by the differences in the tax- benefit systems. The simulation results remain broadly symmetric for most of the selected subgroups.29 The highest response is obtained for low-educated individuals and for married women. The large response of these subgroups is most probably related to the higher personal income taxes for low earners and the very generous maternity benefit system in place in Hungary as compared to the Czech Republic (see Section 2.2). Overall, our results suggest that the major part of the differences in the participation rates for prime-age individuals is related to the tax and transfer systems. Assuming that the hypothetical tax and transfer reform of adopting the Czech system does not alter the composition of the Hungarian workforce and no demand-side constraints prevent adjustment (see Section 3.4 for a discussion of these possibilities), some additional measures would need to be implemented in Hungary for it to reach the Czech participation rate for 25–54 year-olds. This could be achieved, for example, by introducing a lower 5% PIT bracket up to the minimum wage or by shortening the eligibility for unemployment benefits by one month (without touching the longer entitlement period for the elderly), but other measures could also be considered. On the other hand, the gap for the whole working-age population is too large for Hungary to realistically reach the Czech participation rate level by re- parametrising the part of the tax-benefit system that we model explicitly. The “unexplained” part of the difference between the two countries’ participation rates is most likely largely affected by some other elements of the transfer system, most importantly the differences in the oldage and disability pension schemes (see Section 3.4 for a discussion of how these two transfer types are taken into account in our simulations). Although several restrictive measures have been adopted in Hungary to reduce abuse of the disability pension scheme and increase the effective retirement age, the shares of early-retired and disabled individuals are still relatively high in Hungary by international comparison (see Table 1).","Taking the example of two countries, the Czech Republic and Hungary, this paper investigates the extent to which cross-country differences in aggregate participation rates can be explained by the disparities in their tax-benefit systems. We first estimate the same structural and reduced-form labour supply models for Czech and Hungarian individual-level data and use comparable estimates of labour supply at the extensive margin to simulate how the aggregate participation rate would change in one country if it adopted the other country’s tax and social welfare system. Our estimation results yield similar labour supply elasticities for both the structural model and the reduced-form model and for both countries, suggesting that individual preferences are essentially identical in the two countries analysed. Consistently with previous findings, lower educated individuals and married women are the most responsive to tax and transfer changes. The simulation results show that about 2.2–3.0 percentage points out of the total difference of 7.0 percentage points – i.e. about one-third of the total difference – in the participation rates of the 15–74 year-old population can be explained by differences in the tax-benefit systems. The explained part is higher for prime-age individuals: for this group, more than two-thirds of the total difference in participation rates is related to differences in the tax-benefit systems. The simulated effects are quasi-symmetric for almost all subgroups, meaning that if the Czech system was adopted, the Hungarian participation rate of a specific subgroup would increase by about the same number of percentage points as the Czech participation rate would decrease if the Hungarian system were implemented. The largest difference explained by the differences in the tax-benefit systems is identified for low-educated individuals and for married women. These differences are related to the higher personal income taxes for low earners and the more generous welfare system in general and maternity allowances in particular in place in Hungary as compared to the Czech Republic. The “unexplained” part of the difference is also likely to be affected by some elements of the transfer system that we do not explicitly control for, most importantly the differences in the old-age and disability pension schemes. Moreover, in the long run, the tax and transfer system might also influence individuals’ schooling or other life decisions, which might affect the aggregate participation rate through the changing composition of the population. Obviously, these results cannot be directly generalised to other countries. First, it is possible that individual preferences differ more across countries with different cultural and historical backgrounds or institutional structures even within narrowly defined sub-populations. Second, cross-country differences in the composition of the working-age population may lead to different results even following a similar shock to net income levels. Nevertheless, the exercise presented in this paper sheds some light on how important the tax-benefit system can be in explaining the differences in labour force participation between two otherwise similar countries. Finally, it is important to note that we do not suggest that Hungary should adopt the Czech tax and transfer system in order to increase labour force participation. First, even though the Czech participation rate is close to the EU15 average (the usual benchmark for emerging EU member states), the Czech system might not be optimal in every respect. Second, governments’ redistributive preferences may be different and therefore the optimal policy may also differ. Third, given the country’s fiscal constraints, it may not be optimal for Hungary to implement such a large-scale tax- benefit reform. Relative to the 2008 situation, the fiscal impact of implementing the Czech system would be considerable: assuming a constant average consumption share (84%), the reform would impact the budget by 4% of GDP. This is almost double the fiscal impact of the tax and transfer reforms implemented in Hungary between 2008 and 2010, which can already be considered very large (see Benczúr et al., 2018). Moreover, the effectiveness of the reform largely depends on the additional measures chosen to cover the budget deficit. Despite these concerns, the Czech system is still a good benchmark for Hungary, as it provides a realistically achievable alternative for Hungarian policymakers.","This study was funded by the European Commission, Joint Research Centre; and the Czech National Bank Research Project No. D1/12."],["In this paper we consider the problem of testing for the co-integration rank of a vector autoregressive process in the case where a trend break may potentially be present in the data. It is known that un-modelled trend breaks can result in tests which are incorrectly sized under the null hypothesis and inconsistent under the alternative hypothesis. Extant procedures in this literature have attempted to solve this inference problem but require the practitioner to either assume that the trend break date is known or to assume that any trend break cannot occur under the co-integration rank null hypothesis being tested. These procedures also assume the autoregressive lag length is known to the practitioner. All of these assumptions would seem unreasonable in practice. Moreover in each of these strands of the literature there is also a presumption in calculating the tests that a trend break is known to have happened. This can lead to a substantial loss in finite sample power in the case where a trend break does not in fact occur. Using information criteria based methods to select both the autoregressive lag order and to choose between the trend break and no trend break models, using a consistent estimate of the break fraction in the context of the former, we develop a number of procedures which deliver asymptotically correctly sized and consistent tests of the co-integration rank regardless of whether a trend break is present in the data or not. By selecting the no break model when no trend break is present, these procedures also avoid the potentially large power losses associated with the extant procedures in such cases. --------------------------------------------------------------------------------","Macroeconomic series are typically characterised by piecewise linear (or broken) trend functions; see, inter alia, Stock and Watson (1996, 1999, 2005) and Perron and Zhu (2005). Such breaks in the trend function might occur following a period of major economic upheaval or a political regime change. In the univariate setting this has spurred a large literature on testing for an autoregressive unit root when a trend break may be present in the data. The first proper theoretical treatment of this problem was given by Perron (1989) who showed that unit root tests which fail to account for a trend break present in the data have non-pivotal limiting null distributions and are inconsistent under stable root alternatives. Assuming the putative break date to be known, Perron (1989) proposed new unit root tests which avoid these problems by modelling the trend break. However, if a break does not occur this approach loses considerable finite sample power through the inclusion of an unnecessary trend break regressor. Subsequent approaches have focussed on the case where the break date is unknown. Zivot and Andrews (1992) base a test on the most negative of a sequence, taken across all possible break dates, of the Perron (1989) statistics, while Perron (1997) first estimates the trend break location and then uses the Perron (1989) test for the estimated break date. The limiting distributions of the Zivot and Andrews (1992) tests depend on the magnitude of the trend break parameter which renders them infeasible in practice. The Perron (1997) approach is also problematic in that the break point estimator has a non-degenerate limit distribution when no break is present, with the result that the associated unit root test has a different large sample null distribution vis-à-vis the case where a trend break is present. Size-controlled inference can then only be achieved by using so-called conservative critical values corresponding to the case where no break is present, with an associated loss of efficiency where a break is present. As a result, Carrion-i-Silvestre et al. (2009), Harris et al. (2009) and Kim and Perron (2009) advocate approaches based on the use of pre-tests for the presence of a trend break. In the vector time series setting, un-modelled trend breaks cause similar problems for the co-integration rank tests of Johansen (1995). For example, Inoue (1999) documents large losses in finite sample power with the standard trace and maximum eigenvalue tests of Johansen (1995) when an un-modelled trend break is present in the data. As we will show in the simulation results we report in this paper, an un- modelled trend break also causes substantial over-sizing in the standard rank tests, consistent with the findings for standard unit root tests in Perron (1989). Surprisingly then, the literature on testing for co-integration rank in the presence of breaks in the deterministic trend function is relatively sparse compared to the univariate case. In the context of co-integration rank tests of the type considered in Johansen (1995), Johansen et al. (2000) develop likelihood ratio tests, analogous to those considered in the univariate case in Perron (1989), for the case where the break in the trend function occurs at a known point. Like Perron (1989) they consider both level break and trend break models, and extend to allow for multiple breaks in the trend function. Saikkonen and Lütkepohl (2000) for a level break (but no trend break) at a known date, Lütkepohl et al. (2004) for a level break (no trend break) at an unknown point, and Trenkler et al. (2007) for a trend break at a known date, propose further co-integration rank tests, in each case using the pseudo-GLS de-trending method outlined in Saikkonen and Lütkepohl (2000). All of these procedures assume that the autoregressive lag length is known to the practitioner. The approaches taken in the last three of these papers also differ from the approach taken in Johansen et al. (2000) according to how the data generating process [DGP] under consideration is constructed. While they adopt a components DGP, forming the observed process as the sum of the deterministic variables and an indeterministic vector autoregressive [VAR] process, Johansen et al. (2000), follow Johansen (1995) and place the deterministic variables directly into the VAR equation. Finally, Inoue (1999), who also assumes a known autoregressive lag order, develops Zivot and Andrews (1992) type co- integration rank tests by calculating with-break implementations of the Johansen (1995) tests over all possible break dates and basing a test on the most positive of these. The aim of this paper is to address these drawbacks with the existing tests in the literature. In order to focus attention on what we believe to be the empirically most relevant case, we follow Trenkler et al. (2007) and consider only the leading example of the trend break case, but allowing for the possibility of a simultaneous level break. We propose new testing procedures which in spirit generalise the approach taken by Carrion-i-Silvestre et al. (2009), Harris et al. (2009) and Kim and Perron (2009) to the setting of testing for co-integration rank. We consider two possible approaches depending on whether the deterministic component is included additively as in Trenkler et al. (2007) or directly into the co-integrated vector autoregressive [VAR] equation as in Johansen et al. (2000). In either case the first step in the procedure is based on the use of a consistent estimator of the break date. In the context of the component DGP of Trenkler et al. (2007), a multivariate generalisation of the first difference trend break estimator used in Harris et al. (2009) is proposed, along with a corresponding estimator obtained from the levels of the data, while for the Johansen et al. (2000) set-up a maximum likelihood estimator of the break date is used. Based on these break date estimators, for each of the two approaches an information-based method using a Schwarz (1978)-type criterion is then employed to select between the version of the model which includes a trend break (included at the relevant estimated break date) and that which does not. Each of the proposed procedures also employs a Schwarz-type criterion to select the autoregressive lag length. Conventional trace-type co-integration rank tests are then computed appropriate to the model selected by these Schwarz-type criteria. For each of the proposed procedures we establish that: (i) the estimator of the break fraction is consistent for the true break fraction; (ii) the information-based methods based on this estimator consistently select between the with-break and without-break variants of the model, and (iii) the resulting trace tests can be validly compared to known break date critical values in trend break case and to the without break critical values in the no break case. A consequence of our results is that, at least in large samples, the information-based methods we propose allow us to correctly identify whether we need to allow for a trend break in the model or not. This then implies that where a break is not present we will not see the loss in efficiency that is incurred by including a redundant trend break regressor in the model, and at the same time where a trend break is present we will not see the potentially large impact on the size and power properties of the rank tests that result from omitting the trend break. We present Monte Carlo simulation evidence which suggests that the procedure based on the Johansen et al. (2000) set-up is preferred and generally works very well even for a relatively small sample size such that the finite sample performance of this procedure is quite close to that seen for the benchmark rank tests which would obtain with knowledge of whether a trend break was present or not. The key findings of our Monte Carlo simulation exercise are presented here, while a more detailed set of results can be found in the accompanying supplement, Harris et al. (2015).","As discussed in the previous section, we now explore general approaches suggested by the structures of each of (1)–(3). The SC-VECM procedure ~~~~~~~~~~~~~~~~~~~~~ The SC-VECM procedure can then be described as follows. The SC-DIFF procedure ~~~~~~~~~~~~~~~~~~~~~ The SC-DIFF procedure can then be described as follows. The SC-VAR procedure ~~~~~~~~~~~~~~~~~~~~ The SC-VAR procedure is as follows.","Some remarks are in order. Simulation design ~~~~~~~~~~~~~~~~~ The tables of results given here are selected to illustrate the important features of the finite sample properties of the procedures. Nevertheless space constraints mean that results of the full experiment cannot be reported here, but they are made available in the accompanying supplement, Harris et al. (2015). Those results help to explain choices made in the reporting here. For example Tables 2–5 do not include results for SC-VAR because these tests were found to suffer from substantial size distortions when compared to SC- VECM and SC-DIFF, but results for SC-VAR are given in Harris et al. (2015). The supplement also provides finite sample evidence for the choice of 2 as the penalty for the break fraction parameter, rather than the usual 1.","We have focussed on the problem of testing for the co-integration rank in VAR processes of unknown lag order when a break in the deterministic trend component may be present at an unknown point in the sample. In order to simultaneously avoid the size and power problems which can result, even in large samples, from an un-modelled trend break and at the same time guard against the loss of finite sample efficiency which results from allowing for a trend break when no trend break is present, we have outlined an approach based on the use of information criteria. These criteria are used to select the autoregressive lag length and to select between the trend break and no trend break models, using a consistent estimate of the break fraction in the former case. Two possible frameworks were considered depending on whether the deterministic component was included additively in a components representation or directly into the VAR equation, the latter referred to here as the SC- VECM procedure. In each case these procedures were shown to deliver asymptotically correctly sized and consistent tests of the co-integration rank regardless of whether a trend break is present in the data or not. By selecting the no break model when no trend break is present, these procedures were also shown to avoid the potentially large power losses associated with tests which assume that a trend break is known to have occurred, when in fact no break is present. Monte Carlo simulation results were presented which suggest that the procedures generally performed well in practice with the SC-VECM procedure preferred overall."],["Numerous studies have documented the contribution of ICT to growth. Less has been done on the contribution of communications technology, the “C” in ICT. We construct an international dataset of fourteen OECD countries and present contributions to growth for each ICT asset (IT hardware, CT equipment and software) using alternative ICT deflators. Using each country's deflator we find that the contribution of CT capital deepening to productivity growth is lower in the EU than the US. Thus we ask: is that lower contribution due to a lower rate of CT investment or differing sources and methods for measurement of price change? We find that: (a) there are still considerable disparities in measures of ICT price change across countries; (b) in terms of growth-accounting, price harmonisation has a greater impact on the measured contributions of IT hardware and software in the EU relative to the US, than that of CT equipment; over 1996–2013, harmonising investment prices explains just 15% of the gap in the EU CT contribution relative to the US, compared to 25% for IT hardware; (c) over 1996–2013, CT capital deepening accounted for 0.11% pa (6% as a share) of labour productivity growth (LPG) in the US, compared to 0.03% pa (2.5% of LPG) in the EU-13 when using national accounts deflators; and (d) using OECD harmonised deflators, the figure for the EU-13 is raised to 0.04% pa (4% of LPG). --------------------------------------------------------------------------------","Numerous studies have sought to estimate the contribution of ICT to economic growth. Within the field of growth-accounting, studies include those from EUKLEMS (O’Mahony and Timmer, 2009; Van Ark and Inklaar, 2005) and earlier work from Oulton (2002) and Jorgenson (2001) for the UK and US respectively. Those studies showed that: first, in periods of fast technical change, price indices that do not correct for quality potentially (vastly) understate growth in real investment and capital services (Triplett, 2004); and second, better estimation of the ICT contribution requires separating out ICT equipment from plant & machinery in the estimation of capital services.12 However, in general, the growth contribution of the “C” in ICT has received less attention than the larger contributions of IT hardware and software (OECD, 2008). When we consider the role and ubiquity of the internet in business activity, this is potentially important. Papers that have studied and considered CT equipment as a distinct asset include (but are not limited to): Oliner and Sichel (2000), Doms (2005), Corrado (2011b), and Byrne and Corrado (2015). We construct an international dataset including the US and thirteen European economies to study the growth contribution of ICT with a particular focus on telecommunications (CT) capital. We treat all three types of ICT capital (i.e. IT hardware equipment, CT equipment and software) as distinct assets and use these data to review growth-accounting estimates for Europe relative to the US. In doing so, we first apply ICT investment deflators from each country’s national accounts as a benchmark. For all types of ICT capital, we find lower growth contributions in Europe than in the US. Thus we ask: “To what extent is the smaller contribution of telecommunications capital deepening in Europe due to: a) inconsistencies in the measurement of prices for tradeable goods; or b) a lower rate of real (current and past) CT investment?” To answer this we compare results with those based on sets of harmonised deflators constructed by the OECD (Schreyer and Colecchia, 2002).3 The harmonised deflators incorporate some explicit quality adjustment and tend to fall at a faster rate than those from European NSIs, which are typically not constructed using hedonic methods.4 We estimate the contribution of CT capital deepening to be positive for all countries and time periods studied and in the range of zero to 0.24% pa.5 In general, the CT contribution was higher in the late 1990s and early 2000s but has since declined. Applying national accounts deflators, CT capital deepening accounted for growth of 0.03% pa in the EU-13 in 1996–2013, compared to 0.11% pa in the US. Price harmonisation raises the EU contribution to 0.04% pa, implying that only a small proportion of the difference relative to the US is explained by inconsistencies in measures of price change for what are tradeable goods. Instead, the lower EU contribution is explained by weaker growth in CT capital services and a smaller income share, which in turn are driven by the rate of CT investment and the level of the real CT capital stock. The rest of this paper is set out as follows. Section 2 sets out the framework. Section 3 describes features of our dataset and its construction. Section 4 presents our results and finally, Section 5 concludes.","We conduct a sources of growth decomposition for fourteen OECD economies including the US and thirteen European countries.6 The underlying model is set out below. “Aggregation” In addition to our analysis by country, we also present (share-)weighted averages for country groups11 The relation between individual countries and higher-level groups is as follows. Real GDP in the (synthetic) country group grows as in (1) but where the shares for K and L are shares in the aggregate. Or put another way, all components of (1) can be presented as share-weighted averages using that country’s share in the nominal value-added of the group as weights:12 This technique creates a weighted average for the country group. Note that a strict aggregation would require detailed data on internal trade within the EU (or country group) and between the EU (or group) and the rest of the world (RoW). Existing literature ~~~~~~~~~~~~~~~~~~~ This paper is related to a number of strands in the literature. First, a number of researchers have sought to document the contribution of ICT to growth. Earlier studies include those from Jorgenson and Stiroh (2000) and Jorgenson (2001) for the US and Oulton (2002) for the UK, who showed that rapid decline in the prices of semiconductors and ICT capital was enabling a wave of improved productivity growth. Second, other studies have looked to make international comparisons of productivity growth and the contribution of ICT capital. For instance, Van Ark and Inklaar (2005) study the ICT contribution across countries using the EUKLEMS dataset (O’Mahony and Timmer, 2009) which includes harmonised constant-quality price indices for IT equipment (but not CT equipment or software). Other research to consider the deflation of ICT output and its effect on international comparisons includes (Wyckoff, 1995), which concludes that international variation in estimated price change (and therefore output) is largely due to methodological differences rather than real economic phenomena. Schreyer and Colecchia (2002) consider the deflation of ICT investment and construct harmonised deflators for all three types of ICT capital. They find that applying harmonised indices raises real growth in ICT investment and capital services in countries that do (did) not explicitly adjust for quality change. Similarly, Jalava and Pohjola (2008) use US investment price data for IT and CT equipment to re-estimate ICT contributions for Finland. Third, a more recent strand of the literature has sought to improve measurement of CT prices using US data. Doms (2005) and Corrado (2011a) and Byrne and Corrado (2015) each form new estimates of price change for CT equipment but do not conduct growth-accounting. Finally, our paper is also related to a new stream of literature that has sought to better estimate the growth contribution of telecommunications technology using improved measures of prices and capital services. Corrado and Jäger (2014) undertake a sources of growth analysis including CT capital and incorporating utilisation and network effects. Byrne et al. (2017) estimate price change for (constant-quality) ICT services provided via cloud computing and consider the implications of fast growing ICT inputs via the cloud (Byrne and Corrado, 2016), which current national accounting methods may not well measure. Work by Byrne and Sichel (2017) considers the impact of price mismeasurement in the high-tech sector and implications for productivity analysis.","The dataset was built primarily from country-level total economy national accounts data downloaded from OECD.Stat. As documented in the Appendix, where data were incomplete they were supplemented with other sources with some extrapolation or imputation where necessary. Time period ~~~~~~~~~~~ For most countries our data run from 1990 to 2013 but for others the data start later. We therefore conduct most of the analysis over 1996–2013 so that all countries are included and all country groups are balanced.13 Output and gross fixed capital formation (GFCF) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ For output, we use national accounts data on total economy nominal and real gross value added (GVA) at basic prices. Gross fixed capital formation (GFCF) data are nominal and real GFCF in intellectual property products (IPPs, that is: R&D; software; mineral exploration; and artistic originals; with GFCF in the latter two assets combined), computer hardware (IT), telecommunications equipment (CT), other plant and machinery (p), vehicles (v) and (non-residential) buildings (b).14 All specific adjustments to fill in incomplete data are set out in and below Appendix Table 4.15 On ICT investment, we note that measurement issues may exist in cases where different aspects of ICT capital are bundled in the same purchase. In the case of hardware, we note that the convention is that where software is bundled with hardware, and the values cannot be separated, then the investment transaction is recorded in hardware. We assume that the same applies to communications equipment, and that where software is bundled with CT and the values cannot be separated, then the transaction is recorded in CT equipment. Where CT is bundled with IT, we assume the transaction is recorded under IT hardware. However, we note the potential for practice to vary by country, with different countries potentially applying different methods and varying degrees of effort in unbundling various aspects of ICT investment. In Fig. 1 we document nominal GFCF in CT equipment across countries. To ease comparison, estimates are converted to US dollars (USD) using OECD purchasing power parity (PPP) data.16 As expected, nominal CT investment is highest in the US at $102.4bn in 2013. Strikingly, the sum of CT investment in the EU-13 in 2013 is substantially less, at $42.7bn.17 From Fig. 1 we see that, for most countries, nominal CT investment grew relatively strongly in the late 1990s and/or early 2000s but has since slowed or declined. Fig. 2 presents indices (2000=1) of nominal investment (in USD), for the US and EU-13, in: tangible capital excluding ICT equipment (“Tang excl ICT”); intangible capital as defined in the national accounts (“NA Intang”, that is, software, mineral exploration, artistic originals and R&D); IT hardware (“IT”); and communications equipment (“CT”). Prior to 2000 (95-00), nominal investment growth in CT equipment was almost twice as fast in the US (13.2% pa) than the EU-13 (7.7% pa). In the2000′s, prior to the financial crisis (00–07), growth in GFCF in CT equipment turned negative, at -1.8% pa in the US and -2% pa in the EU. In the recession and post recession period (07–13), GFCF CT equipment continued to decline at -1.2% pa in the US and -1% pa in the EU. Data for IT hardware show a similar pattern. In contrast, data for national accounts intangibles show a stronger recovery in investment following the financial crisis and the end of the ICT boom around 2000.18 Of course, although nominal CT investment has been declining in recent years, that does not mean that the real volume of CT investment has declined. That depends on the price per unit of CT investment, which in turn depends partly upon the efficiency characteristic of CT goods and how that has changed over time. In other words, £100 of telecoms investment in 2013 has considerably greater real volume than £100 of investment in years previous. Over the 1990s and early 2000s, investments in fibre optic cable and equipment increased network capacity by a factor of forty (Doms, 2005). Doms (2005) also notes that the pace of progress in fibre-optic capacity was well above that of Moore’s Law, with capacity doubling every year between 1996 and 2001. Clearly an appropriate price index must take this pace of technical change into account. Prices ~~~~~~ National price indices for value-added and GFCF are derived implicitly using current and constant price data and re-referenced to 2005=1 with all volumes re-estimated.19 The key lesson learned from research into the growth contribution of ICT equipment was the required use of constant-quality deflators (see for example Triplett (2004)). Where national practice is to not develop quality-adjusted deflators, researchers (or in some cases, NSIs) have favoured the use of either: a) the US price index (which incorporates some explicit quality adjustment) or some version adjusted for application to that country; or b) sets of harmonised deflators across countries, such as those in EUKLEMS (Timmer et al., 2007). Explicit quality adjustment (i.e hedonics) ensures that estimated changes in price and volume reflect improvements in power and other characteristics, as well as falls in listed prices, such that capital services better reflect changes in the real quantity of capital available for production. Although hedonic adjustment is sometimes regarded as superior, implicit quality adjustment methods (i.e matched models) can yield similar estimates of price change to those uncovered using hedonic techniques, provided the underlying data is sufficiently frequent and granular (Aizcorbe et al., 2003). National practice in EU countries National practice in constructing CT investment price indices varies. Fig. 3 presents national accounts CT GFCF deflators (blue) for each country in our datatset, compared with a harmonised CT deflator (red) from the OECD (Schreyer and Colecchia, 2002).20 From Fig. 3,21 national deflators for the US, Denmark, Portugal, Sweden, the UK and, to a lesser extent, Germany, show fast falling prices.22 In contrast, CT price indices for Austria, Spain, France, Ireland, Italy and the Netherlands are relatively flat with some exhibiting rising prices. National price indices for Belgium, Spain, Finland, France, Ireland, Italy and the Netherlands each have unusual features. That for Belgium exhibits price falls pre-2005, but price rises after. The index for Spain (ICT equipment index), shows rising prices until the 2000s, and gentle price falls for later years. Indices for Finland, Italy and the Netherlands are relatively flat pre-2005, but exhibit stronger price declines thereafter, and in the case of Finland, declines which are stronger than those in the harmonised index. The index for France is relatively flat throughout. Finally, that for Ireland is flat prior to 2005 before falling aggressively at a rate in line with the harmonised index, and then starting to rise such that it diverges strongly from the harmonised index. Since telecommunications equipment consists of tradeable goods, then under competitive conditions we would expect the law of one price to approximately hold, with some allowance for distribution margins, taxes and regulatory differences between countries. Therefore, as argued in Wyckoff (1995), we suggest that such divergence in measured prices reflects methodological differences between NSIs rather than real economic phenomena. Table 1 summarises our knowledge on the methods used to construct CT investment price indices in each country. We are currently unaware of any EU countries that apply hedonics in the estimation of CT investment prices.23 Regarding the US index to which OECD estimates are harmonised, Colecchia and Schreyer (2002) report use of hedonics in the ’switching equipment’ component index. However, this actually refers to the incorporation of research prices indices (Grimm, 1996; 1997) in the historic series through to approximately 1998 only.24 Thus, over the period for which we estimate, the US and therefore harmonised indices are not constructed from hedonic components.25 For most countries, including the US, matched model techniques are used instead, which as noted rely on detailed data in the cross-section and the time-series if they are to adequately account for fast quality change. In the US, the Federal Reserve Board price data, on which the Bureau of Economic Analysis (BEA) index is based, consist of weighted estimates from narrow product classes and are collected every quarter. Data for EU countries which also rely on matched model techniques exhibit far slower price falls, possibly because the data are not sufficiently detailed to observe fast improvements in quality. Harmonisation: Method and detail on underlying US index The red lines in Fig. 3 present the harmonised CT investment price index for each country, based on annual updates of the method described in Schreyer and Colecchia (2002). The method relies on the assumption that relative price changes for ICT assets move the same internationally as they do in the US. US ICT investment price indices are from the BEA. The BEA CT investment price index is a weighted aggregation of underlying indices, one of which is for “telephone switching equipment”, which previously (largely prior to the period considered in this study) incorporated hedonic adjustments (Colecchia and Schreyer, 2002) based on research from the Federal Reserve and others (Byrne and Corrado, 2015). Quality- adjustment since then in the official US data is implicit and based on matched model estimation.27 National price indices vs harmonised indices: Advantages and disadvantages The most obvious criticism of harmonising prices across countries is that differences in recorded price change may reflect genuine differences in the (product) composition of GFCF between countries. Similarly, they may also reflect differences in industrial composition, market structure and/or the degree of competition. Consider software. Some proportion of software investment is own- account (in-house). Harmonisation of software prices with US prices (as in the OECD data) assumes that the composition of software investment across countries is similar to that in the US, which may be a strong assumption (Schreyer and Colecchia, 2002).28 and the argument that investors face the same price across countries is a more difficult one to make. For this reason we are more cautious in the interpretation of our results and the impact of harmonisation in the case of software. In response to these criticisms, we note however that first, as observed in Corrado (2011b), the pattern and profile of price changes in CT equipment is remarkably similar across a diverse range of technologies, products and varieties. Thus, an argument that differences such as those recorded in Table 1 are due to differences in the basket of goods does not appear credible, and the observation from Corrado suggests that harmonised indices provide a better approximation than an index which inadequately accounts for quality change. Second, CT capital equipment consists of tradeable goods. Therefore, with some allowance for margins, tariffs and other frictions, we would expect the law of one price to approximately hold in competitive conditions. Third, if the sources and methods of national statistical institutes (NSIs) fail to adequately account for quality change, estimates of capital input will be biased downward. Harmonising with an index which better accounts for quality change due to more detailed data, such as the US index, helps correct for that bias. It is also more appropriate than using a range of indices with different degrees of known inaccuracy. Fourth, price indices track changes rather than the price level itself. If the price of CT equipment changes at a similar rate across countries, but national price indices record changes at different rates, then harmonised indices will generate estimates of that are more accurate and consistent across countries. Further, if relative price (e.g. CT to non-ICT) changes are similar across countries, then the harmonisation method shown here can exploit this relation.29 Based on these advantages and the assumption of competitive pricing, the use of US ICT price indices has become accepted research practice, and in some cases they are also applied by NSIs in national measurement.30","In this section we set out our growth-accounting results using alternative ICT investment deflators. Sources of growth: EU-13 & US ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Fig. 4 compares a decomposition of US productivity growth with the EU-13, for periods 1996-00 and 2000-13, estimated using national accounts ICT investment deflators. Labour productivity growth is higher in the US, in both periods, as is TFP growth and the contribution of capital deepening. Within capital types, the contribution of CT capital deepening is much less in the EU, as is the contribution of IT capital deepening and intangible capital deepening, where the latter is dominated by software. This is confirmed in Table 2 where we present a decomposition for the EU-13 compared to the US, first using national accounts ICT deflators and second using harmonised ICT deflators. All estimates are averages over 1996 to 2013. Column 1 is growth in value-added per hour worked (LPG); column 2 is the contribution of labour composition (labour services per hour worked); column 3 is the contribution of CT capital deepening (CT capital services per hour worked); column 4 is the contribution of IT capital deepening; column 5 is contribution of software capital deepening; column 6 is the contribution of capital deepening for all other (non-ICT) assets; and column 7 is growth in TFP, estimated as column 1 minus the sum of columns 2 to 6. The memo item in column 8 is the total ICT (IT, CT and software) contribution and that in column 9 is the contribution of CT capital deepening as a share of LPG, estimated as column 3 over column 1. Row 1 presents data for the US, row 2 for the EU-13 using national accounts ICT deflators, and row 3 for the EU-13 using harmonised ICT deflators. From column 1, we see that LPG was higher in the US, at 1.8% pa compared to 1.1% pa in the EU, whilst the contribution of labour composition was higher in the EU (0.21% pa) than the US (0.17% pa). Column 3 presents the contribution of CT capital deepening. Using national accounts deflators, the EU contribution was just over a quarter of the US contribution, at 0.03% pa compared to 0.11% pa in the US. From the third row, we see that applying harmonised ICT prices raises the EU CT contribution to 0.04% pa, but a large disparity relative to the US remains. In the final row, we estimate that of the gap relative to the US, over 1996 to 2013 just 15% of that gap can be explained by different measures of price change. Column 9 presents the share of labour productivity growth explained by CT capital deepening which is 6% in the US, compared to 2.5% in the EU using national accounts deflators, and 3.6% in the EU using harmonised deflators. In column 4, we estimate that the US contribution of IT hardware capital deepening stood at 0.19% pa compared to 0.08% pa in the EU. Applying harmonised deflators raises the EU contribution to 0.12% pa meaning that a larger proportion (25%) of the EU-US IT gap can be explained by different measures of price change. In column 5 we present the contribution of software capital deepening. Using national deflators that was 0.08% pa in the EU compared to 0.19% pa in the US. Applying harmonised deflators raises the EU contribution to 0.13% pa so that 33% of the EU-US gap can be explained by different measures of price change. Column 8 conducts the same exercise for the contribution of total ICT capital deepening. Over 1996–2013, that was 0.47% pa in the US compared to 0.18% pa in the EU-13. Harmonisation of ICT investment prices raises the EU contribution to 0.25% pa, meaning that of the gap relative to the US, 28% can be explained by inconsistent estimates of price change. Thus, for software, investment price harmonisation eliminates one-third of the gap relative to the US, whilst for IT hardware and the ICT aggregate, one-quarter is eliminated. In contrast, for CT equipment, only around one-sixth of the gap is eliminated.31 Harmonisation of ICT investment prices raises the contribution of ICT capital deepening in the EU-13 and therefore lowers the contribution of TFP growth, which was just 0.15% pa in the EU-13 compared to 0.55% pa in the US. Contribution of CT capital deepening Detailed growth-accounting results for each country, country group and asset are set out in Tables 6 and 7 in Appendix Balong with displays comparing growth contributions for each country relative to the US. To simplify, Fig. 5 presents estimates of the average contribution of CT capital deepening in each EU country relative to the US, using each alternative price index. The chart on the left is constructed using national accounts ICT deflators and the chart on the right using harmonised ICT deflators. Each country is presented relative to the US such that positive values represent growth higher than the US and negative values lower. Countries are ordered by country group: Scandinavia; Northern Europe (small); Northern Europe (large); and Southern Europe. For each country we present averages for two intervals: 1996-00 and 2000-13. Consider first the left-hand panel constructed using national accounts ICT deflators. The pattern is fairly uniform with contributions consistently smaller than the US, typically by around 0.1% pa. The exception is Sweden in the late 1990s, where CT capital deepening contributed 0.24% pa of LPG compared to 0.16% pa in the US. In the 2000s the Swedish contribution weakened such that the CT contribution was less than the US, which also fell. Now consider the right-hand panel. With a few exceptions the pattern remains similar, with only a small reduction in the negative gap relative to the US for most countries. The countries for which applying harmonised deflators has the largest impact are Austria (in particular the second period), Portugal (in particular the first period), Ireland and Belgium. Contribution of IT capital deepening Below we present a similar comparison but this time for the contribution of IT (hardware) capital deepening. From the chart on the left, constructed using national accounts ICT deflators, there is a consistently large negative differential relative to the US for all countries in the late 1990s.32 The lower contribution of IT capital deepening in Europe relative to the US is a well- documented finding (see for example, van Ark et al. (2008)). Now moving to the chart on the right, constructed using harmonised ICT deflators, we see that the picture has changed considerably. The negative gap in relative IT contributions in the late 1990s is considerably smaller and for some countries is eliminated entirely. In the 2000s, the majority of countries move from having lower relative IT contributions to having comparable contributions that are in some cases higher than that in the US. Although there is some variation by country, we find that price harmonisation explains a much larger proportion of the gap relative to the US in the case of IT hardware than for CT equipment. Summary ~~~~~~~ Table 3 summarises some other key results. We present data for 2000-13 for broader country groups on: the share of real CT investment in total real investment in 2000 and 2013 (columns 1 and 2); the share of nominal CT capital services in total nominal capital services in 2000 and 2013 (columns 3 and 4); LPG (column 5); the contribution of CT capital deepening (column 6); and the share of LPG accounted for by CT capital deepening (column 7, estimated as column 6 over column 5). All results presented are estimated using harmonised ICT deflators. From Table 3, columns 1 and 2, we see that real CT investment as a share of total real investment was highest in the US, at 4.5% compared to 2.2% in the EU in 2000, and at 5.3% compared to 3.1% in the EU in 2013. Data for EU country groups show that the CT real investment share is fairly similar across the EU, in terms of these broad regions. In 2000, it was highest in Scandinavia, and in 2013, it was highest in Southern Europe. In the final row, we present data for Sweden and note that the CT investment share in Sweden is higher than that for each EU country group and higher than the average for all countries (first row) in both 2000 and 2013. As a share of real investment, over the period CT investment has grown in all groups presented. The contribution of capital inputs depends on the factor income share (elasticity) and growth in capital services. In columns 3 and 4 we present nominal CT capital services (i.e. CT user costs or CT factor income) as a share of total nominal capital services (or gross operating surplus). First we note that, unlike the real investment share, over the period the share of CT user costs in the total has declined for all countries and country groups presented. Second, the CT income share was higher in the US than the EU-13, at 4.7% compared to 2.5% in 2000, and at 2.9% compared to 1.6% in 2013. CT income shares in EU country groups are strikingly similar to those for the entire EU-13 but are higher in Sweden, where it is also higher than that of the US in 2013. Columns 5, 6 and 7 present data on LPG, the contribution of CT capital deepening and the share of growth accounted for by that contribution. Over the period, LPG was strongest in the US, at 1.7% pa compared to 1% in the EU-13. Within the EU-13, LPG was similar across country groups, at 1.1% pa, except for Southern Europe where it was weaker, at 0.6% pa. Again, Sweden is an exception to this pattern with LPG of 1.5% pa, higher than the all country average of 1.3% pa. On the contribution of CT capital deepening, column 6 shows that was highest in the US, at 0.09% pa compared to 0.04% pa in the EU-13. Within the EU-13, the CT contribution was lowest in the large North European countries at 0.03% pa and highest in Scandinavia at 0.05% pa. In the small North European and South European economies it was the same as the EU-13 average at 0.04% pa. Again, the Swedish performance stands out. There the CT contribution was higher than that in EU country groups, and the all country average, and marginally lower than the US, at 0.08% pa. Column 7 presents the share of LPG accounted for by the contribution of CT capital deepening. It is highest in Southern Europe, at 7.1%, due to weak labour productivity growth of just 0.6% pa. In the US, CT capital deepening explained 5.6% of LPG in the 2000s, compared to 3.9% in the EU-13. EU country groups for which the share explained is higher than the EU-13 are Southern Europe (7.1%) and Scandinavia (4.4%).","Numerous studies have documented technological advances in ICT and sought to measure its contribution to growth. However the focus of those studies has generally been on the contribution of total ICT equipment or total ICT including software. That focus on IT hardware (or a group of assets dominated by IT hardware) means that the “C” in ICT has received far less attention. In this paper, we carry out a sources of growth decomposition for the US and thirteen EU countries and estimate the contribution of ICT capital under different assumptions regarding its price, with a focus on the contribution of CT equipment. We treat all three aspects of ICT capital including telecommunications equipment (i.e. IT hardware, CT equipment and software) as distinct assets in the estimation of capital services and test the impact of harmonising ICT investment prices with those from the US. National price indices for each ICT asset vary considerably but for all time periods we find that the contribution of CT capital deepening is positive for all countries and country groups and lies in the range of 0 to 0.24% pa. In general, the CT contribution was highest in the late 1990s and has declined in the 2000s. For the period 1996–2013, the contribution of US CT capital deepening is estimated at 0.11% pa in the US compared to 0.03% in the EU when using official national accounts investment price indices. Re-estimating capital services using harmonised ICT deflators raises the EU CT contribution to 0.04% pa, but a large disparity relative to the US remains. We estimate that of the gap relative to the US over 1996 to 2013, just 15% can be explained by different estimates of price change. In comparison, using national accounts prices, the US contribution of IT hardware stood at 0.17% pa against 0.07% pa in the EU. Applying harmonised prices raises the EU contribution to 0.09% pa meaning that one-quarter of the EU-US gap can be explained by different measures of price change. For software, using national accounts indices the US and EU growth contributions are 0.19% pa and 0.08% pa respectively. Price harmonisation raises the EU contribution to 0.12% pa meaning that a third of the gap in contributions can be explained by differing estimates of price change. Thus, for all forms of ICT capital, we estimate growth contributions in the EU that are considerably lower than those estimated for the US. However, only a small proportion of that gap is eliminated once we harmonise ICT investment prices with US official estimates. Of all ICT assets, that proportion eliminated is smallest in the case of CT capital. Within the EU and over the period 2000-13, growth in real CT investment is strongest in Sweden. As a share of real investment, CT investment is higher in Sweden than the US, as is the CT factor income share which partly depends on the rate of past investment. Similarly, the CT contribution in Sweden was higher than that for all EU country groups and marginally lower than that for the US, at 0.08% pa. In summary, there are still considerable disparities in measures of ICT investment price change across OECD countries. However, these disparities do not explain lower ICT growth contributions in the EU relative to the US. This is particularly so for CT capital, where inconsistencies in estimates of price change only explain approximately 15% of the EU-US gap. The remaining disparity in the contribution is therefore due to a lower rate of CT investment in the EU-13, in the sense of both current and past investments, which determine growth in capital services and the level of the real stock.33 Thus variation between countries is likely affected by issues such as public policy, public finances and public investment in cases where some part or all of the network infrastructure remains in the non-market sector. The timing of network build-out and adoption lags are also likely a cause of variation between countries."],["This study investigates the impacts of workers' remittances on human capital and labor supply by using data for 122 developing countries from 1990 to 2015. This topic has not been explored thoroughly at the aggregate level, mainly due to endogeneity of remittances and the difficulty in finding instruments to resolve this issue. To address the endogeneity of remittances, I estimate bilateral remittances and use them to create weighted indicators of remittance-sending countries. These weighted indicators are used as instruments for remittance inflow to remittance-receiving countries. Results obtained in this study indicate that remittances raise per capita health expenditures and reduce undernourishment prevalence, depth of food deficit, prevalence of stunting, and child mortality rate. Remittances also raise school enrollment, school completion rate, and private school enrollment. Although there is no difference in the impact of remittances on the health outcome of boys and girls, remittances improve the educational outcome of girls more than the educational outcome of boys. Further, remittances decrease the female labor force participation rate but do not affect the male labor force participation rate. --------------------------------------------------------------------------------","Between 1990 and 2015, the number of individuals living outside their countries of birth grew from 153 million people to 244 million people, which corresponds to 2.87% of the world population in the year 1990 and 3.32% of the world population in the year 2015 (United Nations). The total amount of remittances received has risen from $68 billion in 1990 to $553 billion in 2015. The average amount of money each migrant remitted (in 2011 constant dollars) has risen from $688 in 1990 to $2128 in 2015. These amounts include only remittances that have been sent through official channels (the World Bank, United Nations, author's calculations). This surge in the value of remittances has attracted the attention of many researchers. During recent years, different aspects of remittances have been under the scrutiny of researchers. One of the main aspects of remittances is impacts of remittances on remittance-receiving countries. The inflow of remittances can affect remittance-receiving countries in a variety of subjects such as the growth impact of remittance (Rao and Hassan, 2011; Gapen et al., 2009; Giuliano and Ruiz-Arranz, 2009; Ziesemer, 2012); the impact of remittances on poverty and inequality (Acosta, 2006; Adams and Page, 2005; Li and Zhou, 2013; Bang et al., 2016); remittances and financial development (Brown et al., 2013; Aggarwal et al., 2011; Coulibaly, 2015; Chowdhury, 2011); environmental effects of remittances (Li and Zhou, 2015); the impact of remittances on labor productivity (Al Mamun et al., 2015); the impact of remittances on political institutions (Williams, 2017). Although most researchers have investigated the impacts of remittance inflows on remittance-receiving countries, there are few studies on the impacts of remittance outflows on remittance-sending countries (Hathroubi and Aloui, 2016; Alkhathlan, 2013). There is a clear distinction between the remittance inflow to developing countries and the remittance inflow to developed countries. In the year 2015, the average workers' remittances to GDP ratio for developing country was 2% and for developed countries was 0.6%. Remittance to developed countries is not the primary interest of this paper; instead, the impacts of remittances on human capital and labor supply in developing countries is explored in this paper. In the past few years, many researchers have examined the impact of remittances on human capital or labor supply in a specific country or region. The primary barrier in evaluating the impact of remittances on human capital, in case of all developing countries, is endogeneity of remittances. Remittances are endogenous to education, health outcomes, and labor supply of those left behind. Reverse causality, common factors affecting both remittances and human capital, and measurement error are among the sources of endogeneity. Finding instruments to overcome endogeneity of remittances in one country is easier than finding valid and strong instruments for remittances in all developing countries. To address the endogeneity of remittances, this research uses a novel instrumental variable (IV) approach by incorporating three economic indicators of the remittance-sending countries as instruments: per capita Gross National Income (GNI), unemployment rate, and real interest rate. Since each remittance-receiving country has many countries as the sources of remittances, building the instruments requires knowing the bilateral remittances to calculate the weighted average indicators of the remittance-sending countries. Because bilateral remittances are not generally available, I estimate them and use them as weights to build the instruments. This estimation strategy allows me to utilize three valid and strong instruments, which are related to remittance-sending countries, to investigate the impacts of workers' remittances on human capital and labor supply in remittance-receiving countries. I also use this IV approach to study gender-specific impacts of remittances on health, education, and labor supply in developing countries. The results obtained in this paper indicate that remittances raise per capita out-of-pocket health expenditures and per capita total health expenditures and reduce undernourishment prevalence, depth of food deficit, prevalence of stunting, and child mortality rate. Remittances also raise school enrollment, school completion rate, and private school enrollment. As a robustness check, this paper investigates the overall impact of remittances on Human Development Index (HDI) and shows that remittances increase HDI. Another contribution of this paper is showing the gender-specific effects of remittances. Although there is no difference in the impact of remittances on the health outcome of boys and girls, remittances raise the education investment in girls more than in boys. Further, remittances decrease the female labor force participation rate but do not affect the male labor force participation rate. Remittances can lift budget constraints, thereby providing children in remittance- receiving households the opportunity to go to school or have better health outcomes. While remittances can benefit households by lifting liquidity constraints, migration of a family member can have a negative impact on the household's well-being. Migration of a productive family member may have disruptive effects on the life of the household. These observations lead to a fundamental development question: do remittances to developing countries contribute to higher investment in human capital? This research examines this question by exploring the major factors influencing education and health outcomes in developing countries. This paper addresses three main questions: To what extent can the remittances to developing countries improve health and education outcomes in terms of nutrition adequacy, mortality rates, school attendance, and educational attainment? If the remittances can influence the children's health outcome, school attendance, and educational attainment, is there any difference between the genders? What is the impact of remittances on labor force participation for men and women? As major components of human capital, health and education impact long-run economic growth. Therefore, answering these questions is crucial to evaluate the overall effect of workers' remittances on developing countries. The rest of the paper is arranged as follows. Background and literature review is provided in section 2. Section 3 discusses the data used in this paper. Section 4 is devoted to the econometric model. Section 5 presents empirical results, and Section 6 some conclusions.","Early studies on impacts of remittances on human capital (both health and education), financial development, poverty and inequality, and exchange rates did not consider the endogeneity of remittances. However, most of the recent studies about the impact of remittances on different aspects of remittance-receiving households and countries recognize the endogeneity of remittances and include different methodologies to resolve the problem of endogeneity of remittances. The traditional and most popular method of addressing the endogeneity of remittances is to use instrumental variables. While the instruments must be strong, the bigger challenge to researchers is to find valid (i.e. not correlated with the error term) instruments. Generally, three categories of instruments are used by different authors: instruments related to remittance-receiving countries, instruments related to remittance-sending countries, and instruments related to the cost of remittances. Not many papers have addressed the endogeneity of remittances at the aggregate level. Adams and Page (2005) is one of the first studies which addresses endogeneity of remittances in the country-level. They use three instrumental variables to account for endogeneity of the impact of remittances on poverty. The first instrument is the distance (miles) between the four major remittance-sending areas and remittance- receiving countries. The second instrument used by them is the percentage of the population over age 25 that has completed secondary education. This variable is correlated with education, and it is not a valid instrument in this research. The third variable used by them is government stability. The effect of government stability on remittances can be positive (if migrants have investment incentives, they remit more if their home countries are more stable) or negative (if migrants have altruistic incentives, they remit more if their home countries are less stable). Government stability also can be correlated with educational attainment and health outcome. This section is divided into three parts. First, a summary of papers which study the effects of remittances on health outcomes is provided. Then, papers which investigate the impacts of remittances on education are reviewed. Lastly, papers which examine the impacts of remittances on labor supply are summarized. Due to the importance of the endogeneity issue, special attention is paid to the instruments used by different authors. Remittances and health ~~~~~~~~~~~~~~~~~~~~~~ Researchers who have studied the impact of remittances on health outcome mainly focused on three health measurements: health expenditures, nutrition status, and child mortality rates. In the following, I briefly discuss papers who consider the impacts of remittances on these three aspects of health outcome. Note that health expenditure is not an aspect of health outcome per se, but it has a strong impact on health outcome. Ambrosius and Cuecuecha (2013) test the assumption that remittances are a substitute for credit by comparing the response to health-related shocks using Mexican household panel data. They find that the occurrence of serious health shocks that required hospital treatment doubled the average debt burden of households with a member having a health shock compared to the control group. However, households with a parent, child, or spouse in the US did not increase their debts due to health shocks. The authors use weighted linear regression to address the endogeneity problem. Therefore, in the case of Mexican families with a member facing a health shock, it can be concluded that remittances can play a substitute role for credit. Valero-Gil (2009) considers the effect of remittances on the share of health expenditures in total household expenditure in Mexico and finds a positive and statistically significant effect of remittances on the household health expenditure shares. The paper concludes that health expenditure is a target of remittances, and a dollar from remittances is being devoted in part to health expenditures. The author uses the percentage of return migration between 1995 and 2000 at the municipal level as an instrument for remittance. Amuedo-Dorantes and Pozo (2011) find that international remittances raise health care expenditures among remittance-receiving households in Mexico. They select two variables to be included as instruments: the road distance to the US border (from the capital of the Mexican state in which the household resides) and US wages in Mexican emigrant destination states. Their results suggest that 6 Pesos of every 100 Peso increment in remittance income are spent on health. Ponce et al. (2011) study the impact of remittances on health outcomes in Ecuador, and they do not find significant impacts on long-term child health variables. However, they find that remittances have an impact on health expenditures. They use historic state-level migration rates as an instrument for current migration shocks. Some researchers have used a variety of measurements of nutrition status as the factor representing health outcome. Food security level, household food consumption, child nutritional status, and low birth weight are some of the measurements of nutrition status. In the following, I introduce four of these papers. Generoso (2015) analyzes the main determinants of household food security using a survey conducted in Mali in 2005. The author uses a partial proportional odds logit model and shows that the diversification of income sources, through the reception of remittances, constitutes one of the main ways of coping with the negative impact of climate hazards on food security. Combes et al. (2014) explore the role of remittances and foreign aid inflows during food price shocks. They conclude remittance and aid inflows dampen the effect of the positive food price shock and food price instability on household consumption in vulnerable countries. Anton (2010) analyzes the impact of remittances on the nutritional status of children under five years old in Ecuador in 2006. Using an instrumental variables strategy, the study finds a positive and significant effect of remittance income on short-term and middle-term child nutritional status. The author uses the number of Western Union offices per 100,000 people at the province level as the instrument. Frank and Hummer (2002) find that membership in a migrant household provides protection from the risk of low birth weight among Mexican-born infants largely through the receipt of remittances. They did not account for endogeneity of remittances. Many researchers have used child and infant mortality rates as the factor representing health outcome. Reducing child mortality rate is one of the United Nations Sustainable Development Goals (SDG). The goal is by 2030 to end preventable deaths of newborns and children under five years of age, with all countries aiming to reduce neonatal mortality to less than 12 per 1000 live births and under-5 mortality to less than 25 per 1000 live births. In the following, some of these studies are discussed. Terrelonge (2014) examines how remittances and government health spending improve child mortality in developing countries and concludes remittances reduce mortality through improved living standards from the relaxation of households' budget constraints. Chauvet et al. (2013) consider the respective impact of aid, remittances and medical brain drain on the child mortality rate. Their results show that remittances reduce child mortality while medical brain drain increases it. Health aid also significantly reduces child mortality, but its impact is less robust than the impact of remittances. Kanaiaupuni and Donato (1999) consider the effects of village migration and remittances on health outcomes in Mexico and conclude higher rates of infant mortality in communities experiencing intense U.S. migration. On the other hand, mortality risks are low when remittances are high. Zhunio et al. (2012) study the effect of international remittances on aggregate educational and health outcomes using a sample of 69 low- and middle-income countries. They find that remittances play an important role in improving primary and secondary school attainment, increasing life expectancy and reducing infant mortality. Remittances and education ~~~~~~~~~~~~~~~~~~~~~~~~~ To investigate the impact of remittances on the education of those left behind, most of the researchers rely on school attendance as the main indicator of education. Many authors also examine the impact of remittances on child labor. The common belief is that if remittances decrease child labor, then they increase children's education. While school attendance frequently is used to measure investment on education, it represents the quantity of education and fails to capture the quality of education. Therefore, some researchers have used private school enrollment to demonstrate the quality of education. There are two theories about the impact of migration and remittances on education. On the one hand, as Bouoiyour and Miftah (2016) state, parents invest in the education of their children if they believe such investment generates a higher rate of return than the return on savings. Some parents in low-income households invest too little in the education of their children because they cannot afford to finance educational investments regardless of such returns. Hence, it is expected that the reduction of these liquidity constraints will make education more accessible to children from poor families. Therefore, financial transfers from family members can lift the budget constraints and allow parents to invest in the education of their children to the extent that is optimal for them. On the other hand, as Amuedo-Dorantes and Pozo (2010) state, the presence of family members abroad may induce changes in school attendance of children in migrants' households for a variety of reasons. Children may have less time to devote to schooling because they engage in market activities to earn income to defray migration-related expenses of households. Alternatively, children may leave school to perform necessary household chores that the absent migrant no longer attends to. Finally, if children get encouraged to migrate in the future, they may drop out of school if the origin country's education is not generally well-recognized in the destination. Among researchers who consider the impact of remittances on education and address the endogeneity of remittances, two instruments are used frequently. Some studies that recognize the endogeneity of migration and remittances use migration networks as the instrument. Many studies have used economic indicators of remittance-sender countries as instruments for remittances. The following papers have used migration networks as the instrument for remittances. Bouoiyour and Miftah (2016) use the history of migration networks and remittances costs as instruments to address the endogeneity of remittances. They find that the receipt of remittances in rural areas of southern Morocco has a significant positive effect on school attendance, especially for boys. Acharya and Leon-Gonzalez (2014) explore the effects of migration and remittance on the educational attainment of Nepalese children. Their results indicate that remittances help severely credit-constrained households enroll their children in school and prevent dropouts. Also, remittances help households that face less severe liquidity constraints increase their investment in the quality of education. By using migration networks as instruments, the authors gain similar results. Acosta (2011) studies the effects of remittance receipt on child school attendance and child labor in El Salvador. The paper uses El Salvadorian municipal-level migrant networks and the number of international migrants who returned two or more years ago as instruments to address the endogeneity of remittances. The author concludes insignificant overall impact of remittances on schooling; a strong reduction of child wage labor in remittance recipient households; and an increase in unpaid family work activities for children in those households. Moreover, while girls seem to increase school attendance after remittance receipts by reducing labor activities, boys, on average, do not benefit from higher education. Alcaraz et al. (2012) study the effects of remittances from the U.S. on child labor and school attendance in recipient Mexican households. They use distance to the U.S. border along the 1920 rail network as an instrument for the membership in the remittance recipient group. By using the differences-in-differences method, they find that the negative shock on remittance receipts in 2008–2009 caused a significant increase in child labor and a significant reduction of school attendance. As stated by Salas (2014) the common belief is that private schools provide better education compared to that provided by public schools. Besides increasing school enrollment, remittances can affect the choice of school type. The main limitation to access private education in developing countries is its cost. Remittances can lift the budget constraints and allow parents to invest in the quality of education of their children, as well as its quantity. Salas investigates the effect of international migration on children left behind in Peru. Historical department-level migration rate is used as an instrument for remittances. The model analyzes the role of international remittances on the investment decision between sending children to a public school or to a private school and finds that international remittances have a positive effect on the likelihood to send children to private schools. While using migration networks to address the endogeneity of migration allows us to assess the impact of migration on educational attainment, it is not the best possible instrument if we are interested in impacts of remittances on educational attainment. Using economic indicators of remittance-sending countries is more useful if we want to evaluate the impacts of remittances on investment in education. The following studies use economic indicators of remittance-sending countries to address the endogeneity of remittances. Calero et al. (2009) investigate how remittances affect human capital investments through relaxing resource constraints. They use information on source countries of remittances and regional variation in the availability of bank offices that function as formal channels for receiving remittances. They show that remittances increase school enrollment and decrease the incidence of child work, especially for girls and in rural areas. Besides increasing school enrollment, remittances affect the choice of school type. Remittances lead to a net substitution from the public to private schooling, hence increasing the quality of human capital investments in children. Bargain and Boutin (2015) explore the effects of remittance receipt on child labor in Burkina Faso. They use economic conditions in remittance-sending countries as instruments for remittances. The authors conclude while remittances have no significant effect on child labor on average, they reduce child labor in long-term migrant households, for whom the disruptive effect of migration is no longer felt. Fajnzylber and Lopez (2008) argue that in 11 Latin American countries with the exception of Mexico, children of remittances-receiving families are more likely to remain in school. The positive effect of remittances on education tends to be larger when parents are less educated. The authors use two external instruments based on the real output per capita of the countries where remittances originate. Amuedo-Dorantes et al. (2010) address the endogeneity of remittance receipt and find that the receipt of remittances by the households in Haiti lifts budget constraints and raises the children's likelihood of being schooled, and the disruptive effect of household out-migration imposes an economic burden on the remaining household members and reduces children's school attendance. They use two variables as instruments: weekly earnings of workers in the United States who are similar to potential Haitian remitters, and unemployment in those geographic areas in which the household is likely to have migrant networks. Some other papers use other instrumental variables or do not address the endogeneity of remittance. Bansak and Chezum (2009) examine the impact of remittances on educational attainment of school-age children in Nepal, focusing on differences between girls and boys. They use past literacy rates and political unrest by the district as instrumental variables to address the endogeneity. Their results indicate that positive net remittances increase the probability of young children attending schools. Yang (2008) concludes that favorable shocks in the migrants' exchange rates (appreciation of the host counties' currency versus Philippine peso) lead to enhanced human capital accumulation in origin households, a rise in child schooling and educational expenditure, and a fall in child labor. Edwards and Ureta (2003) examine the effect of remittances from abroad on households' schooling decisions using data for El Salvador. They examine the determinants of school attendance and find that remittances have a large and significant effect on school retention. The main issue with their paper is that they do not recognize the problem of endogeneity. Acosta et al. (2007) explore the impact of remittances on poverty, education, and health in eleven Latin American countries. While remittances tend to have positive effects on education for all Latin American countries with the exception of Jamaica and the Dominican Republic, the impact is often restricted to specific groups of the population (e.g. the positive effect of remittances on education tends to be larger when parents are less educated). Remittances and labor supply ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Researchers who study the impact of remittances on labor supply in remittance-receiving countries mainly focus on two aspects of labor supply: labor force participation rate and hours worked per week. As argued by Cox-Edwards and Rodriguez-Oreggia (2009), an important concept underlying the labor force participation decision is the notion of the reservation wage. The reservation wage is the lowest wage rate at which a worker would be willing to accept a particular type of job. A job offer involving the same type of work and the same working conditions, but at a lower wage rate, would be rejected by the worker. An increase in the reservation wage would reduce the probability that an individual participates in the labor force. One of the determinants of the reservation wage is non-labor income, which for an individual is a function of his or her own assets and the amount of income of other household members. The higher the level of income of the rest of the household (e.g. remittances), the higher the reservation wage of the individual, and the lower the probability that he or she participates in the labor force. Many researchers have looked at remittances as an additional non-labor income for recipient households and hypothesized that the presence of remittances would lead to a reduction in labor force participation among recipient household members. In studying the impact of migration and remittances on labor supply in recipient-households and receiving-countries, migration and remittances are endogenous. Migration and remittances are correlated with the error term due to a variety of factors. Reverse causality, omitted variable bias (uncontrolled common factors affecting both migration and remittances and labor supply) and measurement errors are among the potential sources of endogeneity. If the migrant's decision on how much to remit depends on whether the family members are looking for a job or not, reverse causality exists. Omitted variable bias may exist if remittances are related to the remittance- recipient family members' wealth or ambition, which may be correlated to the labor supply by the recipient. Many researchers use instrumental variables to address the endogeneity of migration and remittances. Particularly, migration networks are used frequently to address the endogeneity of migration and remittances. In the following, I review some papers which used this approach. Nguyen and Purnamasari (2011) apply an instrumental variable estimation method, using historical migration networks as instruments for migration and remittance receipts and find that, in Indonesia, migration reduces the working hours of remaining household members. Hanson (2007), based on rural households in Mexico in 2000, concludes individuals are less likely to participate in the labor force if their household either has sent migrants abroad or received remittances from abroad. The author also finds that women from high-migration states become less likely to work relative to women from low-migration states. The author argues that the unobserved characteristics of households that affect labor supply are also likely to affect whether households choose to send migrants abroad, and uses historical emigration rates as one of the approaches to address the endogeneity issue. Acosta (2006) controls for household wealth and uses selection correction techniques and concludes that, in El Salvador, remittances are negatively related to the adult female labor supply. However, on average adult male labor force participation remains unaffected. The author recognizes the endogeneity issue and uses the village migrant networks and the number of international migrants who returned two or more years ago as instruments for remittance receipts. Many authors use other instruments to address the endogeneity of remittances. For example, Amuedo-Dorantes and Pozo (2006) assess the impact of remittances on Mexicans labor supply. They use per capita count of Western Union offices in the state as an instrument for addressing the endogeneity of remittances. They conclude while overall male labor supply does not vary because of changes in remittance income, its composition by type of employment does. Unlike men, the overall female labor supply decrease due to changes in remittance income, although only in rural areas. Acosta et al. (2008) conclude that in all 10 Latin American and Caribbean countries examined, remittances have a reducing effect on the number of hours worked per week. This negative effect is present both in urban and rural areas. Afterward, remittances are instrumented with the share of remittance- receiving households in the country interacted with household characteristics that affect their likelihood to migrate. After instrumenting for endogeneity, their negative impact on labor force participation ceases to be significant in a number of cases. They also find that the reductions in labor supply caused by remittances tend to be much smaller among individuals with higher levels of schooling. Some researchers use other methods to address the endogeneity of remittances, or they do not address this problem at all. Guha (2013) states that with an increase in the non-wage income, such as remittances, the households will work less and have more leisure, as the loss in the wage income gets compensated by the remittance income. The author uses international remittances to Bangladesh as the reference case, and applies Dynamic Stochastic General Equilibrium (DSGE) model, to show that a sudden increase in foreign remittances results in a fall in the labor supply to the traded sector and a decline in the output in the traded sector. Cox-Edwards and Rodriguez- Oreggia (2009) use propensity score matching to separate persistent from sporadic remittances. They find limited evidence of labor force participation effects of persistent remittances. This implies that remittances are an integral part of remittance recipient households' income generation strategy. Acosta et al. (2009) find that in El Salvador an increase in remittance flows leads to a decline in labor supply. Kim (2007) studies the reasons for the coexistence of high unemployment rates with increasing real wages in Jamaica. The author concludes that households with remittance income have higher reservation wages and reduce labor supply by moving out of the labor force. Funkhouser (2006) uses longitudinal data from 1998 and 2001 in Nicaragua to examine the impact of the emigration of household members on household labor market integration and poverty. The author concludes that households with emigrant had a reduction in labor income compared to otherwise similar households.","This study examines 122 developing (poor and middle-income) countries1 from the six major developing world regions: Latin America and the Caribbean, sub-Saharan Africa, the Middle East and North Africa, East Asia and the Pacific, South Asia, and Europe and Central Asia. The income and regional classification in this paper follow the conventions of the World Bank. This panel of countries is sufficiently diverse, which means that these results are internationally applicable. The annual data has a time span of 1990–2015. All dollar values in this paper are constant 2011 US dollars. Data on remittances come from the International Monetary Fund (IMF), which has defined remittances as the sum of two components: personal transfers (workers' remittances) and compensation of employees. The World Bank has adopted the same definition. Compensation of employees refers to the income of border, seasonal, and other short-term workers who are employed in an economy where they are not resident and of residents employed by nonresident entities. Being present for one year or more in a territory or intending to do so is sufficient to qualify one as being a resident of that economy (IMF, 2009). Between 1990 and 2015, personal transfers constituted 66% of remittances, and compensation of employees constituted 34% of remittances. In the year 2015, these ratios were 65% and 35% respectively. As Chami et al. (2009a) and Chami et al. (2009b) show, this aggregation is not appropriate since the two different types of transfers have different properties and respond differently to economic shocks. Thus, I follow Combes et al. (2014) in using only workers' remittances as the measure of remittances. Therefore, the narrow definition of remittances (workers' remittances) is used in this paper to record remittances. Remittances are divided by remittance-receiving countries' populations and converted in constant 2011 international dollars. The income variable in equation (1) is per capita Gross Domestic Product (GDP) in constant 2011 international dollars adjusted for purchasing power parity (PPP). In this paper, I report the impacts of remittances on two types of human capital: health outcomes and educational attainment. After investigating the impacts of remittances on health outcomes and educational attainment, as a robustness check, I examine the effect of remittances on HDI. Finally, I estimate the impact of workers' remittances on labor force participation. To fully investigate the impacts of workers' remittances on all aspects of health and education outcomes in developing countries, nine health measurements and nine education indicators have been used as dependent variables. Using different health and education measurements is an excellent robustness check and ensures that demonstrated positive impacts of remittances on health and education outcomes do not happen by chance. Nine health measurements are examined in this research: per capita out-of-pocket health expenditure, per capita total health expenditure, life expectancy at birth, undernourishment prevalence, depth of food deficit, prevalence of stunting, neonatal mortality rate, infant mortality rate, and under-five mortality rate. Although per capita out-of-pocket health expenditure and per capita total health expenditure are not measures of health, because they have direct effects on health outcomes, they are included as two of the health outcome measurements. Seven explanatory variables are employed in health regressions as the covariates: per capita GDP, population over 65, urban population share, health expenditure to GDP ratio, food production index, government expenditure on education, and newborns protected against tetanus. Nine educational measurements are chosen in this paper: pre-primary gross enrollment rate, primary gross enrollment rate, secondary gross enrollment rate, tertiary gross enrollment rate, percent of children out of primary school, private primary enrollment rate, secondary compliment rate, primary compliment rate, and gender parity index for primary education. Two new explanatory variables are employed in the education section: Labor force participation and government expenditures on education as a percentage of GDP. Table 1 describes all dependent and explanatory variables alongside their descriptive statistics. Table 2 provides descriptive statistics for gender-specific variables. Except for remittances and HDI, data on all dependent and explanatory variables are from the World Development Indicators. HDI data are from the United Nations Development Programme (UNDP). This study uses three instruments to address the endogeneity of workers' remittances: weighted average per capita GNI, unemployment rate, and real interest rate of remittance-sending countries. Data on GNI and populations are from the United Nations. Data on unemployment rates and real interest rates are from the World Development Indicators. The real interest rate is defined as lending interest rate adjusted for inflation as measured by the GDP deflator. Data on out-of-pocket health expenditures and total health expenditures cover the years between 1995 and 2014. The share of out-of-pocket health expenditures in total health expenditure in developing countries has fluctuated in recent years. In 1995 and 2014 it was 40.8% and 36.2% respectively. However, there has been a stable increasing trend in the amount of per capita total health expenditure and per capita out-of-pocket health expenditure: they went up from $131.4 and $53.6 in the year 1995 (expressed in constant 2011 US dollars adjusted for PPP) to $530.3 and $191.9 in the year 2014 respectively. In recent years there has been a remarkable decline in undernourishment prevalence, depth of food deficit, and child mortality rates across the developing world. Undernourishment prevalence data cover years 1991–2015. Average undernourishment prevalence went down from 23.7% in the year 1991 to 12.67% in the year 2015. Depth of food deficit data cover years 1992–2015. Average depth of food deficit decreased from 175.6 kilocalories in the year 1992 to 91.9 kilocalories in the year 2015. In 1990, average neonatal, infant, and under- five mortality rates per 1000 live births were 39.3, 69.1, 99.7 respectively. In 2015, average neonatal, infant, and under-five mortality rates declined to 20.8, 34.6, and 46.4 respectively, in the developing world. As UNDP states, HDI is a summary measure of average achievement in key dimensions of human development: a long and healthy life, being knowledgeable, and having a decent standard of living. The HDI is the geometric mean of normalized indices for each of the three dimensions. The health dimension is assessed by life expectancy at birth, the education dimension is measured by the mean of years of schooling for adults aged 25 years and more, and expected years of schooling for children of school age. The standard of living dimension is measured by gross national income per capita. The HDI uses the logarithm of income, to reflect the diminishing importance of income with increasing GNI. The scores for the three HDI dimension indices are then aggregated into a composite index using geometric mean. Data on HDI are adopted from UNDP. While HDI data cover the years 1990–2015, male and female HDI data cover the years 2000, 2005, and 2010 to 2015. In recent years there has been a slight but smooth and stable decreasing trend in labor force participation rates in developing countries. In 1990, total, male, and female average labor force participation rates in developing countries were 68.1%, 82.7%, and 53.4% respectively. In 2015, total, male, and female average labor force participation rates in developing countries were 63.5%, 77.9%, and 48.9% respectively.","for i = 1, …, N and t = 1, …, Ti where H is a measure of health, education, HDI, or labor supply in country i at time t, β0 is the intercept, Rit is the per capita remittances received by country i at year t, Iit is the mean per capita income of country i at year t, Xit is a vector of other variables that potentially affect Hit, δi is region dummy and eit is the error term. Since all variables are in logarithm, β1 can be interpreted as the elasticity of human capital with respect to per capita remittances and β2 can be interpreted as the elasticity of human capital with respect to per capita income. The reason that region dummies are used in this research rather than country dummies is that for many countries, there is just one year of data available, and for using country- specific effects, those countries had to be dropped from the regression. To avoid dropping those countries, I used region-specific fixed effects rather than country-specific fixed effects. All six major developing world regions are included: Latin America and the Caribbean, sub-Saharan Africa, the Middle East and North Africa, East Asia and the Pacific, South Asia, and Europe and Central Asia. The income and regional classification in this paper follow the conventions of the World Bank. I followed Adams and Page (2005) in using region-specific effects as fixed effects. However, in those cases that more than one year of data were available for each country, as a robustness check, I also estimated the regressions by using country-specific effects. The results were similar in terms of sign and significance of the parameter estimates but magnitudes were slightly different. Correlation matrix between selected variables is provided in Appendix I. Several correlation coefficients are high, therefore, in order to avoid multicollinearity effect arising from that, as a general rule, I try to not include highly correlated explanatory variables in a same regression equation. For example, in regressions corresponding to the impact of remittances on child mortality rate, rather than the dollar value of health expenditures, the ratio of health expenditures to GDP is used. Also, in all regression equations for the case of the impact of remittances on education, rather than the dollar value of government expenditure on education, its ratio to GDP is used.2 Some of the variables used in the regression models, like per capita GDP, are non-stationary in the level, but they are first-difference stationary. Since there are non-stationary variables in the regression models, the results of the regressions might be spurious. To check whether the results of the regressions are spurious or not, for the regressions with non- stationary dependent variables, I implement panel cointegration tests. If cointegrations exist, then it can be concluded that the error term is stationary, which means the results of the regressions are not spurious. Two tests are used to check whether panel cointegrations exist: the Kao Residual Cointegration Test proposed by Kao (1999) (also known as the residual-based Augmented Dicky Fuller test) and the Johansen Fisher Panel Cointegration Test proposed by Johansen (1988). In both tests, the null hypothesis is that cointegration does not exist. I implemented both tests for all models with non-stationary dependent variables and all p-values were less than 0.05. Therefore, based on the results of both tests, the null hypothesis is rejected, which means that a cointegration relationship exists among the variables. In other words, the error terms are stationary, and the results of the regressions are not spurious. Another econometric issue that arises here is that the error terms are autocorrelated and also they do not have constant variance, which shows the existence of heteroskedasticity. To overcome autocorrelation and heteroskedasticity, Newey-West Heteroskedasticity and Autocorrelation Consistent (HAC) standard errors are used. In other words, since I worry that there might be some serial correlation in the error terms even after controlling for fixed effects and endogeneity of remittances, I adjust standard errors to account for heteroskedasticity and autocorrelation by using autocorrelation HAC standard errors. Another issue that needs to be addressed is that whether standard errors should be adjusted for clustering. There is no cluster in the population of interest that is not represented in the sample. Moreover, except for the case of Ordinary Least Squares (OLS) results which are provided in the tables only for comparison purposes, fixed effects are included in the regressions and there is no need to adjust standard errors for clustering. (Abadie et al., 2017) Remittances may be endogenous to education or health outcomes in the remittance-receiving countries. The endogenous relationship between workers' remittances received and schooling decisions or health outcomes can be explained via three reasons. First, there exists a reverse causality between human capital (education attainment and health outcomes) and remittances. An individual may decide to migrate and send remittances because he or she has school-age or sick and undernourished children, and in turn, remittances affect educational investment or health outcomes by loosening liquidity constraints. The decision to live outside the country and send remittances is determined simultaneously with the decision regarding health and education expenditures. Second, there exist unobserved characteristics included in the error terms that may be correlated with both the decision to send remittances and the decision to send children to a school, or how much the household spends on their nutrition and health (e.g. ability or ambition). Third, measurement error is another source of endogeneity. Officially recorded remittances do not include remittance in kind, unofficial transfers through kinship or through informal means such as hawala operators, friends, and family members. The negative impact of transactions costs on remittances encourages migrants to remit through informal channels when costs are high. Transfer costs are higher when financial systems are less developed. Evidence from household surveys also shows a sizable informal sector (Freund and Spatafora, 2008). Since poor countries usually are less financially developed, migrants from poor countries are more likely to remit through informal channels. Because income is also correlated with human capital investment, measurement error is also one of the sources of endogeneity. Therefore, the least squares estimates of the impact of remittances may be biased as remittance is endogenous. The traditional way, which is followed in this paper, is to resolve the endogeneity problem by using instrumental variables. In the following, the instruments suggested by other researchers are discussed and then I introduce the instruments that are used in this paper. Three kinds of instruments are proposed by scholars. The first category of variables is related to remittance-receiving countries. The main problem with using these instruments is that they can easily be correlated with education, health or labor supply of the same country in a way other than through remittances or covariates. Therefore, they can be invalid instruments. The second category of variables is related to the cost of remittances (e.g. the number of branches of Western Union). These instruments are usually strong and valid, but data on these variables are mainly unavailable for most developing countries. The third category of variables is related to remittance-sending countries. Since these variables are not related to remittance-receiving countries, after controlling for some covariates, they are valid and strong. However, the main problem is that each remittance-receiving country receives remittances from many remittance-sending countries, and due to lack of information on bilateral remittances, the weights of remittance-sending countries are unknown. Aggarwal et al. (2011) use economic conditions in the top remittance-source countries as instruments for the remittances flows. They argue economic conditions in the remittance- source countries are likely to affect the volume of remittance flows that migrants are able to send, but are not expected to affect the dependent variables in the remittance receiving countries in ways other than through its impact on remittances or covariates. However, because bilateral remittance data are largely unavailable, they identify the top remittance-source countries for each country in the sample, using 2000 bilateral migration data. The dataset identifies the top five OECD countries that receive the most migrants from each remittance-recipient country. They construct instruments by multiplying the per capita GDP, the real GDP growth, and the unemployment rate, in each of the top five remittance-source countries by the share of migration to each of these five countries. The technique used in this paper is similar to Aggarwal et al. (2011). However, I use an estimation of bilateral remittances and use them as weights of remittance-sending countries. In this research, I introduce three instruments that are correlated with remittances and uncorrelated with the dependent variables unless through explanatory variables. Workers' remittances, by definition, are money sent by migrants from host countries (remittance-sending countries) to their home countries (remittance-receiving countries). The value of remittances hinges on both remittance-sending and remittance- receiving countries economic variables. Since the dependent variables belong to remittance-receiving countries, in order to ensure the instruments are valid, we should select a number of economic variables from remittance-sending countries. Three variables used in this research as instruments are per capita GNI, unemployment rate, and real interest rate. If the remittance-sending country's per capita GNI increases, it means migrants' income has increased, which means they have more money available to spend and remit. Therefore, it can be expected that remittances increase in response to a rise in the remittance-sending country's per capita GNI. If the unemployment rate goes up in the remittance-sending country, some migrants will lose their jobs, and the total remittances sent by migrants will decrease. If the real interest rates go up in the remittance-sending country, migrants have more incentive to invest in the host country rather than the home country, and they will remit less. Table 3 shows the results of regressing per capita remittances on these three instrumental variables while in each column different covariates have been used. Note that in all seven columns the parameter estimate for the impact of per capita GNI in remittance-sending countries on remittances is significant and positive, the parameter estimate for the impact of unemployment rate and real interest rates in remittance-sending countries on remittances is significant and negative. The IMF provides workers' remittances data annually. Data are available at the aggregate level for each country. Unfortunately, the IMF does not provide bilateral remittance data. Bilateral remittance means how much money a remittance-receiving country receives from each specific remittance-sending country. Therefore, although we know how much remittance each remittance-receiving country receives in any year in total, we do not know how much of the received remittances comes from a specific remittance-sending country. Fortunately, the UN has provided a comprehensive bilateral migration data from 1990 to 2015. This means we know how many immigrants live in each host-country, and we also know how many of them are from each specific home-country. I used this dataset,4 per capita GNI of remittance- sending countries, and per capita GNI of remittance-receiving countries to estimate bilateral remittances. I then use estimated bilateral data to construct weighted averages of per capita GNI, unemployment rate, and real interest rate of remittance-sending countries. These weighted average indicators are used as instruments to address the endogeneity of remittances. Before discussing the results of the regressions, I will provide some evidence about the correctness of the estimated bilateral remittance data. Data on bilateral remittances are mostly unavailable. Even where bilateral remittances are reported, they may not be accurate, because funds channeled through international banks may be attributed to a country other than the actual source country (Ratha, 2005). One reason why bilateral remittance data is inaccurate is that financial institutions that act as intermediaries often report funds as originating in the most immediate source country. For example, the Philippines tends to attribute a large part of its remittance receipts to the United States because many banks route their fund transfers through the United States (Ratha, 2005). Only a few papers have used bilateral remittances data. and only for limited recipient-country sender-country pairs (Lueth and Ruiz-Arranz, 2008; Frankel, 2011; Docquier et al., 2012). I compare the estimated bilateral remittances used in this paper with the corresponding part of the datasets they used in their studies. The actual data consists of 12 receiving countries, 16 remittance-sending countries, and 1744 observations.5 The Pearson Correlation Coefficient between actual data and estimated data is 0.666 and is statistically significant, indicating there is a strong positive linear correlation between actual data and estimated data. This study uses Kolmogorov-Smirnov's two-sample test to check if the distribution of the two samples is the same or not. The null hypothesis is that the two distributions are not statistically different from each other. Using the panel data, the p-value for the test is 0.17, and I fail to reject the null. Two main concerns of using instrumental variables method, to address the problem of endogeneity, are validity and strength of the instruments. The reason that in this paper bilateral remittances are estimated, and economic variables related to remittance-sending countries are used is to guarantee the validity of the instruments. More specifically, selecting economic variables from remittance-sending countries, rather than from remittance-receiving countries, as instruments is to ensure the validity of instruments. My identifying assumption is that per capita GNI, unemployment rate, and real interest rate in remittance-sending countries do not affect health outcomes, educational attainment, and labor force participation in remittance-receiving countries other than through remittances, per capita GDP of remittance-receiving countries, or other covariates included the regressions. The other concern about the instruments is whether they are strong or not. The whole process of estimating bilateral remittances and use them as the weights of remittance-sending countries is to ensure that instruments are strong. To investigate the strength of the instruments, first stage regressions are examined. The results of the first stage regressions are provided in Table 3. Note that all F-statistics for weak instruments are greater than 50, which indicates the instruments are very strong. There is a clear advantage in using estimated bilateral remittances to construct the instruments over using top 5 or 10 remittance-sending countries (the method used by other researchers). This advantage can be demonstrated by the substantial improvement in the F-statistic for weak instruments of this paper and similar papers. For example, F-statistic for two of the main instruments used in this paper, GNI per capita and unemployment rate in remittance-sending countries, are 342 and 98 respectively. These two F-statistics are remarkably higher than F-statistic in similar studies which use GNI per capita and unemployment rate in just top one or three or five remittance-sending countries as instruments.","This section is divided into four subsections. At first, I investigate impacts of remittances on health outcomes in Tables 4–6 and gender-specific impacts of remittances on health outcome in Tables 7 and 8. Then, I explore the impacts of remittances on education in Tables 9–11 and gender-specific impacts of remittances on education in Tables 12–14. Afterward, I look into the impact of remittances on HDI in Table 15. Finally, I examine the impacts of remittances on labor supply in Table 16. In all regressions, the results of the IV approach are reported alongside the results of the OLS model and Fixed Effect (FE) model. FE model and OLS model are similar in all aspects except one. Unlike OLS model, FE model includes region-specific effects and addresses time-invariant cross-country effects. However, both OLS and FE results are inconsistent and biased. They are presented in this paper for comparison purposes. To address both endogeneity and time-invariant cross- country effects simultaneously, all IV estimates also include region-specific effects. Remittances and health ~~~~~~~~~~~~~~~~~~~~~~ This study closely follows the literature on the choice of health measurements. Four categories of health measurements are used here as dependent variables: health expenditures, life expectancy at birth, nutrition status, and child mortality rates. Per capita out-of-pocket health expenditure and per capita total health expenditure are the variables representing health expenditures. Undernourishment prevalence, depth of food deficit, and prevalence of stunting are three variables representing nutrition measurements. Neonatal, infant, and under-five mortality rates represent child mortality rates. The estimates for the impacts of remittances on out-of-pocket per capita health expenditures, per capita total health expenditures, and life expectancy are presented in Table 4. For many countries, just one year of data is available, and for using country- specific effect, those countries had to be dropped from the regression. Due to this data limitation, this study controls for region-specific effect rather than country-specific effect. However, for those cases where more data are available, at least two years of data for each remittance-receiving country, it is possible to control for country-specific effect. As a robustness check, I use country-specific effect in those regressions and investigate the impact of remittances on health outcomes. The results are the same as regressions with region-specific dummies in terms of sign and significance of remittances, but the magnitudes are slightly different.6 In the first regression, per capita out-of- pocket health expenditure is used as the dependent variable. Population over 65 and urban population share are used as explanatory variables alongside per capita remittances and per capita GDP. Both OLS and FE models underestimate the impacts of remittances on health expenditures and the effect captured by the IV is larger than the effects captured by the OLS or the FE. The results of IV regression suggest that, on average, a 10% increase in per capita remittances and per capita GDP will lead to a 1.5% and 6.4% increase in per capita out-of-pocket health expenditure, respectively. In the second regression, per capita total health expenditure is used as the dependent variable. The results of IV regression suggest that, on average, a 10% increase in per capita remittances and per capita GDP will lead to a 1.1% and 8.9% increase in per capita total health expenditure, respectively. Note that the impact of per capita remittances on per capita total health expenditure is smaller than the impact of per capita remittances on per capita out-of- pocket health expenditure. The reason is that total health expenditure is the sum of private and public health expenditure and remittances have almost no effect on public health expenditure. In the third regression, life expectancy at birth is used as the dependent variable. The results of IV regression suggest that, on average, a 10% increase in per capita remittances and per capita GDP will lead to a 0.3% and 0.4% increase in life expectancy at birth respectively. The impacts of per capita remittances, per capita GDP, alongside other variables, on nutrition variables are provided in Table 5. Undernourishment prevalence, depth of food deficit, and prevalence of stunting are used as the fourth, fifth, and sixth measurements of health outcomes. Due to restrictions in the availability of data, the number of observations in each regression varies considerably. For example, in the regression with undernourishment prevalence as the dependent variable, the number of observations is 1629, but in the regression with prevalence of stunting as the dependent variable, the number of observations is 412. Food production index and urban population share are used as explanatory variables alongside per capita remittances and per capita GDP. The results of IV regressions suggest that, on average, a 10% increase in per capita remittances and per capita GDP will lead to, respectively, a 1.5% and 4.3% decline in undernourishment prevalence, a 1.9% and 5.9% decline in depth of food deficit, and a 1% and 2.7% decline in prevalence of stunting. Note that the magnitude of IV coefficients for per capita remittances are substantially larger than those from the OLS and FE estimations in all three regressions which means OLS and FE models underestimate the impacts of workers' remittances on undernourishment prevalence, depth of food deficit, and prevalence of stunting. The impacts of per capita remittances, per capita GDP, alongside other variables, on child mortality rates are provided in Table 6. In order to investigate the impact of remittances on child mortality rate, health expenditure to GDP ratio, urban ratio, and newborns protected against tetanus are chosen as control variables. Since high-income countries are primary funders of the World Health Organization (WHO), their economic conditions can affect WHO's budget, which in turn affects child mortality rates in developing countries through different global vaccination programs implemented by WHO. Controlling for the variable newborns protected against tetanus can assure the validity of the instruments used in the IV regressions. Both OLS and FE models underestimate the impacts of remittances on reducing child mortality rates. Remittance has negative and statistically significant effects on child mortality rates. Based on the IV results, on average, a 10% increase in per capita remittances causes a 1% decrease in neonatal mortality rate, a 1.7% decrease in infant mortality rates, and a 1.9% decrease in under five mortality rates. Also, on average, a 10% increase in per capita GDP leads to a 3.9% decrease in neonatal mortality rates, a 4.2% decrease in infant mortality rates, and a 4.7% increase in under five mortality rates. Gender-specific data are available only for four health outcome variables: life expectancy at birth, prevalence of stunting, infant mortality rate, and under-five mortality rate. Table 7 provides the gender-specific impacts of per capita remittances on life expectancy at birth and prevalence of stunting. Remittance has positive, statistically significant, and almost similar effects on male and female life expectancy. On average, a 10% increase in per capita remittances causes a 0.3% increase in male and female life expectancy at birth. Also, on average, a 10% increase in remittances leads to a 1.3% decrease in male and 1.2% decrease in female prevalence of stunting. Table 8 provides the gender-specific impacts of per capita remittances alongside per capita GDP, per capita health expenditure, urban population share, and newborns protected against tetanus on infant mortality rates and under-five mortality rates. There are negligible differences on impacts of remittances on child mortality rates between two genders: on average, a 10% increase in the amount of per capita remittances leads to a 1.31% and 1.35% reduction in infant mortality rates for males and females, respectively. Also, on average, a 10% increase in the amount of per capita remittances leads to a 1.43% and 1.49% reduction in under 5 mortality rate for males and females, respectively. These results show that when households in developing countries receive remittances from their family members living abroad, they spend the received money almost equally on the health status of their boys and girls. As we will see, this is not the case for the impact of remittances on education as remittances increase the female education more than male education. Remittances and education ~~~~~~~~~~~~~~~~~~~~~~~~~ In all regressions with education measurements as dependent variables, per capita remittances, per capita GDP, labor force participation rate, and government expenditure on education are used as explanatory variables. Pre-primary enrollment rate, primary enrollment rate, and secondary enrollment rate are the first three education measurements. The impacts of per capita remittances on these three variables are displayed in Table 9. A rise in workers' remittances (which is represented by per capita remittances) or in income measurement (which is represented by per capita GDP) can loosen the budget constraints for households in developing countries and motivate them to invest more in their children's education. The results provided in Table 9 show how worker’ remittances can increase children's school attendance in developing countries. On average, a 10% rise in per capita remittance increases pre-primary and secondary enrollment by 3.5% and 0.6%, respectively. However, based on the IV model, while a 10% rise in per capita remittance increases primary enrollment by 0.16%, this impact is not statistically significant. Also, on average, a 10% rise in per capita GDP increases pre-primary, primary, and secondary enrollment by 6.6%, 0.7%, and 3.2%, respectively. When labor force participation increases, parents have more incentive to invest in their children's education. If parents predict that their children, especially their girls, will not enter the job market, then they will have less incentive to spend on their children's education. Therefore, it is expected that the labor force participation rate has a positive impact on education. On average, a 10% rise in the labor force participation rate increases pre-primary, primary, and secondary enrollment by 15.9%, 4.8%, and 3.5%, respectively. Also, when the government expenditure on education increases, the marginal rate of return on households' investment in education increases. The increased marginal rates of returns, in turn, promote households to spend more on their children's education. On average, a 10% rise in government expenditures on education increases pre-primary, primary, and secondary enrollment by 3.8%, 0.6%, and 2.5%, respectively. The substantial difference between the impacts of per capita remittances on pre-primary enrollment rate, primary enrollment rate, and secondary enrollment rate can be explained by considering the mean these variables: 39.3%, 100.6%,7 and 60.7%, respectively. Households in developing countries consider pre- primary education as a luxury commodity. However, they consider primary education, which is required in many countries, as a necessity. Therefore, the elasticity of pre-primary enrollment rate with respect to per capita remittances or per capita GDP is expected to be higher than the elasticity of primary enrollment rate. However secondary education is less of a necessity in comparison with primary education, and tertiary education is less of a necessity in comparison with secondary education. Therefore, it is reasonable that secondary enrollment rate has a higher elasticity than primary enrollment rate, and tertiary enrollment rate has a higher elasticity than secondary enrollment rate. The impacts of per capita remittances on tertiary enrollment rate, primary age out of school rate, and private primary enrollment rate are shown in Table 10. The results presented in Table 10 show that, on average, a 10% rise in remittance per capita increases tertiary enrollment rate by 1.1%. However, based on the IV model, while a 10% rise in per capita remittance decreases out of school rate of primary age children by 1%, this impact is not statistically significant. The last dependent variable used in Table 10 presents the impact of workers' remittances on the quality of children's education rather than the quantity. On average, a 10% rise in per capita remittances increases enrollment in private primary schools by 2.8% which means that a rise in remittances leads to a net substitution from public to private primary schools, hence increasing the quality of children's education. In contrast, a rise in per capita GDP leads to a net substitution from private to public primary schools. One explanation is that as per capita GDP of developing countries increase, the quality of public schools also increase and parents will have less incentive to send their children to private schools. As Table 11 shows, on average, a 10% rise in per capita remittance increases primary completion rate by 0.6% and secondary completion rate by 0.9%. Once again, the difference in the impacts of per capita remittances on these two measurements of education can be explained by taking into account the average primary completion rate and lower secondary completion rate: 80% and 60%, respectively. Households in developing countries consider primary education a necessity and secondary education as less of a necessity than a luxury commodity. Therefore, it is reasonable that secondary completion rate has higher elasticity than primary completion rate with respect to per capita remittances and with respect to per capita GDP. Based on the descriptive statistics, boys' average primary enrollment rate is 103.4%8 and girls' average primary enrollment is 97.4%. This means the average households' investment in girls' primary education is less than their investment in boys' primary education in developing countries. Therefore, the marginal rate of return on primary education is higher for girls than boys. Hence, the inflow of remittances, which loosens the budget constraints for households in developing countries, is invested more heavily in girls' primary education than boys. This, in turn, raises the primary gender parity index, though just by a small amount. On average, a 10% rise in per capita remittance increases the primary gender parity index by 0.2%. This impact is statistically significant in all three models: OLS, FE, and IV. Workers' remittances increase the gender parity index. The rise in gender parity index means when households in developing countries receive remittances from the migrants, they spend a higher proportion of it on girls' primary education than on boys. This is particularly important because, as Bansak and Chezum (2009) state, if remittances do affect human capital positively, then not only will remittances affect long-run growth in developing countries, but the opportunities for women should improve as the female population becomes more educated in these countries. Table 12 provides gender- specific effects of per capita remittances, per capita GDP, government expenditure on education, and gender-specific labor force participation rate on male and female pre- primary enrollment rates and primary enrollment rates. On average, a 10% increase in per capita remittances increases male and female pre-primary enrollment rates by 3.4% and 3.5%, respectively. In contrast, while per capita remittance has no statistically significant impact on male primary enrollment rate, on average, a 10% increase in per capita remittances increases female primary enrollment rate by 0.3%. This asymmetric impact of remittances on female and male primary enrollment rate is in line with the positive impact of remittances on the gender parity index of primary education. There is a remarkable heterogeneity between the impact of the male labor force participation rate on male education and the impact of the female labor force participation rate on female education. While male labor force participation has no effect on male education investment, female labor force participation has a positive and significant effect on female education investment. On average, a 10% increase in the female labor force participation rate increases female pre-primary and primary enrollment by 3.1% and 0.6%, respectively. The impacts of the same explanatory variables on secondary and tertiary enrollment rates are provided in Table 13. Results provided in Table 13 show that per capita remittance has a greater impact on the female secondary and tertiary enrollment rates than on the male rates. On average, a 10% increase in per capita remittances increases male and female secondary enrollment by 0.7% and 0.9%, respectively. On average, a 10% increase in per capita remittances increases male and female tertiary enrollment by 0.6% and 1.3%, respectively. The impacts of the explanatory variables on male and female primary and secondary completion rates are provided in Table 14. The results provided in Table 14, confirm that remittance has a greater effect on the female primary and secondary completion rates than on male rates. On average, a 10% increase in per capita remittances increases the male and female primary completion rates by 0.5% and 0.9%, respectively. On average, a 10% increase in per capita remittances increases the male and female lower secondary completion rate by 0.8% and 1.2%, respectively. Comparing the elasticities obtained for male and female education with respect to per capita remittances in Tables 12–14 shows that households in developing countries, upon receiving remittances, invest in girls' education more than boys' education. Remittances and HDI ~~~~~~~~~~~~~~~~~~~ As a robustness check for the impact of remittances on human capital, I examine the effect of remittance on HDI in developing countries. Table 15 provides the impacts of per capita remittances, per capita GDP, government expenditures on education, and per capita health expenditures, on total, male, and female HDI. As Table 15 shows, on average, a 10% rise in per capita remittances increases total HDI, male HDI, and female HDI by 0.35%, 0.3%, and 0.4%, respectively. Note that only limited amounts of gender-specific HDI data are available. The number of observations in the first regression is 1069, and in the second and third regressions, 394 and 393, respectively. The positive impact of remittance on HDI confirms previous findings in this paper. Workers' remittances do improve human capital in developing countries. Remittances and labor force participation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ This section examines the impact of workers' remittances on the labor force participation rate in the remittance-receiving countries. Remittances can affect labor force participation in remittance-receiving countries mainly through two channels: human capital and reservation wages. As stated by Amuedo-Dorantes and Pozo (2006) remittances and the error term in equation (1) may be correlated. Remittances may be endogenous and their coefficient estimates may be biased. There are three potential sources for this endogeneity. Omitted variable bias may exist if remittances are related to wealth which, in turn, may be correlated to the labor supply by the recipient. Additionally, there is the potential of reverse causality as labor supply may influence migrants' decisions to send remittances home. The third source of endogeneity can be measurement errors. Therefore, in this section, I use the same five instrumental variables to address the endogeneity of remittances. So far, we have concluded remittances increase human capital, especially health and education outcomes, in developing countries. On the one hand, increases in human capital can increase labor force participation. On the other hand, remittances, as a non-labor source of income, increase reservation wages in developing countries. Increases in the reservation wage can cause a reduction in the labor force participation rate. Therefore, the net impact of remittances on labor supply (labor force participation) is a priori unclear. The important question is: does the rise in the reservation wage and its negative impact on labor force participation outweigh the positive impact of remittances on labor force participation through raising human capital? If so, are the impacts of remittances on male and female labor force participation rates symmetric? The impacts of remittances on total, male, and female labor supply are provided in Table 16. In the IV model provided in Table 16, I use instruments to address the endogeneity of workers' remittances and control for the level of income, education, the ratio of the population living in urban areas, and geographic region. The results of the impact of remittances on labor force participation rate suggest that, on average, a 10% increase in remittances leads to a 0.17% decline in total labor force participation rate, a 0.03% decline in the male labor force participation rate, which is not statistically different than zero, and a 0.33% decline in the female labor force participation rate. These results show that while workers' remittances have no impact on the male labor force participation rate, they reduce the female labor force participation rate in developing countries. There is an enormous heterogeneity in labor force participation between men and women in developing countries. This heterogeneity is clearer when we review the descriptive statistics provided in Table 2. While average labor force participation for men is 76.6% in developing countries, it is just 50.9% for women. Due to cultural features of many developing countries, reservation wage for women is more sensitive to non-labor income than reservation wage for men. Therefore, upon receiving a greater amount of remittances some working-women will stop working. This is not the case for men in developing countries.","Remittances can increase investment in human capital. This impact is especially important since human capital is a crucial factor in the promotion of growth in developing countries in the long-run. This paper uses a dataset on workers' remittances, health outcomes, and educational attainment of 122 developing countries from 1990 to 2015 and an innovative instrumental variables estimation strategy to examine the impact of workers' remittances on human capital and labor supply in the developing world. Also, the gender-specific effects of remittances on health, education, HDI, and labor force participation are investigated. The results obtained in this paper show that remittances have an improving and statistically significant impact on health outcomes in developing countries. I use instruments to address the possible endogeneity of workers' remittances and control for level of income, some covariates, and geographic region. The results of the impact of remittances on health outcomes suggest that, on average, a 10% increase in per capita remittances will lead to a 1.5% increase in out-of-pocket per capita health expenditures, a 1.1% increase in total per capita health expenditures, 0.3% increase in life expectancy at birth, 1.5% decline in undernourishment prevalence, 1.9% decline in depth of food deficit, and 1% decline in prevalence of stunting. Also, remittances have a strong and statistically significant effect on reducing child mortality rates in developing countries. On average, a 10% increase in per capita remittances will lead to a 1% decline in neonatal mortality rate, 1.7% decline in infant mortality rate, and 1.9% decline in under-five mortality rate. The gender-specific results show that the effects of workers' remittances on reducing infant and under-five mortality rates are almost the same for male and female children. Remittances also have a positive and statistically significant impact on education in the developing countries. I use instrumental variables to address endogeneity of workers' remittances and control for the level of income, government expenditures on education, labor force participation, and geographic region. The results of the impact of remittances on education outcomes suggest that, on average, a 10% increase in per capita remittances will lead to a 3.5% increase in pre-primary enrollment rate, 0.7% increase in secondary enrollment rate, and 1.1% increase in tertiary enrollment rate. Remittances not only have a positive effect on enrollment rates but also have a positive effect on quality of education: on average, a 10% increase in per capita remittances will lead to a 2.3% increase in enrollment in private primary schools. In addition to raising the enrollment rate, remittances increase school completion rates: on average, a 10% increase in per capita remittances will lead to a 0.6% increase in primary completion rates and 0.9% increase in secondary completion rates. Unlike the impacts of remittances on reducing child mortality rates which were symmetric between genders, remittances improve girls' education outcomes more than boys. On average, a 10% increase in per capita remittances will lead to a 0.2% increase in the primary gender parity index, the ratio of girls to boys enrolled at the primary level. On average, a 10% increase in per capita remittances will lead to a 0.3% increase in girls' primary enrollment rate while it has no statistically significant effect on boys' primary enrollment rate. These results are in compliance with the positive effect of remittances on primary gender parity. Also, on average, a 10% increase in per capita remittances will lead to a 1.3% increase in girls' tertiary enrollment rate while it has no statistically significant effect on boys' tertiary enrollment rate. The positive effect of remittances on girls' education is greater than the positive effect of remittances on boys' education in all case of school completion rates: 10% increase in per capita remittances rises male and female primary completion rate by 0.5% and 0.9%, respectively and rises male and female lower secondary completion rate by 0.8% and 1.2%, respectively. Since remittances improve health and education outcomes in developing countries, it is expected to have a positive impact on HDI as well. On average, a 10% increase in per capita remittances will lead to a 0.3% and 0.4% increase in male and female HDI, respectively. The net impact of remittances on labor force participation is a priori unclear. On the one hand, workers' remittances can increase human capital investments of poor households, hence increasing the labor supply. On the other hand, remittances can loosen budget constraints, raise reservation wages, and, through the income effects, reduce the labor force participation of remittance-receiving individuals. The results obtained in the last section of this paper show that remittances have a negative and statistically significant impact on labor force participation in developing countries. I use instrumental variables to address the possible endogeneity of workers' remittances and control for the level of income, education, the percentage of population living in urban area, and geographic region. The results of the impact of remittances on labor force participation suggest that, on average, a 10% increase in per capita remittances will lead to a 0.17% decrease in labor force participation. The results of gender-specific effects of remittances on labor force participation can be more informative. While remittances have no statistically significant effect on male labor force participation, they reduce female labor force participation: on average, a 10% increase in per capita remittances will lead to a 0.3% decline in female labor force participation. This study contributes to the literature in a number of ways. First, this paper uses data from 122 developing countries over the period from 1990 to 2015. I am not aware of any other studies which explore the impact of remittances on human capital and labor supply and uses such a large dataset. Second, this study uses an innovative new approach for building instruments to address the endogeneity of remittances. This study estimates bilateral remittances uses them to create weighted indicators of remittance-sending countries. These weighted indicators are used as instruments for remittance inflow to remittance-receiving countries. selecting economic variables from remittance-sending countries, rather than from remittance-receiving countries, as instruments is to ensure the validity of instruments. Using estimated bilateral remittances and to construct the instruments is to ensure that instruments are strong. This novel approach to instruments can be applied to address the endogeneity problem while exploring the impact of remittances on many socioeconomic variables such as education, health, labor supply, poverty, inequality, growth, and financial development. Moreover, this approach can be used in both country-level studies and household-level studies. Finally, this study, in addition to investigating the overall impact of remittances on human capital and labor supply, explores the impacts of remittances on male and female human capital and labor supply using a wide range of gender-specific variables. This allows researchers to compare the impacts of remittances by gender. In terms of policy recommendation, the high-income countries should take further steps to reduce the current transaction costs of remitting money to remittance-receiving countries. Based on Remittance Prices Worldwide (RPW) reports, the average global cost of remittances went down from 9.12% in the first quarter of 2012 to 7.52% in the third quarter of 2015. Unfortunately, since the third quarter of 2015, the cost of remittances has not been reduced any further. The global cost of remittances in the first quarter of 2017 was 7.45%, staying 4.45 percentage point above the 3% UN SDG goal. The high transaction costs act as a type of tax on emigrants from developing countries who are often poor and remit small amounts of money with each remittance transaction. Lowering the transaction costs of remittances would not only encourage a larger share of remittances to flow through formal financial channels but will also help to improve health and education outcomes in developing countries."],["The lack of housing in areas where young adults have greater opportunities to study and get work complicates young adults’ entry into the adulthood. Difficulties in accessing housing may therefore delay childbearing and may negatively have an effect on education opportunities. To increase housing accessibility, some municipalities have earmarked apartments for young adults. These “youth dwellings” are criticized for being small and not necessarily facilitating family formation and fertility, better suiting students’ needs. We have in this paper compared the long-term pattern of childbearing and education for young adults that entered their housing market through small cheap youth housing with those youngsters that received a rental apartment from the ordinary housing stock. To be able to draw the conclusion that differences in fertility and educational pattern between these two groups comes from the different housing situation and not from differences in in preferences when it comes to childbearing or individual prerequisites for higher education, we have used a geocoded data and information on the individual's family background as well as a matching technique to create a comparison group that are similar to the treatment group in several aspects. The present results indicate that building affordable housing that is small and space efficient is sufficient and positive if the aim is to promote higher education. Affordable housing is on the other hand not enough to promote childbearing, instead, it seems to inhibit childbearing until there is a possibility of moving on in the housing career. Our result also indicates that the next step need not necessarily be homeownership, as earlier research has indicated. Entering the housing market via youth housing and then being able to move on to rental accommodation in the ordinary housing market also seems to have a positive effect on overall childbearing, although moving to cooperative housing or owned housing has an even larger effect. --------------------------------------------------------------------------------","The increase in adult children living with their parents has raised important questions regarding household formation. Furthermore, the lack of housing in areas where young adults have greater opportunities to study and get work complicates young adults’ entry into the housing market. Earlier research has found that housing and childbearing are closely connected and difficulties accessing the housing market may possibly lead to delayed childbearing (see, e.g., Mulder, 2006, 2013; Pinnelli, 1995; Castiglioni & Dalla Zuanna, 1994; Krishnan & Krotki, 1993; Clark, 2012) and may negatively influence education opportunities (see, e.g., Cunningham, Harwood, & Hall, 2010; Dworsky, 2008; Garriss-Hardy & Vrooman, 2005; Crowley 2003; Conley, 2001; Rosenbaum, 1995). But is affordable housing sufficient to promote childbearing and education? Leaving the parental home occurs for various reasons, the role of available housing likely differs for each reason. Young adults who want to leave home for education have little latitude for postponement and are likely to move even if they have to accept substandard housing (Mulder, 2006, p. 406). Those who want to leave the parental home for household formation have more latitude to wait until they have found suitable or/and affordable housing. For example, it has been demonstrated that higher housing costs are associated with lower probabilities of leaving the parental home to live with a partner, an association not found for those leaving the parental home to live alone (Mulder & Clark, 2000). Family formation has been demonstrated to be connected to homeownership tenure. Studies have also found that the decisions to become a homeowner and have children are made simultaneously (Malmberg, 2010, Enström Öst, 2012a, Kulu & Steele, 2013). Furthermore, it has been demonstrated that the likelihood of having children is greater for homeowners and that the transition to first-time homeownership often occurs in anticipation of parenthood (Mulder & Wagner, 2001; Feijten & Mulder, 2002, Mulder, 2006, Kulu, 2008, Holland, 2012). However, having small amounts of equity in one´s home reduces the ability to realize a desired move (Ferreira, Gyourko, & Tracy, 2010). An affordable housing market may enable smooth entrance into the housing market, perhaps via a small cheap apartment, enabling later progress in the housing career to higher-quality housing (Mulder, 2006, 2013). If young people succeed in leaving the parental home and enter the housing market in a small dwelling, this may positively affect childbearing if subsequent access to high-quality housing is easy. However, if such housing is scarce, prices are high, and/or mortgage providers are strict, young people might postpone childbearing until they find a house suitable for family formation, which may reduce the number of children born (cf. Chiuri & Jappelli, 2003). Similarly, Vignoli, Rinesi, and Mussino (2013) find that women who feel more secure about their housing conditions are more likely to plan to have their first child. Simon and Tamura (2009) and Clark (2012) have explicitly investigated the effect of housing costs on childbearing. Both studies show that first birth is significantly delayed in an expensive housing market. Liu and Clark (2016) demonstrates that an increase in the cost of renting is predicted to decrease the number of children born by renting households. However, the effect of higher house price is ambiguous and depends on the initial holdings of housing and the willingness to substitute between children and other goods. The housing market in Sweden, especially in the capital Stockholm, is increasingly difficult for young adults to enter (Bokriskommittén, 2014; Swedish union of tenants, 2015). This market is characterized by increasing prices and housing costs and by housing construction that has lagged behind population growth (The Swedish National Board of Housing, Building and Planning, 2013). Queues to obtain rental apartments are growing and cooperative apartment prices are high. Studies indicate that a generation of young people may have little or no chance of accessing the housing market unless they are rich, well paid, or/and have generous and wealthy parents (cf. Enström Öst, 2012b). Fertility research commonly relates the relatively high Swedish fertility to the characteristics of the Swedish welfare regime that, for example, may promote female labour-market attachment by making it easier to combine work and family life (see, e.g. Andersson & Scott, 2007). However, along with indications of an inaccessible housing market for young people, the average age at which women bear their first child has increased in Sweden by approximately one year over a five-year period (cf. Andersson, 1999, 2000; Statistics Sweden, 2011). Recent years have also seen reports of students in higher education being forced out of larger cities, such as Stockholm, because of difficulties finding accommodation. To try to increase the accessibility of the housing market, housing projects targeting young adults have been started in Sweden. Several municipalities have earmarked small apartments for young adults to help them compete for an access to rental apartments. By examining a housing project for young adults in Stockholm, initiated as early as 1996, this study advances our understanding of whether building small apartments for young adults may solve problems related to housing shortage for young. This study will investigate the causal effects of youth housing on higher education and parenthood. To our knowledge, no earlier studies have this focus. The dataset for this study contains information on young people who gained access to the housing market in 1996 via a particular youth housing project in Stockholm. This study explores the development of their fertility pattern and education level during a time period of 14 years after they entered the housing market and compares them with the fertility and educational pattern of a matched group of young people similar in several respects. Matching is used to evaluate the effect of youth housing by comparing those young adults that moved into the youth house in 1996 with those young adults that moved into an ordinary rental apartment the same time period. The goal of this matching procedure is, for every young adult that moved into the youth house, to find at least one young adult with similar observable characteristics against whom the effect of youth housing can be assessed. With the data available for this project we can match on explanatory variables five years before the youth house was defined. By this matching procedure, it enables a comparison of outcomes among young adults moving into youth housing and young adults that moves into the ordinary rental housing stock to estimate the effect of youth housing reducing bias due to confounding. The present result indicates that having access to small youth housing reduces the probability of becoming a parent. However, if the rest of the housing market is mobile, i.e., enables young adults to move on in their housing careers, the total effect on becoming parent is positive. The effect from living in youth house on completing higher education is positive and independent of the mobility of the rest of the housing market. This result confirms that the housing market has repercussions for both the fertility and the educational patterns. The paper is organized as follows: The next section discusses the theory of household formation and the Swedish housing market during the study period and describes the studied youth housing case. Section 3 presents the data and the empirical strategy and Section 4 presents the results. The paper ends by presenting the conclusions and discussing the policy implications of the present findings. Theory of housing and household formation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A large body of empirical work emphasizing the role of life course events on mobility, such as leaving the parental home and get birth to a child (Clark & Dieleman, 1996). Transition to life stages involving higher levels of commitment, such as parenthood, lead to requirements for long-stay housing and to changing preferences over dwelling attributes. However, desires for mobility may be prompted by a wide range of life events, but desires cannot always be realized. Income and wealth constraints, transactions costs that vary with tenure, social ties, the supply of dwellings and the functioning of the housing and mortgage markets all affect whether a desired move will in fact be realized (Kiel, 1994; Linneman & Wachter, 1989; Stein, 1995; Venti & Wise, 1984; Wheaton, 1990; Helderman et al. 2004; Belot & Ermisch, 2009; Ermisch & Washbrook, 2012). Mulder (2006) has explored the relationship between housing and household formation. Obviously, to move from parental homes and form households, young people need somewhere to live. In a well- functioning housing market, housing demand equals housing supply and all households can access housing that meets their needs. However, housing prices, housing supply, and the ability to obtain housing loans are limiting factors, especially for young adults who have had limited time to accumulate savings for home down payments. Some young adults may therefore postpone household formation if they cannot find suitable affordable housing. The degree to which the availability of housing affects household formation, however, probably depends on the urgency with which people want to form new households. Young adults who need to move, for example, for work or higher education, might have little latitude for postponement and may therefore move even though they must accept substandard housing. However, those moving to cohabit, marry, or have children may have more latitude to wait until they find suitable housing. Furthermore, young people who succeed in leaving parental homes, for example, to live in student housing or smaller apartments, might still postpone childbearing if they think housing of a certain quality is a precondition (see, e.g., Ineichen, 1979, 1981; Ström, 2010). Lovenheim and Mumford (2013) also find a statistically significant positive relationship between house price changes and fertility for homeowners. However, no statistically significant negative relationship between house price changes and fertility was found for renters. Dettling and Kearney (2014) find that high housing prices have a negative effect on the fertility of renters but a positive effect for home owners. The Swedish housing market ~~~~~~~~~~~~~~~~~~~~~~~~~~ The Swedish housing market has three dominant tenure forms: single-family housing, cooperative multi-family housing, and multi-family rental housing. All rental housing units are subject to rent control. In the case of cooperative housing, the property is owned by a cooperative association. Each resident owns a share of the cooperative and occupies an apartment with tenancy rights nearly as strong as those of full ownership. Cooperative housing is traded on an ordinary free housing market and in practice is regarded as a form of owner-occupied housing, although this is not precisely correct from a judicial perspective. The standard of housing is generally high in Sweden, irrespective of tenure type. The housing market in Sweden has undergone several gradual and substantial changes in recent decades. The government housing policy of granting substantial subsidies to all new housing construction changed in the early 1990s, and subsidies were gradually phased out over a decade without being replaced with other investment incentives. This has resulted in the low production of new housing of all types, and housing construction has lagged population growth since 1991, as well as very high house prices and housing costs. Swedish house prices increased by over 200% in the 17 years from 1996 to 2013, a period when general consumer prices increased by only 60%. Furthermore, the price level of houses in the first quarter of 2015 was 12 percent higher than in the first quarter of 2014. However, this trend is even stronger in the market for cooperative apartments where the prices was 17 percent higher during the first quarter of 2015 than in the corresponding quarter of 2014 (The Swedish National Board of Housing, Building and Planning, 2015). After the 1991 tax system reform, which led to sharply increased rents (Englund, Hendershott, & Turner, 1995), the vacancy rate in the rental sector started to climb and was quite high by 1998. Since then, vacancy rates have decreased and today several counties in Sweden report a housing supply shortage (The Swedish National Board of Housing, Building and Planning, 2015). In 2015, only 40 percent of Sweden’s young adults lived in a home of their own. This is the lowest percentage ever measured in Sweden (Swedish union of tenants, 2015). Furthermore, in 2015, 20% of young adults who had children lived with a relative, in a rented room or in a student residence. Stockholm is the largest and most dynamic regional housing submarket in Sweden; it is Sweden’s most heterogeneous in terms of tenure forms, price variation, and neighbourhood structure, though, from an international perspective, it is comparatively small, homogenous, and easy for housing consumers to conceive of and analyse. The official recommendation to those who want a rental apartment in inner Stockholm is to register as an applicant and stay on a waiting list. Over the last two decades, in-migration to Stockholm has increased substantially. With population growth substantially exceeding housing supply growth, the Stockholm housing market is now suffering from a pronounced housing shortage (Andersson & Söderberg, 2012). About 520,000 people were waiting for an apartment in Stockholm County in January 2016 and the average waiting time for a rental apartment in the region exceeds eight years. When it comes to student housing, approximately 80,000 students are waiting for only 12,000 student dwellings.1 “Youth house” ~~~~~~~~~~~~~ In 1995 a housing company in Stockholm decided to earmark an entire 150-apartment property in inner Stockholm for young adults aged 18–25 years. These apartments had formerly been earmarked for nurses working at the nearby hospital. The apartments were all small, i.e., 30 square metres in floor area comprising one room with a kitchenette, and were deemed fit for youth by the housing company. Earmarking these apartments for young adults therefore required no major or costly renovations by the housing company. The property, located in inner Stockholm and built in the 1950s, has been subject to no default-enhancing renovations since completion. To obtain one of the units, which are all rental apartments, young adults must apply at the housing company or at the housing service that allocates vacant rental apartments in Stockholm. Applicants aged 18–25 years who register interest and wait the longest will receive a vacant apartment. The rents for these youth apartments could be considered quite low, especially relative to the costs of the housing alternatives available to these young adults. Most rental apartments in that same area but in the ordinary housing stock, i.e., not earmarked apartments, have undergone default- enhancing renovations with the result that their rents have increased significantly. The waiting time for such apartments is over 15 years, indicating that this is an attractive area to which young adults without the possibility of obtaining an earmarked apartment would normally have difficulties gaining access, unless they can buy an apartment, which is very expensive. A one room rental apartment in the ordinary housing stock is also on average 10 square metres larger than the apartments in this Youth House. Data and sample ~~~~~~~~~~~~~~~ The dataset is extracted from a database provided by Statistics Sweden. This database contains information on all individuals who have resided in Sweden, including their demographic and socioeconomic situation as well as geocoded data with coordinates and neighbourhood area codes for where the individuals live. Using this database, it is possible to link records between individuals and generations, because the data include a household identity code and, for every individual born after 1932, a specific identity code for the individual’s mother and father. The sample used in this study consists of young adults aged 18–25 years who moved into the property earmarked for young adults (“Youth House”) in Stockholm in 1996. These individuals were defined by identifying the geographic coordinates of the property, which is located near a large hospital and is surrounded by a green area. No other residential properties are located immediately adjacent to the Youth House. Since we in our data also have the geographical coordinates for everyone, we have been able to define the young adults living a maximum of 30 m from the property geographical coordinates in 1996, i.e., those 112 young adults aged 18–25 years who moved into Youth House in 1996. This group of young adults will constitute the treatment group of this study. We will follow them in our data until they have a first child, complete their higher education, or fourteen years have elapsed (i.e., until 2010). Their childbearing and education patterns will be compared to those of a matched group of young adults, i.e., a comparison group of young adults similar to members of the treatment group in several respects but who did not move into youth housing. The matching procedure for defining the comparison group of young adults is explained in the next section. Propensity score and matching ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We are in this paper interested in comparing the long-term pattern of childbearing and education for young adults that entered their housing market through small cheap youth housing with those who did not and instead received a rental apartment from the ordinary housing stock. However, those individuals that choose to live in small cheap youth housing may differ from those that prefer other rental apartments in terms of different attitudes towards education and family formation, for example. A simple comparison of their education and childbearing pattern would then be biased by these confounding variables. This confounding by indication is almost invariably present in non-randomised studies. Rosenbaum and Rubin (1983) proposed the use of propensity scores as a method for allowing for confounding by indication. Propensity may be defined as an individual’s probability of being treated with the intervention of, in this case, entering the housing market through youth housing, given the information available about that individual. The propensity score provides a single metric that summarises all the information from the explanatory variables. Individual subjects may have the same or similar propensity scores, yet some will start their housing career in youth housing and others will not. An assumption of propensity score analysis is that a fair comparison of treatment outcomes can be made between subjects with similar propensity scores who either did or did not start their housing career in youth housing. The propensity score may be estimated for each subject from parametric regression in which the treatment is the dependent variable. A feature of this approach is that explanatory variables are selected based on their ability to predict exposure to the intervention of interest. Matching is a popular approach employed to include propensity scores in the analysis, and empirical examples can be found in very diverse fields of study. Matching requires that each treated individual is matched with one or several individuals having the same or similar propensity score. By contrasting the outcomes between treated and untreated sets of individuals with similar propensity for treatment, we can estimate the average treatment effect. Strategy used To create a balanced (see, e.g., Rosenbaum and Rubin 1983) comparison group of young adults in this study, we use the following strategy. First, we identify where the individuals who moved into the Youth House in 1996 grow up, i.e. we identify the geographical coordinates for the individuals’ parental homes five years before they moved into the Youth House, i.e. in 1991. We then identify all young adults under 25 years old, with no children and with no higher education, who lives within 100 m of each individual in the treatment group in 1991. The propensity scores, i.e. the probability of entering the housing market through Youth housing, are estimated for all these youngsters as well as for the treated. The propensity score has been estimated using a Probit model in which the dependent variable is the probability of moving in to the Youth House in 1996.2 The covariates included in the model are sex, age, and final primary school grade as well as family income, parental homeownership status, and whether the parents live in single- or multifamily housing. All covariates refer to the situation in year 1991, five years before the Youth House was defined. After having estimated the propensity scores for all youngsters, we perform a 10 nearest-neighbor matching to identify those young adults under 25 years old with no children and no higher education living close to each individual in the treatment group with a similar propensity score and five years before the Youth House was created. Nearest-neighbour matching identifies a control case with a propensity score closest to that of each treatment case (see, e.g., Gu & Rosenbaum, 1993). Here, we match with replacement, meaning that after a control case is used as a match, it is put back into the sample, and can be used again to match other treated units as well. After this matching procedure, we have tested whether the treated and control groups are balanced with respect to the observed characteristics. The result is presented in Table 1. Table 1 show the means of the variables for the treatment and matched control groups and our balancing test of the explanatory variables cannot reject equality of means between the treated and the matched comparisons; the t-values indicate no significant differences between these groups, indicating the presence of assessed balance. With 112 individuals in the treatment group and 985 in the comparison group, 170 of the comparisons appear more than once in the matched data. For a sample to be considered sufficiently balanced, Rubin (2001) recommends that Rubin’s B, i.e., the absolute standardized difference of the mean linear index, of the propensity score between the treated and (matched) control groups, be less than 25 and that Rubin’s R, i.e., the ratio of the variance of the propensity score index, between the treated and (matched) comparison groups, be 0.5–2.0. Both those recommendations are fulfilled, so the matching procedure is deemed satisfactory. Empirical strategy ~~~~~~~~~~~~~~~~~~ We will also include variables indicating when an individual moved from Youth House and to which tenure type he or she moved, i.e., other rental apartment, cooperative apartment, or owned housing. 78 percent of the individuals in the treatment group moved from the Youth House during the observation period. About 40 percent of those moved to cooperative or owned housing. Note that, after matching a sample, one can simply use the difference between the treatment and control groups to estimate the average treatment effect. However, even if we have a matched comparison group, we still want to check whether other variables not included in the matching procedure may, over time, affect the childbearing decision or the decision to become educated, differently between the treated and the compared subjects. These variables are welfare benefits, income from work, and income from capital. However, including these covariates did not alter the result for the treatment variable or the conclusions of this study. Young adults who remain childless or do not complete higher education are censored at the end of the observation period, i.e., in 2010. When modelling a Cox proportional hazard model, a key assumption is the proportional hazards. Accordingly, we will perform some diagnostic tests for non-proportionality and estimate parametric models, i.e., the exponential and Weibull models. Becoming parent ~~~~~~~~~~~~~~~ Here we report the results of the Cox proportional hazard models. The tables present the hazard rates rather than the coefficients themselves. A hazard rate above 1 indicates increased risk and a hazard rate less than 1 decreased risk. The variable of main interest in all tables is Youth House, a dummy variable indicating whether the individual moved into Youth House in 1996, i.e., whether the individual belongs to the treatment group. Other variables of interest are, as earlier mentioned, a dummy variable indicating the time point when an individual first moved from Youth House (i.e., Youth House_moved) as well as dummy variables indicating the type of tenure to which the individual first moved (i.e., Youth House_moved to rental = moved to another rental apartment, Youth House_moved to coop = moved to a cooperative, and Youth House_moved to own = moved to owned housing). Table 2 presents the results of the model with becoming parent as the outcome variable. Models 1 and 2 present the model results including a variable indicating the time the individuals move from Youth House. The results indicate that living in Youth House has a large significant and negative effect on childbearing comparing to those that instead moved into the ordinary housing stock, a result that is stable even when including covariates in the model. The hazard of becoming parent while in Youth House is 0.12. However, moving from Youth House greatly increases the hazard (a hazard of about 17) of becoming a parent, indicating the importance of a mobile housing market. This result remains when controlling for the type of tenure to which subjects move when they leave the Youth House (Models 3 and 4). The hazard of becoming parent are, as expected, largest for ownership. However, it is also surprisingly higher for those that move to cooperative apartments than for those that move into a rental apartment in the ordinary housing stock, a hazard of about 26 compared to 15 for rentals. The results are stable though the sample is quite small and even when including covariates in the models. Higher education ~~~~~~~~~~~~~~~~ Table 3 presents the results of the education model, i.e., the outcome of completing higher education. Here, unlike in the childbearing case, we observe a significantly increased risk of completing higher education for the treatment group, i.e., a positive effect of Youth House on higher education, corresponding to a hazard of approximately 2.6. However, the variable indicating a move from Youth House was not significant in this case, indicating that a mobile housing market seems unimportant to the promotion of higher education. This result remains when controlling for the type of tenure to which subjects move (not presented here). None of those variables were significant. Sensitivity tests ~~~~~~~~~~~~~~~~~ All the models were also estimated using the exponential and Weibull models. The Weibull model is more general and flexible than is the exponential model and allows for hazard rates that are non-constant but monotonic. The results of both these models were in line with those of the Cox model presented in Sections 4.1 and 4.2. As already mentioned, a key assumption when testing such models concerns the proportional hazards. We accordingly performed the Schoenfeld residuals test, developed by Therneau and Grambsch (2000), to test the proportional hazard assumption. We found no evidence that the models with higher education as the outcome violated the proportional hazard assumption, nor that the models with the outcome childbearing that did not include the variable indicating a move from Youth House violated the same assumption. For the other models, the results of the Schoenfeld residual test indicated non-proportional hazards. One solution to this potential problem is to interact the variables displaying signs of non-proportional hazards with the natural log of time (Box-Steffensmeier & Zorn, 2001, p. 978) to explicitly allow the effect of the variable to vary across time. We did this for all variables that displayed signs of non-proportional hazards; this, however, did not alter the results. All the variables of interest were still highly significant, with the same estimate size as with the Cox model.","The point of departure of this study was to examine the effect of youth dwellings, i.e., apartments earmarked for young adults, on young adults’ childbearing and education patterns. Youth dwellings are often small and lack proper kitchen facilities, and therefore do not necessarily facilitate family formation. If access to higher-quality housing is difficult because of scarcity, high prices, and strict mortgage provision, young people might postpone childbearing if housing of a certain standard is desirable when forming a family. However, young adults who leave the parental home for education may have little latitude for postponement and may therefore obtain an education despite living in substandard housing. We have in this paper compared the long-term pattern of childbearing and education for young adults that entered their housing market through small cheap youth housing with those youngsters that received a rental apartment from the ordinary housing stock. To be able to draw the conclusion that differences in fertility and educational pattern between these two groups comes from the different housing situation and not from differences in preferences when it comes to childbearing or individual prerequisites for higher education, we have used a geocoded data and information on the individual’s family background as well as a matching technique to create a comparison group. An assumption of propensity score analysis is that a fair comparison of treatment outcomes can then be made between subjects with similar propensity scores who either did or did not start their housing career in youth housing. The present results indicate that gaining access to a small, low-rent inner-city apartment earmarked for young adults may promote higher education but negatively affect childbearing unless the rest of the housing market is mobile, i.e., enables young adults to advance in their housing careers. However, the effect of such housing on higher education is independent of mobility in the rest of the housing market. So, what are the policy implications of our finding? The key policy implications of these results are that: Building affordable housing that is small and space efficient is sufficient and positive if the aim is to promote higher education. Affordable housing is on the other hand not enough to promote childbearing, instead, it seems to inhibit childbearing until there is a possibility of moving on in the housing career. Our result indicates also that the next step need not necessarily be homeownership, as earlier research has indicated. Entering the housing market via youth housing and then being able to move on to rental accommodation in the ordinary housing market also seems to have a positive effect on overall childbearing, although moving to cooperative housing or owned housing has an even larger effect. These results may lead to a deeper understanding how policy makers may promote childbearing and higher education through the housing market, something that seems to be a necessity to attract and retain younger residents. Small apartments for younger people in an area where most elderly or families with children live may also create a good dynamic and a variety of people. It also creates a broader service and store range and a good basis for public transport. However, the present results do not allow us to determine the driving force of the results achieved by youth housing. Is it that young adults do not find the general concept of youth housing compatible with childbearing, or is it that the dwellings are small and lack proper kitchen facilities, i.e., the housing size and standard do not meet the quality norms required for childbearing? Or does the relatively low rent make alternative accommodations seem too expensive, creating a lock-in effect that postpones childbearing? Answering these questions requires more data and further research into youth housing."],["We develop a location analysis spatial model of firms’ competition in multi-characteristics space, where consumers’ opinions about the firms’ products are distributed on multilayered networks. Firms do not compete on price but only on location upon the products’ multi-characteristics space, and they aim to attract the maximum number of consumers. Boundedly rational consumers have distinct ideal points/tastes over the possible available firm locations but, crucially, they are affected by the opinions of their neighbors. Proposing a dynamic agent-based analysis on firms’ location choice we characterize multi-dimensional product differentiation competition as adaptive learning by firms’ managers and we argue that such a complex systems approach advances the analysis in alternative ways, beyond game-theoretic calculations. --------------------------------------------------------------------------------","Product characteristics are typically considered as given when economists study firms’ strategies and behaviour. But firms in industries with product differentiation actually choose the features of their products based on consumers’ preferences. The role and effects of the demand side in product differentiation processes have been neglected by the relevant literature (Coombs et al., 2001; Mueller et al., 2015) and only recently have started to receive attention (Andersen et al., 2011). The importance of consumer behaviour in shaping markets’ properties is oversighted and in most cases the simplistic perspective of perfect rational homogeneous agents is implicitly adopted in the literature (Valente, 2012). Among the works highlighting the relevance of demand for the emergence of new products, Witt (2001), Saviotti (2001), Malerba et al. (2007), Windrum et al. (2009), Nelson and Consoli (2010), Valente (2012), Markey-Towler and Foster (2013), Mueller et al. (2015), Markey-Towler (2016) and Schlaile et al. (2017) can be included. However, no attention has been paid to advancing a generalized evolutionary model of demand for the effect of consumers’ social networks on their consumption decisions and, in turn, on firms’ process of developing new characteristics for their products. The aim of this paper is to address this gap in the literature by suggesting a way to understand the relevance of consumers’ social networks in shaping demand behaviour in such markets, which in turn frame the process of product-embodied innovations. We do so using complex networks, which mathematically are described by graph objects. We develop a stylized computational model to examine the issue of product differentiation in a multi-dimensional space with consumers’ choices emerging from consumers’ opinion-based multilayered networks. The model that we propose is designed to analyze and illustrate potential effects of consumers’ heterogeneity, particularly regarding their preferences, their bounded rationality and the role of their social networks on the development of new product characteristics. We ask two questions: First, how consumers’ social networks affect their decisions on which products to buy and how this decision leads to market-share inequality at the firm level? Second, how firms respond to these decisions by developing new characteristics for their products defined over a multidimensional characteristic space? Assuming that bounded rational consumers with heterogeneous preferences can have an active role in the product- embodied innovation process, and studying the influence of a change in their purchasing opinion over time due to their participation on social networks calls for new modelling approaches, namely complex systems analysis and agent-based simulation techniques (Pyka and Grebel, 2006; Pyka and Fagiolo, 2007; Gilbert, 2008; Farmer and Foley, 2009; Tubaro, 2011; Schlaile et al., 2017; Muller, 2017). Agent-based simulation approaches focus on the rules and elements that constitute a system and the interactions of the players/agents within the system and have gained increasing momentum in many scientific disciplines (Schlaile et al., 2017; Mueller and Pyka, 2017; Vermeulen and Pyka, 2016; Namatame and Chen, 2016; Hamill and Gilbert, 2016; Boero et al., 2015; Wilensky and Rand, 2015; de Marchi and Page, 2014; Kiesling et al., 2012). In our case, through agent-based simulation we are able to account for the economic agents’ heterogeneity and the related implications of their interactions by representing the economic system in a more realistic fashion overcoming the simplistic models limited to representative agents (Bonabeau, 2002; Macal and North, 2005, 2006, 2010, 2014; Page, 2011; Schalaile et al., 2017). The agent-based simulation model described in this paper is based on the works by Laver, (2005), Laver and Sergenti (2011), Valente (2012), Markey-Towler and Foster (2013), Halu et al. (2013), Mueller et al. (2015), Schlaile et al. (2017), Muller (2017) and represents a first attempt to address the issue of neglected consumers’ social networks in the process of developing new products. In particular, we focus on the effects of heterogeneous and boundedly rational consumers’ social networks on firms’ multi-dimensional product differentiation competition.1 Our model captures only parts of the consumers’ demand-side, namely ‘consumption as voting’ (Dickinson and Carsky, 2005; Shaw et al., 2006; Moraes et al., 2011) whereby consumers choose which firms/suppliers and new products/services to support.2 In this sense, our model could also apply to political competition, as studied by Downs (1957) and many others after him. More specifically, it could be perfectly placed within the scope of agent-based models described in Laver and Sergenti (2011). However, in this work we will stick to the firm affairs’ language. In more detail, what we propose here is a multi-dimensional agent-based model of firms’ locational choice in the product- characteristics space that describes a finite number of firms competing for customers.3 This model is inspired by the New Consumer Theory literature dealing with the non-price aspects of consumer behaviour (Lancaster, 1966, 1971, 1975, 1979; Ironmonger, 1972; Ratchford, 1975; Earl, 1983, 1986) and has its roots in the behavioral economics literature having to do with boundedly rational consumers making decisions under the presence of multiple attributes (Selten, 1998). From consumers’ perspective, we begin with a very basic idea. Consumers’ decision on purchases cannot be different from most other decisions people make in their daily lives, in the sense that some process for acquiring information and evaluating it is necessary. The model also allows for heterogeneity of demand based on heterogeneous consumers having individual preferences for the particular characteristics of the products, like the colour, the size, the brand, the shape, their functionality etc.: consumers have preferences that place them on their ideal points in a multi-characteristics space.4 For every firm we model the network of consumers’ opinions about its product (i.e. for every firm there is a corresponding opinion network describing consumers’ opinion about the firm’s product: the nodes represent the consumers and a link between two nodes represents their exchange of opinion about the corresponding firm’s product), on which (opinion network) a contagion dynamics can take place.5 Consumers are represented on every network and can be active only in one of the networks (purchase of one product from one firm) at the end of a given time period (at the moment of the purchase). Each consumer has also the option not to purchase, and in that case she will be inactive in all networks. Before the end of each time period, and while the consumers think about the purchase, they are more likely to change opinions and examine different options. They have some form of product characteristics expectations at the beginning, which is reflected by their position in the product characteristics space, but their decision on which product to purchase is crucially affected by the opinion of their peers: at the time of the purchase, they tend to be active in the network where the majority of their peers are also active.6 In addition, as they gain experience and acquire more information about the products, their opinion can change, but it stabilizes over time (see Hoeffler and Ariely, 1999; West et al., 1996). This “uncertainty reduction”, reflecting both consumers’ learning and changes in their environment, is captured by dynamics that slow down until the purchase moment when they become completely frozen. These dynamics are implemented in this work with the simulated annealing algorithm.7 An illustration of a possible final configuration of the described system is shown in Fig. 1. At the end of the simulated annealing calculation we have the number of consumers that opt to purchase from each firm. This is affected by the average connectivities of the consumers’ opinion networks. Hence, we observe “regions” in the average connectivity space where the active nodes belonging to some opinion networks with high average connectivity percolate the system, while nodes of the remaining opinion networks with lower average connectivities are concentrated in disconnected clusters indicating high market share inequality. Nevertheless, “regions” where the average connectivities of the consumers’ opinion networks are comparable manifest a market that sustains low market share inequality (see also Halu et al., 2013). To the best of our knowledge, this paper is the first study that incorporates the role of opinion exchange and processing on a multilayered network, in an attempt to shed light on how the market-share inequality and the competition in a multi- characteristics space are affected by the presence of consumers’ interactions.8 Taking into account these multiple layers is crucial, as the considerable interest in various multilayered systems demonstrates (Buldyrev et al., 2010; Vespignani, 2010; Parshani et al., 2011; Baxter et al., 2012; Bashan et al., 2012; Gomez et al., 2013; Radicchi and Arenas, 2013; De Domenico et al., 2013; Radicchi, 2014; Garas, 2016). On the other hand, to make significant advances in understanding consumers’ decisions, we must device a method for studying this process. Such method would allow us to observe consumers’ behaviour from up close, to dig below the surface and watch consumers as they try to exchange information about a myriad of alternative products while refining the overwhelmingly volume of information of their multiple characteristics. Behavioral decision theory guides the process-oriented, complex systems-framework we present in the next sections in an effort to develop a new set of measures for studying market proceedings. The results reported in the subsequent sections are intended to demonstrate that such a realistic complex systems approach provides a plausible basis for understanding market-share inequality and firms’ competition in the product- characteristics space and can advance the analysis of these topics in interesting and alternative ways, beyond game-theoretic calculations. The rest of the paper is organized as follows. In Section 2 the agent-based model (ABM) is introduced, section 3 gives the simulation strategy and discusses the simulation results and Section 4 concludes. Consumers’ behaviour ~~~~~~~~~~~~~~~~~~~~ Traditionally the assumed heuristic for consumers is a maximization rule over some utility function defined over the set of products’ quantities. However, this rule is inconsistent with extensive evidence presented in behavioral economics and marketing literature. Here we assume that each consumer’s preferences can be characterized by an ideal economic position in some k-dimensional product characteristics space, and we look closely on the decisions made by consumers when choosing which firm’s product to buy. Having switched from formal analysis to computation, we depart from the classical analytically tractable models first, by assuming as baseline decision rule that consumers are affected by the opinion of their peers and second, by making the appropriate behavioral assumptions following the relevant literature of behavioral economics, discussed below. Behavioral decision theory is psychological in its orientation, beginning with the view of humans as limited information processors or, perhaps more accurately, as “boundedly rational information processors” (Simon, 1955, 1956, 1957, 1959). Humans have developed a large number of cognitive mechanisms to cope with the overwhelming volume of information in the modern societies. These mechanisms are adopted automatically without any conscious and are cognitive shortcuts for making certain judgments and inferences with considerably less alternatives than those dictated by rational choice, focusing attention on a small subset of all possible information. Kahneman and Tversky (1973, 1974, 1984) and Tversky and Kahneman (1973, 1974) have identified three general cognitive heuristics that decision makers adopt in the process of information gathering and analysis: (a) decomposition, which refers to braking a decision down into its component parts, each of which is presumably easier to evaluate than the entire decision; (b) editing or pruning, which refers to simplifying a decision by eliminating (ignoring) otherwise relevant aspects of the decision; (c) decision heuristics, which are simplifying the choice between alternatives thus providing cognitive efficiency. These heuristics have direct application to consumers’ choices. Consumers face a myriad of alternative products and there is compelling evidence which suggest that consumers simplify their decisions with a consider- then choose decision process in which they first identify a set of products, the consideration set, for further evaluation and then choose from the consideration set. In seminal observational research Payne (1976) identified that consumers use consider-then- choose decision processes. This heuristic is firmly rooted in both the experimental and prescriptive marketing literature (e.g. Bronnenberg and Vanhonacker, 1996; Brown and Wildt, 1992; DeSarbo et al., 1996; Hauser and Wernerfelt, 1990; Jedidi et al., 1996; Mehta et al., 2003; Montgomery and Svenson, 1976; Paulssen and Bagozzi, 2005; Roberts and Lattin, 1991; Shocker et al., 1991; Wu and Rangaswamy, 2003). Evolution dynamics during consumers’ purchasing process ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Initially, active nodes in each network j are distributed according to consumers’ consideration sets [Eq. (3)], and we allow the consumers to change opinion at any time before they make their final choice. These dynamics are implemented using the simulated annealing algorithm, which works as follows. Starting from a relatively high initial temperature, T,11 i.e. a large number of potential conflicts [a large-number of non- binding constraints given by Eq. (4), hence a high H given by Eq. (5), which means that before the purchase time, consumers can be active in more than one network, i.e. they can have positive opinion for more than one firm/product] due to the stochastic way initial opinions are distributed, we use a Monte Carlo dynamics which will reach an equilibrium following the equation (5). As the time for the final choice approaches, i.e. as the simulated annealing progresses, the effective temperature T decreases and the consumers tend to have less and less opinion conflicts with their network peers until they reach to zero conflicts at equilibrium, when H = 0. Firms’ ‘behaviour’ ~~~~~~~~~~~~~~~~~~ Following Laver and Sergenti (2011) we assume that firms’ managers use an adaptive decision rule to set product characteristics on a multidimensional product-characteristics space at any given period t. We assume an adaptive rule that models a manager who constantly modifies product characteristics in the search for more customers, and the manager cares only about the firm’s market share. If the manager’s decision for the characteristics of the product at time t was rewarded by an increase in market share, then the firm makes a unit move at time t + 1 in the same direction as the move at t. If not, the manager reverses direction and makes a unit move on a heading randomly selected within the half-space being faced now. In other words, the firm’s manager relentlessly forages in the product-characteristics space, always searching for more customers and never being satisfied, changing strategic planning in the same direction as long as this is rewarded with more sales, but casting around for a new strategic planning when the previous one was punished with falling or static sales. Now that we have structurally defined the evolution process of the complex system as a whole, we can proceed to simulations.","For clarity in what follows we consider two firms, hence two layers in the consumers’ opinion network, in the simulation experiments. Though, our code can be straightforwardly generalized to consider multiple firms/networks. In our setting, consumers are assumed to be intrinsically interested in product characteristics and to have ideal points in a product-characteristics space (again, in order to aid visualization, the simulated version of the model is implemented in two dimensions, i.e. in the Euclidean plane, but the model can be implemented in any number of dimensions). Firms compete with each other by offering products with varied characteristics to consumers. Fig. 2 summarizes the model and its evolution process, which was programmed in R. Firm system dynamics ~~~~~~~~~~~~~~~~~~~~ Initiation of the model randomly distributes a discrete set of firms’ locations and consumers’ ideal points across the product-characteristics plane.13 Consumers are initially present and active in the opinion-networks/firms which lie within their consideration sets. As discussed above (see Section 2.3), we allow the consumers to change opinion at any time before they make their final choice, and we model their dynamics using the simulated annealing aiming to reduce to zero the number of conflicts with their network peers (the number of violated mathematical constraints given by Eq. (4) for all consumers) when reaching equilibrium. More precisely, following Halu et al. (2013), the algorithm starts with an initial temperature T = 1 and at every time step we select a node at random from either one of the two networks with equal probability and we change it from active to inactive or vice versa. After this change, the Eq. (5) is recalculated, and if the difference with respect to its previous value, ΔH, is negative i.e., the number of conflicts in the system is reduced, the change is accepted. If ΔH > 0, the change can still be accepted but with a small probability given by e−ΔΗ/kT. This random selection process is repeated 2N time steps, in order to update the whole system on average once and then we advance the system time by one Monte Carlo step. The whole process is repeated by slowly reducing the temperature until we reach equilibrium at H = 0, where there are no more conflicts in the network. The final configuration of the “uncertainty reduction” process just described is depicted in Fig. 1. The simulation results visualized in Fig. 3 are in accordance with the findings of Halu et al. (2013) and imply that the firm with the most connected customers gets the lion’s share of the market: Case 1: In the regions where the mean degree of both networks is smaller than one, there are no giant components in the networks. This means that the networks are fragmented and the opinion of a consumer is not affected by peer opinions. In this case each firm possesses only a marginal market share and noticeable market share inequality is unlikely. Case 2: In the regions where the mean degree of either one of the two networks is smaller than one and for the other network bigger than one, the giant component in the latter network emerges and market share inequality occurs. Case 3: In the regions where the mean degree of both networks is greater than one, giant components emerge in both networks. In these regions we have the pluralism solution of the consumers’ choices and hence, no noticeable market share inequality is apparent. Once consumers have purchased products, firms’ managers adapt their locations to reflect the pattern of consumers’ preferences. Firms’ managers are assumed to use an “unconstrained” adaptive rule that constantly modifies product characteristics in the search for more customers. The rule searches for customers using a “win-stay, lose-shift” algorithm (Nowak and Sigmund,1993; Bendor et al., 2003; Laver, 2005): if the previous move increased market share the manager makes another unit move in the same direction. If the previous move did not increase market share, the manager makes a unit random move in the opposite direction chosen randomly within the half space toward which it now faces.15 Managers use no information whatsoever about the global geography of the product-characteristics plane. They have no knowledge of the ideal point of any consumer but applying recursively the limited feedback from their local environment they pick up effective clues about the best direction in which to move. Once firms’ managers have adapted firms’ positions, consumers readapt and once more are active in the opinion- networks/firms which lie inside their consideration sets. Then they are again affected by peers’ opinions till their final decision on which product to purchase, then firms readapt to the new configuration of consumers’ preferences and the process iterates continuously. Location choice ~~~~~~~~~~~~~~~ In the previous section we analyzed consumers’ behaviour when their opinions are affected by the opinions of their network peers and the effect of this process in firms’ market share. Next, we consider firms’ location choice on the product-characteristics plane with origin (0,0) assuming that consumer ideal points’ distribution is normal on both product- characteristics’ dimensions.16 We check whether certain product differentiation patterns emerge. Summarizing, Fig. 4 discloses the following scenarios for the different phases of the model: Scenario I:〈k〉A ≫ 〈k〉B. In this case, firm B moves toward the center of the product-characteristics plane searching for customers while the firm A with better connected consumers can afford to roam anywhere in the product characteristics plane. Scenario II:〈k〉A ≈ 〈k〉B. In this case, no firm has a significant advantage over its competitor due to the presence of the opinion-network. Therefore, very much along the lines predicted by the traditional spatial competition model, both firms systematically move toward the center of the product-characteristics plane. Scenario III:〈k〉A ≪ 〈k〉B. In this case, it is firm B that can afford to roam anywhere in the product-characteristics plane when the firm A with the poorly connected underlying opinion-network is hunting customers moving towards the highest density of consumers.","This paper contributes to the literature of product-characteristics competition at the firm-level by proposing an evolutionary economic approach that focuses on the role of consumers’ social networks in the creation of new product characteristics. In the evolutionary framework, firms’ competition for market share depends on the preferences/tastes of many economic agents that interact and exchange knowledge/ideas/opinions, thereby forming complex social networks which condition the evolution of the process. The influence of consumers’ social networks in the development of new products is neglected in the traditional literature with ‘passive’ and homogeneous consumers: consumption is aggregated into economy’s demand side and it is reduced to the purchase decision only, excluding social interactions. On the supply side, firms’ managers incorporate new characteristics to their products that consumers can either buy or not without taking into account consumers’ heterogeneous preferences and the role of their social networks. Filling this gap, we put forth a product-characteristics competition spatial computational model based on complex systems science. Moving away from the forward-thinking strategic game theoretic models, we argue that consumers’ boundedly rational behaviour within their opinion exchange networks can result in concentrated power nodes in the network structure of the market. We confirm Halu et al. (2013)’s finding that the key feature to get the higher share of the overall product-purchases is the connectivity of the consumers’ opinion networks corresponding to different firms. Regarding firms’ location, we use agent-based modeling to study multi-characteristics competition in the evolving market. We are interested in location competition among multiple firms in a multidimensional product-characteristics environment in which consumers and firms care for more than one product characteristic. We show that for the case of strong inequality in the density of firms’ underlying opinion-networks, the firm with the highly connected network is roaming anywhere in the product-characteristics plane away from its origin, despite the fact that this is the location of the highest density of consumers. The opposite is true, with the firms approaching the center more often, when their underlying networks are weak or when all underlying networks are strong and no firm has a relative advantage on this. Most of the relevant literature considers consumers’ preferences as exogenous, referring to behavioral heuristics as determinants that lie outside the sphere of economics (Bowles, 1998). Valente (2012) argues that this assumption is used as a justification for aggregating towards a representative consumer and avoiding the economic analysis of heterogeneous agents with different preferences. Experimental and cognitive psychology empirical evidence shows that preferences are not exogenous but rather seem to be constructed during their elicitation (Shafir et al., 1993) and, furthermore, are influenced by the preferences of other agents (Tversky and Kahneman, 1981; Kahneman et al., 1982). This means that an interested party, like a firm competing for customers, might be able to shape consumers’ decisions towards a specific option by influencing their preferences or affecting the opinions of their peers. Besides, it is well known that a very large share of companies’ expenditures is devoted to marketing usually designed to press consumers to adopt a particular perspective/opinion of the product. In this paper, we demonstrate how consumers’ social networks can be considered the link between the supply side of the market (managers’ decision on marketing and strategic directions) and the demand side (consumers’ preferences influenced by managers’ strategies). In this way, we capture the feedback loop between consumers and firms, namely consumers’ preferences guiding firms’ innovation strategies and on the other way around, firms shaping consumers’ preferences by pursuing the consolidation of dense underlying social networks. The consumers’ social networks are appealing to organizations not only as a low cost way of reaching an audience, but also for their increasing penetration into everyday life and decision making due to the acceleration of the social media in the web (Sadovykh et al., 2015). Our paper is also related to the recent works by Vitell (2003), Vitell (2015), Schlaile et al. (2016), Schlaile et al. (2017), Muller (2017) on the role of consumers in responsible innovation and demand. This strand of literature illustrates that consumers’ heterogeneity and bounded rationality plays an important role in the creation and diffusion of responsible innovations. The results suggest that when all consumers focus solely on negative characteristics and demand only “responsible” products and services then the complex interplay between demand and supply can create situations that are inferior to scenarios with fewer responsible consumers. Along these lines, our model could easily be expanded to include negative product characteristics allowing the identification of central actors that have the power of influencing their peers’ opinions towards consuming more responsibly. Our complex networks approach may also be appropriate to address the issue of public and stakeholder engagement for stimulating responsible innovation (Jackson et al., 2005; Chilvers, 2008; Delgado et al., 2011; Owen and Goldberg, 2010; Blok et al., 2015). Though we miss the analytical tractability of the relevant theoretical models, our simulated results demonstrate the feasibility of agent-based techniques to describe and explore firms’ locational choices, as well as the capability of the proposed model to further advance the analysis in alternative ways beyond the game- theoretic framework. The proposed model should be seen only as a starting point for further research on the role of consumers in shaping markets since the real world is far more complex than our model with various aspects and factors influencing consumers’ decision making. One interesting direction for future research might be the extension of our model towards considering the role of committed agents i.e. the situation in which a fraction of consumers always remains active in one of the networks, never changing its opinion (Galam and Moscovici, 1991; Galam and Jacobs, 2007; Xie et al., 2011; Mobilia et al., 2007; Halu et al., 2013; Masuda, 2015). Likewise, it might be also interesting to see whether the assumption of some “iconoclast consumers”, i.e. consumers who intentionally behave opposite to the tendencies of their peers, alters our results."],["The aim of this study was to identify the most important variables determining current differences in physical stature in Europe and some of its overseas offshoots such as Australia, New Zealand and USA. We collected data on the height of young men from 45 countries and compared them with long-term averages of food consumption from the FAOSTAT database, various development indicators compiled by the World Bank and the CIA World Factbook, and frequencies of several genetic markers. Our analysis demonstrates that the most important factor explaining current differences in stature among nations of European origin is the level of nutrition, especially the ratio between the intake of high-quality proteins from milk products, pork meat and fish, and low-quality proteins from wheat. Possible genetic factors such as the distribution of Y haplogroup I-M170, combined frequencies of Y haplogroups I-M170 and R1b-U106, or the phenotypic distribution of lactose tolerance emerge as comparably important, but the available data are more limited. Moderately significant positive correlations were also found with GDP per capita, health expenditure and partly with the level of urbanization that influences male stature in Western Europe. In contrast, male height correlated inversely with children's mortality and social inequality (Gini index). These results could inspire social and nutritional guidelines that would lead to the optimization of physical growth in children and maximization of the genetic potential, both at the individual and national level. --------------------------------------------------------------------------------","The increase of height in the industrialized world started only ∼150 years ago. In the past, adult stature reflected fluctuations in environmental factors influencing physical growth (mainly climate, diseases, availability of food and population density) and considering that these conditions have almost never been optimal, it is not surprising that height could not approach maximal genetic limits. Paradoxically, the tallest people in Europe before the start of the industrial revolution may have been Early Upper Paleolithic hunters from the Gravettian culture that emerged at least ∼36,000 calibrated years ago (Prat, 2011) and is connected with the migration from the Near East that brought Y haplogroup (male genetic lineage) I-M170 to Europe (Semino et al., 2000). The estimated stature of Gravettian men is thought to be tall in the whole of Europe: 176.3 cm in Moravian localities (n = 18) and 182.7 cm in the Mediterranean (n = 11) (Dočkalová and Vančata, 2005). These exceptional physical parameters can be explained by a very low population density and a diet rich in high quality animal proteins. However, Late Upper Paleolithic and especially Mesolithic Europeans were noticeably smaller; Formicola and Giannecchini (1999) estimated that the average height of Mesolithic males was 173.2 cm (n = 75) in the Carpathian Basin and Eastern Europe, and only 163.1 cm (n = 96) in Western Europe. Perhaps the lowest values in European history were recorded in males from the Late Neolithic Lengyel culture (in the Carpathian Basin during the 5th millenium BC), who were only 162.0 cm tall (n = 33) (Éry, 1998). A new increase (up to ∼169 cm) came during the 3rd millenium BC with the advance of Eneolithic cultures (Corded Ware culture, Bell-Beaker culture) (Vančata and Charvátová, 2001; Dobisíková et al., 2007) and it is tempting to speculate that it was just in this period, when beneficial genes of lactose tolerance spread in Central and Northern Europe, because their origin is placed in the Neolithic era (Itan et al., 2009). During the next millenia, stature of Europeans changed mostly due to climatic conditions and usually remained within the range of ∼165–175 cm (Hermanussen, 2003; Steckel, 2004; Koepke and Baten, 2005; Mummert et al., 2011 etc.). The modern increase of stature in the developed world is closely tied with the positive effect of the industrial revolution. Ironically, these socioeconomic changes were accompanied by a new temporary decline of body size in the late 18th century and then again in the 1830s, which probably resulted from bad living conditions in overcrowded cities, a worsening of the income distribution and the decline of food available per capita1 (Komlos, 1998; Steckel, 2001; Zehetmeyer, 2011). These negative circumstances were overcome at the end of the 19th century and since that time the trend of height increase in Europe has been practically linear (Baten, 2006; Danubio and Sanna, 2008; Baten and Blum, 2012). Before WWII, its pace was faster in the northern and middle zone of industrialized Europe, while after WWII, it was mainly countries in Southern Europe that experienced the biggest increments (Hatton and Bray, 2010). Approximately since the 1980s, a beginning deceleration or even stagnation started to be apparent in some nations. After being the tallest in the world for 200 years, the US was overtaken by many Northern and Western European countries such as Norway, Sweden and the Netherlands (Komlos and Lauderdale, 2007). Various authors tried to identify the most important environmental variables that contributed to this dramatic increase of height in the industrialized world. Besides GDP per capita (as the main indicator of national wealth), the most frequently mentioned factors are children's mortality, the general quality of health care, education, social equality, urbanization and nutrition (Komlos and Baur, 2004; Baten, 2006; Bozzoli et al., 2007; Baten and Blum, 2012; Hatton, 2013 etc.). However, the number of investigated European countries is usually limited to the western half of the continent and to former European colonies, because information from the rest of Europe is less accessible. In the present work, we fill this gap. The aim of this study can be summarized in four basic points: Mapping the current average height of young males in nations of European descent and describing the actual dynamics of the height trend. Identifying main exogenous (environmental) factors that determine the final male stature. Creating recommendations that could lead to the maximization of the genetic potential in this regard. Investigating the role of genetics in the existing regional differences in male height. Collection of anthropometric data ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In order to have the most representative basis for our comparisons, we collected a large number of anthropometric studies of young European males from the whole continent (including Turkey and the Caucasian republics), and also data on young white males from Australia, New Zealand and USA. Whenever possible, we preferred large-scale nationwide surveys that have been conducted within the last 10 years and included at least 100 individuals. Forty-six out of 48 values in Table 1 were based on studies that were conducted between 2003 and 2013. The remaining two studies (from Australia and Lithuania) were done between 1999 and 2001. In case we found two or more representative sources from this period, we preferred those with higher measured values. In general, information from the southeastern part of Europe was the most problematic, because large-scale, nationwide anthropometric research is often virtually non-existent. The only country from which we had data on less than 100 individuals was Georgia (n = 69).2 In seven cases, we found the most recent data only on local populations (Bosnia and Herzegovina, Georgia, Iceland, Latvia, Romania, Sweden and United Kingdom). The urban surveys from Iceland and Latvia may increase the real national mean, but the study from Iceland was undertaken in the greater area of Reykjavik that encompasses ∼70% of the Icelandic population. The mean for the United Kingdom (177.7 cm) was computed from the average height in England, Scotland and Northern Ireland, weighted by the population size. A weighted average of six adult, non- university samples from Bulgaria was 175.9 cm (n = 643), which agrees with a rounded mean 176 cm (n = 1204) reported in the age category 25–34 years in a recent nationwide health survey CINDI 2007 (P. Dimitrov–pers. communication). In the case of Romania, we used an unweighted average of 3 local surveys (176.0 cm). Similarly, the average height in Bosnia and Herzegovina (182.0 cm) was computed as an unweighted mean of two local studies in the city of Tuzla, Northeastern Bosnia (178.8 cm), and in Herzegovina (185.2 cm), respectively. Because of the large (6+ cm) difference between these regions, this estimate must be considered only as very rough. However, considering that Herzegovinians used to be the tallest in the country and inhabitants of Tuzla the shortest (Coon, 1939), it appears reasonable.3 The average of Herzegovina was further corrected by +1 cm for unfinished growth in 17-year olds, which increased the average of Bosnia and Herzegovina to 182.5 cm. The addition of +1 cm to the height of 17-year old Ukrainian boys increased the Ukrainian mean to 176.9 cm.4 In another two countries (Armenia and Moldova), male height was estimated, based on highly representative data of 20–24 year old women from demographic and health surveys (DHS): 159.2 cm for Armenia (n = 1066) and 162.3 cm for Moldova (n = 1099). These female means were multiplied by 1.08, which was the average ratio of male and female height in 36 countries, where data on both sexes were available from the same study. The obtained values—171.9 cm for Armenia and 175.3 cm for Moldova—appear reasonable, because male and female height in these 36 studies mutually highly correlated (r = 0.91; p < 0.001). No information on measured height was available from Luxembourg, Malta and Wales. Exogenous and genetic variables ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Data from these anthropometric surveys were subsequently used for a geographical comparison of European nations (Fig. 1) and for a detailed analysis of key factors that could potentially play a role in the current trends of male height in Europe, primarily with the help of statistics from the CIA World Factbook (the Gini index), FAOSTAT (food consumption) and especially the World Bank, which contained all other variables that we examine (GDP, health expenditure, children's mortality, total fertility, urbanization).5 Since the final adult stature can be influenced by changes in living standards up to the cessation of growth, it was not surprising that long-term averages of these variables routinely produced higher correlation coefficients than data from the latest available year. However, data from all these sources were available only since 2000 and our analysis was therefore limited to long-term means from the period 2000–2012. The only exception was the Gini index (which was based on the latest available year from 1997 to 2011) and nutritional statistics (FAOSTAT 2000–2009), where the mean of Montenegro was computed only from the period 2006–2009. Genetic profiles of European nations (the phenotypic incidence of lactose tolerance, Y haplogroup frequencies in the male population) were obtained from various available sources and the most numerous samples were always preferred. Only samples with at least 50 individuals were included (see Appendix Tables 3a–d). Frequencies of Y haplogroups (male genetic lineages) E-M96, G-M201, J-P209, R1a-M420 and I-M170 were available from 43 countries (with only Australia and New Zealand missing), but the information on R1b-U106 and R1b-S116 was more limited, because these two lineages were discovered relatively recently (Myres et al., 2011). Nevertheless, most of the missing information concerned the eastern and southeastern parts of Europe, where their frequency is often close to zero and thus largely unimportant for this study. A similar limitation applied for lactose tolerance, where we found data only from 27 countries. Statistical analyses ~~~~~~~~~~~~~~~~~~~~ The analysis of the data collected was performed by the software package Statistica 12. Statistical relationships between variables were first investigated via Pearson's linear correlation coefficients. Subsequently, we identified the most meaningful factors via a series of stepwise regression analyses. Distribution of average male height in Europe ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The average height of 45 national samples used in our study was 178.3 cm (median 178.5 cm). The average of 42 European countries was 178.3 cm (median 178.4 cm). When weighted by population size, the average height of a young European male can be estimated at 177.6 cm. The geographical comparison of European samples (Fig. 1) shows that above average stature (178+ cm) is typical for Northern/Central Europe and the Western Balkans (the area of the Dinaric Alps). This agrees with observations of 20th century anthropologists (Coon, 1939; Coon, 1970; Lundman, 1977). At present, the tallest nation in Europe (and also in the world) are the Dutch (average male height 183.8 cm), followed by Montenegrins (183.2 cm) and possibly Bosnians (182.5 cm) (Table 1). In contrast with these high values, the shortest men in Europe can be found in Turkey (173.6 cm), Portugal (173.9 cm), Cyprus (174.6 cm) and in economically underdeveloped nations of the Balkans and former Soviet Union (mainly Albania, Moldova and the Caucasian republics). If we also took single regions into account, the differences within Europe would be even greater. The first place on the continent would belong to Herzegovinian highlanders (185.2 cm) and the second one to Dalmatian Croats (183.8 cm).6 The shortest men live in Sardinia (171.3 cm; Sanna, 2002). Large differences in height exist even within some countries. For example, the average height of Norwegian conscripts in the northern areas (the Finnmark county) is only 176.8 cm, but reaches 180.8 cm in the southern regions of Agder and Sogn og Fjordane,7 which is a very similar value like in Denmark and Sweden. In Finland, recruits are 178.6 cm tall on average, but men from the southwest of the country reach 180.7 cm (Saari et al., 2010). Both in Norway and Finland, the national average is pulled down by the northernmost regions that are inhabited by the Saami and are usually characterized by below-average values around 177 cm. In Germany, the tallest men can be found in the northwest, in areas adjacent to the Netherlands and Denmark (∼181 cm), while the shortest men come from the southeast, east and southwest (∼179 cm) (Hiermeyer, 2009). Unusually large differences are typical for Italy, where we can detect a 6.6 cm gap between Sardinia and the northeastern region of Friuli-Venezia Giulia (Sanna, 2002). In the western Balkans, geographical changes in body size are similarly striking, especially in Bosnia and Herzegovina. The current state of the height trend ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The trend of increasing height has already stopped in Norway, Denmark, the Netherlands, Slovakia and Germany. In Norway, military statistics date its cessation to late 1980s. Since that time, the average height of Norwegian conscripts aged 18–19 years has fluctuated between 179.4 and 179.9 cm.8 In Germany, the height of recruits has not changed since the mid-1990s, when it reached ∼180 cm (J. Komlos–pers. communication). Denmark Statistical Yearbooks also show a stagnation of the height of conscripts throughout the 2000s.9 In the Netherlands, national anthropometric surveys done in 1997 and 2009 documented the same height in young men aged 21 years (184.0 cm and 183.8 cm, respectively; Schönbeck et al., 2013). According to preliminary results of the Slovakian survey from 2011 (L. Ševčíková–pers. communication), the stature of Slovak males also remained the same, 179.4 cm in 2001 and 179.3 cm in 2011. In contrast, the positive trend of height still probably continues in Sweden and Iceland. The stature of Swedish conscripts was growing throughout the 1990s and reached 180.2 cm in 2004, which was a 0.7 cm increase in comparison with 1994 (Werner, 2007). In 2008–2009, Swedish boys of Scandinavian origin aged 17–20 years (n = 2408) from the Göteborg area were 181.4 cm tall (Sjöberg et al., 2012). Although these results may not be perfectly comparable, it is noteworthy that the mean in the latter study would decrease to 180.8 cm, if we included boys with immigrant origin. These made up 13.4% of the whole sample and their mean height was only 177.7 cm. This observation can lead to a speculation that the marked deceleration of the height trend in some wealthy European countries may also be due to the growing population of short-statured immigrants. Unfortunately, precise information on the “immigration effect” is not routinely addressed in nationwide anthropometric studies. Nevertheless, population means from an English health survey (2004) allow us to estimate that immigrants and their descendants decreased the height of men in England by ca. −0.3 cm, from 177.5 to 177.2 cm.10 A certain stagnation or even a reversal of the height trend can also be observed in other parts of Europe, but considering that it concerns relatively less developed countries, with short population heights (e.g. Azerbaijan), it stems from momentary economical hardships. The onset of the economic recession in the early 1990s was particularly harsh in Lithuania, where it may have affected the youngest generation of adult men, but not women: While in 2001, the average height of 18-year olds was 181.3 cm in boys (n = 195) and 167.5 cm in girls (n = 278) (Tutkuviene, 2005), in 2008 the means in the same age category in the area of Vilnius were 179.7 cm in boys (n = 250) and 167.9 cm in girls (n = 263) (Suchomlinov and Tutkuviene, 2013). Since the economic situation in Lithuania has dramatically improved during the last ∼15 years, the contemporary generation of teenage boys could potentially exceed the value from 2001. In contrast, the fastest pace of the height increase (≥1 cm/decade) can be observed in Ireland, Portugal, Spain, Latvia, Belarus, Poland, Bosnia and Herzegovina, Croatia, Greece, Turkey and at least in the southern parts of Italy. Interestingly, the adult male population in the Czech Republic also continues to grow at a rate of ∼1 cm/decade. Our own observations were supported by a recent cross-sectional survey of adolescents from Northern Moravia (Kutáč, 2013), in which 18-year old boys reached 181.0 cm (n = 169). In Estonia (Kaarma et al., 2008) and Russia (RLMS survey), the pace of height increments in males can be estimated at 0.5 cm/decade. In Slovenia, the upward trend in 21-year olds has been slowing down, from 0.5 cm during 1992–2002 to only 0.3 cm during 2002–2012 (G. Starc–pers. communication). In many of the wealthiest countries (Austria, France, Switzerland, United Kingdom and USA) the increase is mostly steady, but slow. For example, the height of Austrian conscripts increased by 0.5 cm in the period between 1991–1995 and 2001–2005 (E. Schober–pers. communication). Similarly, the stature of Swiss conscripts increased by 0.6 cm between 1999 and 2009 (Staub et al., 2011). Approximately the same pace is typical for white American men (Komlos and Lauderdale, 2007). In England, the height of men aged 25–34 years stagnated at ∼176.5 cm during the 1990s, but then started to grow again around 2000 and reached 177.6 cm in 2008–2009. Information from the southeastern part of Europe is much scarcer, but DHS surveys show a moderate upward trend (∼0.8 cm/decade) in Albanian men (Albania DHS survey, 2008–2009) and about 0.5 cm/decade in women from Armenia (Armenia DHS survey, 2005) and Moldova (Moldova DHS survey, 2005). Positive relationship between GDP per capita and the height trend in Europe ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As expected, the large regional differences in male height outlined in the Section 3.1 correlated with average GDP per capita (in USD, by purchasing power parity) according to the World Bank (2000–2012) (r = 0.38; p = 0.011) (Table 2; Fig. 2). However, this moderately strong relationship actually consists of two far narrower correlation lines in the former communist “Eastern block” (r = 0.54; p = 0.008) and non-communist “Western countries” (including Cyprus and Turkey) (r = 0.66; p < 0.001). The difference between them is about 20 000 USD. In other words, some of the former “Eastern” countries reach the same height like former “Western countries” despite ∼20,000 USD lower GDP and at the same level of GDP, they are ca. 3 cm taller. These facts could imply that the trend of increasing height in Europe continued uninterrupted on both sides of the “Iron curtain”, irrespectively of the increasing gap in national wealth, and today's differences in height were largely determined by regional differences existing in Europe at the end of the 19th century. To test this hypothesis, we compared documented averages of young males from the end of the 19th century with present values in 21 countries/regions (Fig. 3). This comparison shows that the height of males in the eastern half of Europe (and even in Southern Europe and the Netherlands) grew much faster than in the highly industrialized Western Europe. For example, since the 1880s male height has increased by ca. 11 cm in Great Britain and France, by 7.4 cm in Australia and by 8.1 cm in USA. In contrast, the height of Czech and Dutch men has increased by 16.6 cm, Slovak men by 16.7 cm, Slovenian men by 17.2 cm, Polish men by 17.5 cm and Dalmatian men by at least 18.3 cm. Furthermore, the height of young men in these countries is currently higher than in Western Europe, often very markedly. Apparently, this phenomenon can’t be explained by any economic statistics. A very eccentric example emerged especially in former Yugoslavia, where Bosnia and Herzegovina, Montenegro and Serbia reach much higher values of mean height than their GDP would predict. After the exclusion of 6 successor states of former Yugoslavia, the correlation with GDP per capita in the remaining 39 countries would steeply increase to r = 0.58 (p < 0.001) and it would even reach r = 0.80 (p < 0.001) in 17 countries of the former “Eastern block”. Remarkably, the huge decrease of height across the border of Montenegro and Albania (at least -9 cm), at similar values of GDP per capita, has hardly any parallel in other parts of the world, except a similar difference across the borders of malnourished North Korea/South Korea (Kim et al., 2008; Pak, 2010) and USA/Mexico (Del- Rio-Navarro et al., 2007; McDowell et al., 2008). Again, this points to the strong role of some specific local factors. A very different situation can be observed in former “Western” countries like Norway and Switzerland that have a much higher GDP per capita than we would assume on the basis of the correlation line. Interestingly, the same applies for the white population of USA, which is a phenomenon that is discussed (Komlos and Lauderdale, 2007). It seems that economic growth in these countries must have outpaced the positive height trend, but this fact still cannot explain, why the trend remains quite slow or non-existent during the last decades. Health expenditure ~~~~~~~~~~~~~~~~~~ Besides GDP per capita, another moderately significant relationship (r = 0.35; p = 0.018) was found between height and average health expenditure per capita (in USD, by purchasing power parity) according to the World Bank (2000–2011). The general pattern was very similar to that in Fig. 2, with two parallel and much stronger lines of correlation consisting of the former “Eastern” (r = 0.59; p = 0.003) and “Western” (r = 0.56; p = 0.007) countries (Fig. 4). Wealthy countries spend more money on healthcare than poorer countries, and USA is an anomaly in this regard. Relatively low expenditures on healthcare are typical for the Balkans, Caucasian republics, the Mediterranean and former USSR republics. Children's mortality under 5 years ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A closely related factor, average children's mortality under 5 years (per 1000 live births; according to the World Bank, 2000–2012), showed the highest impact on physical growth out of all examined socioeconomic variables (r = − 0.61; p < 0.001) that would persist even after the exclusion of an outlier—Azerbaijan (r = − 0.58; p < 0.001) (Fig. 5). Indeed, Hatton (2013) found that children's mortality (or a disease-free environment, respectively) has had a far bigger influence on the positive height trend in Europe since the mid-19th century than GDP per capita and other socioeconomic stats. Our data nevertheless indicate that in contemporary Europe, children's mortality is a serious issue only in some regions of the former USSR and the Balkans (mainly Azerbaijan, Georgia, Turkey, Moldova, Armenia, Albania and Romania). When the death rate falls below 10 cases per 1000 live births, this factor loses its predictive power, as shown by the sharp difference between correlation coefficients in the former “Eastern” block (r = − 0.75; p < 0.001) and “Western” countries (r = − 0.48; p = 0.022). Gini index ~~~~~~~~~~ The Gini index of social inequality (the latest available year after the CIA World Factbook, 1997–2011) also appeared to be a significant, negative variable (r = − 0.40; p = 0.007) (Fig. 6). The lowest values of the Gini index (i.e. the highest levels of social equality) are typical for Northern/Central Europe, while the highest Gini indices (the biggest social differences) can be found mainly in the Balkans, former USSR, and also in the United Kingdom and USA. A higher Gini average (33.2) indicates that social differences in countries of the former “Eastern block” are somewhat more accentuated, but they have no relationship to male stature (r = − 0.30; p = 0.16), most probably because of outliers in the least developed regions, where the Gini index recedes to the background at generally high levels of poverty. In contrast, the Gini index in “Western” countries is lower on average (31.5), but it does play some role in this regard (r = − 0.52; p = 0.013), mainly because of the wide polarity between egalitarian societies of Northern/Central Europe on one hand, and more socially stratified countries like USA, UK, Portugal and Turkey on the other hand. In general, social inequality tended to increase with decreasing GDP per capita (r = − 0.31; p = 0.038) and it also correlated with higher children's mortality rates (r = 0.39; p = 0.008). It is understandable that when lower social classes suffer from an uneven distribution of wealth, they are likely to reach lower height on average and pull the national mean down. A huge role of socioeconomic differences can be demonstrated on some examples from the Balkans, where it concerns mainly the contrast between urban and rural populations. For example, male height in Romania can range from 168.2 cm in rural areas of Suceava county up to 179.5 cm in middle-sized towns of the same region (Vasilov, 2001). Nevertheless, male height in Lithuania, Latvia and Bosnia and Herzegovina is still above-average, despite rather high Gini indices. Urbanization (% urban population) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The effect of urbanization (measured as the proportion of the urban population according to the World Bank, 2000–2012) was quite low (r = 0.30; p = 0.046), when Europe is viewed as a whole, but its role increases dramatically, when only former “Western countries” are taken into account (r = 0.55; p = 0.007), or when former Yugoslavia is excluded (r = 0.56; p < 0.001) (Appendix Fig. 1). This is also reflected by the high association between urbanization and GDP per capita (r = 0.69; p < 0.001) (Appendix Table 1). In the former “Eastern block”, urbanization in the above mentioned period was much lower (60.5% vs. 77.0% in the “West”) and has no relationship to stature (r = 0.14; p = 0.53). Based on our data, it seems that the level of urbanization starts to affect stature only after more than ∼70% population is concentrated in cities. Total fertility ~~~~~~~~~~~~~~~ The analysis of fertility (total number of children per woman) was based on the statistics from the World Bank (2000–2011). The presumption is that if the number of children is low, the expenditure of parents per a child can increase, and hence living conditions for its development improve. This factor apparently played an important role in the acceleration of the height trend in Europe in the past (Hatton and Bray, 2010), but this does not seem to be the case in contemporary Europe (r = − 0.11; p = 0.47) and especially in “Western” countries (r = 0.12; p = 0.60) (Appendix Fig. 2). The correlation coefficient was medium high only in former “Eastern” countries (r = − 0.47; p = 0.024), largely due to higher birth rates in economically disadvantaged regions of the Balkans and former USSR. Considering that birth rates are already very low in all regions of Europe, height differences among nations are unlikely to be influenced by this variable. To the contrary, it is possible that in the future, height will correlate positively with higher fertility, as indicated by the positive relationship of fertility with GDP per capita (r = 0.33; p = 0.025), health expenditure (r = 0.36; p = 0.016) and urbanization (r = 0.40; p = 0.006) (Appendix Table 1). Nutrition ~~~~~~~~~ The available data on protein consumption from the FAOSTAT database (2000–2009) show that male stature generally correlates positively with animal proteins (r = 0.41; p = 0.005), particularly with proteins from milk products in general (r = 0.47; p = 0.001), cheese (r = 0.44; p = 0.002), pork meat (r = 0.42; p = 0.004) and fish (r = 0.33; p = 0.028).11 In plant proteins, the effect was clearly negative (r = − 0.50; p < 0.001), and was particularly strong in proteins from wheat (r = − 0.68; p < 0.001) (Fig. 7) and cereals in general (r = − 0.59; p < 0.001), but partly even in rice (r = − 0.38; p = 0.01) and vegetables (r = − 0.35; p = 0.017) (Table 3). Total protein consumption is quite an insignificant dietary indicator, because its relationship to male stature is very weak (r = 0.19; p = 0.20). When we combined the intake of proteins with the highest r-values, the correlation coefficients further increased, as evidenced by the relationship betwen height and the total consumption of proteins from milk products and pork meat (r = 0.55; p < 0.001), or from milk products, pork meat and fish (r = 0.56; p < 0.001) (Fig. 8). All these comparisons still failed to explain the extreme height of the Dutch. The secret of the tall Dutch stature emerged only after we plotted the total consumption of proteins from milk products, pork meat and fish against the consumption of wheat proteins. This “protein index” was the strongest predictor of male height in our study, when all 45 countries were included (r = 0.72; p < 0.001) (Fig. 9).12 Without former Yugoslavia, the correlation further increased to r = 0.77. The ratio between milk and wheat proteins gave almost the same result (r = 0.71; p < 0.001). These findings are logical, because according to the amino acid score (AAS, the proportion of the worst represented amino acid), wheat contains one of the poorest proteins among all kinds of food due to a severe deficit of lysine, whereas animal proteins contain one of the very best proteins (see Appendix Table 2 and further e.g. Milford, 2012; FAO Expert Consultation, 2013). Furthermore, the digestibility of protein from plant foods is between 80% and 90%, while the digestibility of animal protein is higher, usually ∼95%, and can reach 100% in boiled meat.13 Therefore, the PDCAAS (Protein Digestibility Corrected Amino Acid Score, i.e. true protein quality) of plant proteins is even lower and it is mainly the ratio between the best and worst proteins that determines the overall level of the diet.14 Implications for rational nutritional programs ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The fundamental role of nutrition among exogenous variables influencing male stature has very important implications. While many economic and social factors are usually beyond an individual's direct control (personal income, total fertility and living in urban areas being rather exceptions), nutrition is a matter of personal choice. It is true that the accessibility of certain foodstuffs is limited by national wealth (GDP per capita), but the relationship is not as high as we would expect. GDP per capita has the most direct influence on the total consumption of animal protein (r = 0.82), meat (r = 0.78), cheese (r = 0.75) and exotic fruits like pineapples, oranges and bananas (r≥0.67), but not so much on the “protein index” (r = 0.57; p < 0.001). This shows that nutrition in all countries could be significantly improved via rational dietary guidelines, irrespectively of economic indicators. In this context, it should be noted that the quality of proteins (expressed as the “protein index”) in the wealthiest nations of European descent has a tendency to deteriorate during the last 2–3 decades (Fig. 9 and Appendix Fig. 3). This trend is characterized by a partial decrease in the consumption of milk products, beef and pork meat, and increasing consumption rates of cheese, poultry and cereals. In contrast, the quality of nutrition in the Mediterranean and some countries of the former Eastern block has been improving (Appendix Fig. 4). In the light of these findings, it is understandable, why other wealthy nations haven’t reached the height standard seen in the Netherlands and why their positive height trend slowed down or stopped. We suspect that these negative tendencies are due to the combination of the inadequate “fast-food” nutrition with some misleading dietary guidelines such as “modern healthy eating plates” of the Harvard School of Public Health. These are currently promoted by certain public initiatives in Czech schools and empasize vegetables and whole grains at the expense of animal proteins, while recommending to limit milk intake.15 Although some of these recommendations could certainly improve overall health in adults, our calculations based on the FAOSTAT data show that they should be viewed as detrimental for the healthy development of children. The possible health risks of a diet based on high-quality animal proteins seem to be exaggerated as well and we intend to address all important issues related to height and nutrition in a more detailed paper, based on our own research using actual statistics of the incidence/prevalence of cancer and cardiovascular diseases in Europe and long-term averages of FAOSTAT data (1993–2009). Y haplogroups I-M170 and R1b-U106: possible genetic determinants of extreme tallness in Europe ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Although the documented differences in male stature in European nations can largely be explained by nutrition and other exogenous factors, it is remarkable that the picture in Fig. 1 strikingly resembles the distribution of Y haplogroup I-M170 (Fig. 10a). Apart from a regional anomaly in Sardinia (sub-branch I2a1a-M26), this male genetic lineage has two frequency peaks, from which one is located in Scandinavia and northern Germany (I1-M253 and I2a2-M436), and the second one in the Dinaric Alps in Bosnia and Herzegovina (I2a1b-M423).16 In other words, these are exactly the regions that are characterized by unusual tallness. The correlation between the frequency of I-M170 and male height in 43 European countries (including USA) is indeed highly statistically significant (r = 0.65; p < 0.001) (Fig. 11a, Table 4). Furthermore, frequencies of Paleolithic Y haplogroups in Northeastern Europe are improbably low, being distorted by the genetic drift of N1c-M46, a paternal marker of Ugrofinian hunter-gatherers. After the exclusion of N1c-M46 from the genetic profile of the Baltic states and Finland, the r-value would further slightly rise to 0.67 (p < 0.001). These relationships strongly suggest that extraordinary predispositions for tallness were already present in the Upper Paleolithic groups that had once brought this lineage from the Near East to Europe. The most frequent Y haplogroup of the western and southwestern part of Europe is R1b-S116 (R1b1a2a1a2-S116) that reaches maximum frequencies in Ireland (82%), Britain, France and on the Iberian peninsula (∼50%), and is in all likehood tied with the post-glacial expansion of the Magdalenian culture from the glacial refugium in southern France/Cantabria (Fig. 10b).17 As already mentioned, Mesolithic skeletons in Western Europe were typically petite (∼163 cm) (Formicola and Giannecchini, 1999), which could imply that this haplogroup would be connected with below average statures. When we omitted 14 countries of Eastern and Southeastern Europe, where the frequencies of R1b-S116 are mostly close to zero, the relationship between R1b-S116 and male height reached statistical significance (r = − 0.48; p = 0.049) (Appendix Fig. 5). Therefore, the frequency of R1b-S116 may limit the future potential of physical stature. A related branch, R1b-U106 (R1b1a2a1a1-U106), apparently has a very different history and is typical for Germanic speaking nations, with a frequency peak in Friesland (∼43%). Its correlation with male height in 34 European countries was rather moderate (r = 0.49; p = 0.003), but increased to 0.60 (p = 0.007) without Eastern and Southeastern Europe. Visually, the distribution of this haplogroup resembles that of I1-M253 and I2a2-M436 in Northern and Central Europe, which would indicate that they shared a common local history, very different from R1b-S116. This assumption is also supported by the fact that out of all examined genetic lineages in Europe, combined frequencies of R1b-U106 and I-M170 have the strongest relationship to male stature (r = 0.75; p < 0.001 in 34 countries) (Fig. 11b and Appendix Fig. 6). The fourth most important Y haplogroup endogenous to Upper Paleolithic Europe, R1a-M420, is typical for Slavic nations (with ∼50% frequency in Belarussians and Ukrainians, and 57% in Poles) and has a steeply curvilinear relationship with height that peaks in the case of Poland (178.5 cm) and then abruptly decreases, largely due to a rising proportion of I-M170 in the genetic pool of tall nations. As a result, it is not a good predictor of male height (r = 0.19; p = 0.21) (Appendix Fig. 7), but combined frequencies of I-M170 and R1a-M420 reach statistical significance (r = 0.52; p < 0.001) (Appendix Fig. 8). After the exclusion of N1c-M46, the frequency of R1a-M420 in the Baltic states would rise above neighbouring Slavic nations (up to 66.5% in Latvia), but its predictive power still would lag behind I-M170 (r = 0.25; p = 0.11). Such a result implies that this lineage may have a moderate position in Europe, as for genetic predispositions for height. Y haplogroups E-M96, G-M201 and J-P209 represent the most important non-indigenous, post-Mesolithic lineages that spread to Europe mainly during the Neolithic expansion from the Near East (6th millenium BC). Expectably, their combined frequencies are the highest in the southeastern part of Europe (>40%) and decrease below 5% in the most remote areas of the continent such as Ireland, Scotland and the Baltic region (Appendix Fig. 9). Their relationship to male stature was markedly negative (r = − 0.64; p < 0.001) (Appendix Fig. 10) and remained significant even after an adjustment for GDP (r = − 0.57; p < 0.001) or the “protein index” (r = − 0.40; p = 0.009). This result indicates that the extremely small statures of early agriculturalists stemmed both from severe undernutrition and genetics. Multiple regression analyses of socioeconomic, nutritional and genetic data ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Considering that data on the frequencies of I-M170, R1a-M420 and three post-Mesolithic lineages (E-M96, G-M201 and J-P209) were available from 43 countries (except Australia and New Zealand), we included these 43 countries into a series of multiple regression analyses. When we examined the relationship between male height and 6 socioeconomic variables previously discussed above, a forward multiple regression identified children's mortality and the Gini index as the only factors contributing to a fully saturated model (R = 0.63, adj. R2 = 0.37; p < 0.001). Nutritional factors included into a forward stepwise regression consisted of the “protein index” and 9 significant (p < 0.05) protein sources with a daily intake higher than 3.0 g protein per capita (see Table 3). Except pork meat, all of them played an important role (R = 0.89, adj. R2 = 0.73; p < 0.001). Among three genetic factors, only two (I-M170 and E + G + J) turned out to be meaningful (R = 0.76, adj. R2 = 0.56; p < 0.001). Subsequently we made a standard regression analysis that included all 17 meaningful variables significantly related to stature (p < 0.05): two genetic factors (I-M170, E + G + J), five socioeconomic factors and 10 nutritional factors. Here, five variables turned out to be statistically significant (Table 5a), but the low values of tolerance indicated a very high degree of multicollinearity. A simpler, fully saturated model based on a forward stepwise regression (R = 0.89, adj. R2 = 0.75, p < 0.001) consisted of eight variables (Table 5b).18 The “protein index” determined in this study was by far the most important, followed by Y haplogroup I-M170 and cheese (the source of the most concentrated milk proteins).19 Interestingly, this model also explained the seemingly outlier position of the Dinaric countries and the most important factor here was the inclusion of the genetic component (I-M170) (Appendix Table 4 and Fig. 11). When examining the causes of the unexpected gap in height between “Eastern” and “Western” countries in Fig. 2, we observed that out of all variables examined in this study, Y haplogroup I-M170 and the “protein index” had the most marked effect on the increase of R-values (to R = 0.71, adj. R2 = 0.48 and R = 0.73, adj. R2 = 0.50, respectively) and hence a narrowing of this gap, when added as independent variables into a multiple regression that incorporated male height and GDP per capita. In other words, the tallest countries of the former Eastern block enjoy a relatively high level of nutrition that is often better than in “Western” countries with similar or higher GDP values. In addition, their means of male height are further elevated due to genetic factors, which especially applies for former Yugoslavia. The role of lactose tolerance: another possible genetic predictor of height in Northern and Central Europe? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The fundamental role of milk in the nutrition of European nations stems from the high prevalence of lactose tolerance—a very valuable genetic trait allowing a sufficient intake of milk even during adulthood. Although many scientists are still not certain about the evolutionary role of lactose tolerance in Europe, the findings of this study show that the ability to consume milk must have been a tremendeous advantage, because it opened up an additional source of nutrition that not only guaranteed survival in the periods of food shortage, but contained nutrients right of the highest natural quality. Interestingly, the correlation of lactose tolerance with the consumption of whole milk (r = 0.10; p = 0.62) and milk products in general (r = 0.59; p = 0.001) was weaker than that between lactose tolerance and male height (r = 0.71; p < 0.001) (Fig. 12). This is the exact opposite of what we would expect. Considering that lactose tolerance is a genetic trait, it is possible that it may correlate with certain genes that determine tall stature. When examining the relationship between lactose tolerance and frequencies of Y haplogroups, we noticed a very high association of lactose tolerance with the combined frequency of I-M170 and R1b-U106 (r = 0.74; p <0.001 in 23 countries) (Appendix Fig. 12). The correlation coefficient further increased to r = 0.78 (p < 0.001, 21 countries), when R1b-S116 was added, although it overestimated lactose tolerance in R1b-S116-rich areas such as the British Isles, France, Spain and Italy (Appendix Fig. 13 and 14). Based on these results, we hypothesize that the spread of lactose tolerance started from the Northwest European territory between Belgium and Scandinavia, which is the presumed area of origin of the Funnelbeaker culture during the 4th millenium BC (Appendix Fig. 15a). These populations probably carried Y haplogroups I-M170 (I1-M253 and I2a2-M436), R1b-U106 and partly R1b-S116. The successor Corded Ware populations most likely expanded from present-day Poland during the 3rd millenium BC and their genetic composition may have been somewhat different. In any case, contrary to some assumptions, which are also based on limited archeogenetic data (Haak et al., 2008), the distribution of lactose tolerance does not seem to be related to Y haplogroup R1a-M420 (r = − 0.10; p = 0.62 in 26 countries), and this picture would hardly change even when missing data from Latvia and Belarus were added. Obviously, if Corded Ware populations consisted of R1a-M420 males, they were not closely related to the earliest carriers of the lactose tolerance genes. Alternatively, they may have belonged to some minor subbranches of R1a-M420.20 Besides that, the spread of lactose tolerance in Western Europe and in the British Isles must have been tied with some other cultural circle, perhaps the Bell-Beaker culture during the second half of the 3rd millenium BC (Appendix Fig. 15b). Similarly, lactose tolerance can’t be connected with Near Eastern agriculturalists, i.e. Y haplogroups E-M96, G-M201 and J-P209 (r = − 0.73; p < 0.001), despite computer simulations showing that genes of lactose tolerance were originally acquired from agricultural populations in Central Europe (Itan et al., 2009).","The dramatic increase of height in Europe starting during the late 19th century is closely linked with the beneficial effect of the industrial revolution. Our comparisons show that this effect is manifested in generally higher standards of living, better healthcare, lower children's mortality, lower fertility rates, higher levels of urbanization, higher social equality and access to superior nutrition containing high-quality animal proteins. In the past, some of these factors may have been more important for a healthy physical growth than they are today and some of them are important only in certain regions, but in general, the most important exogenous factor that impacts height of contemporary European nations is nutrition. More concretely, it is the ratio between proteins of the highest quality (mainly from milk, pork meat and fish) and the lowest quality (i.e. plant proteins in general, but particularly wheat proteins). Besides that, we discovered a similarly strong connection between male height and the frequency of certain genetic lineages (Y haplogroups), which suggests that with the gradual increase of living standards, genetic factors will increasingly be getting to the foreground. Even today, many wealthy nations of West European descent are ca. 3 cm smaller than much poorer countries of the former Eastern block with the same nutritional statistics. Another evidence for this genetic hypothesis recently appeared in the study of Turchin et al. (2012), who found systematically higher presence of alleles associated with increased height in US whites of North European ancestry than in Spaniards. Remarkably, the quality of nutrition in the wealthiest countries shows signs of a marked deterioration, as indicated by the decreasing values of the “protein index”. This can illuminate the recent deceleration/cessation of the positive height trend in countries like USA, Norway, Denmark or Germany, which was routinely explained as the exhaustion of the genetic potential. In our opinion, this assumption is still premature and with the new improvement of nutritional standards, some increase still can be expected. A very specific case is that of Dinaric highlanders, who are as tall or even taller than the wealthiest nations, in spite of considerably lower living and nutritional standards. However, their seemingly outlier position can be explained by strong genetic predispositions. In the near future, we want to focus on this problem in detail via our planned field project directly in the Dinaric Alps, in collaboration with the University of Montenegro."],["Recent work on collective intertemporal choice suggests that non-dictatorial social preferences are generically time inconsistent. We argue that this claim conflates time consistency with two distinct properties of preferences: stationarity and time invariance. While time invariance and stationarity together imply time consistency, the converse does not hold. Although non-dictatorial social preferences cannot be stationary, they may be time consistent if time invariance is abandoned. If individuals are discounted utilitarians, revealed preference provides no guidance on whether social preferences should be time consistent or time invariant. Nevertheless, we argue that time invariant social preferences are often normatively and descriptively problematic. --------------------------------------------------------------------------------","Many important decisions in economic life require groups of people with heterogeneous time preferences to implement a collective consumption plan. Examples abound: families must decide on savings and intra-household resource allocation, partners in a firm must decide how to distribute profits between payouts to themselves and capital investments, communities with property rights over a natural resource must decide on an extraction plan, and resource rich countries must decide how to consume the proceeds from their sovereign wealth funds. In each of these examples an asset is held in common and is consumed dynamically over time, and the stake-holders in the decision very often have heterogeneous time preferences (Frederick et al., 2002). How should such decisions be evaluated and made, given people's different attitudes to time? Recent theoretical and applied work (Jackson and Yariv, 2014, 2015; Adams et al., 2014) suggests that any attempt to make such collective intertemporal choices in a non-dictatorial fashion is doomed to confront a time inconsistency problem (see also earlier observations by Marglin, 1963; Feldstein, 1964; and Zuber, 2011). This paper argues that this finding is due to a conflation of time consistency with two a priori distinct properties of preferences: Stationarity and Time Invariance. Stationarity, introduced by Koopmans (1960), is an independence property of preferences, while time invariance requires preferences over future consumption streams to be the same in each evaluation period (this does not rule out preference reversals – quasi-hyperbolic time preferences, for example, are time invariant). Jackson and Yariv (2014, 2015), for example, show a necessary conflict between collective intertemporal choice and stationarity, but assume that social preferences must be time invariant. Yet while time invariance and stationarity together imply time consistency, the converse does not hold; time invariance is not required for time consistency. If the differences between these three properties of intertemporal preferences are made explicit it becomes clear that time consistent collective intertemporal choice is possible. The differences between time consistent and time invariant social preferences are particularly subtle when individuals' preferences are discounted utilitarian, as is commonly assumed in the literature on the aggregation of time preferences. In this case, time consistency and time invariance are equivalent properties of individuals' preferences1 (Halevy, 2015), but not of social preferences. We show that if individuals have heterogeneous discount factors, and social preferences are utilitarian and non-dictatorial, these preferences may be time consistent or time invariant, but not both. Since these properties are indistinguishable for discounted utilitarian individuals, but not for social preferences, the choice to impose one or the other property at the social level is ultimately a normative one – neither choice is more or less consistent with individuals' preferences. We argue however that time invariance is likely to be a problematic feature of social preferences – both normatively and descriptively – for within-group intertemporal choice. Time invariance may however be more descriptively plausible for between-group choice, as might occur in intergenerational decision-making. We relate these observations to the empirical literature that seeks to test the time consistency of household behaviour, suggesting that the definition of time consistency it employs is too restrictive.","To make our argument we begin by providing definitions of three distinct properties of intertemporal preferences: time consistency, time invariance, and stationarity. Once we have defined these properties, it will immediately be clear that any two of them implies the third, as observed in a simplified setting by Halevy (2015). We then show that the notion of time consistency used by e.g. Jackson and Yariv (2014, 2015); Adams et al. (2014) is in fact the conjunction of time invariance and stationarity. Time invariance and stationarity imply time consistency, but the converse does not hold. “[Stationarity] does not imply that, after one period has elapsed, the ordering then applicable to the ‘then’ future will be the same as that now applicable to the ‘present’ future. All postulates are concerned with only one ordering, namely that guiding decisions to be taken in the present.” – Koopmans et al. (1964) Halevy (2015) has observed in a simplified setting3 that any two of the three properties in Definition 1 implies the third. Inspection of the definitions shows that this observation holds in general.","If utilitarian social preferences are time consistent and non-dictatorial, the following interpretations of (7) are equivalent: Individuals have forward-looking welfare measures, and social preferences have τ-dependent welfare weights (8). Individuals experience lifetime utility, and welfare weights are constant.","We have argued that there is no conflict between time consistency and non-dictatorial collective intertemporal choice, provided that time invariance is abandoned. Should we nevertheless insist that social preferences be time invariant, thus ruling out the possibility of time consistency? Given that revealed preference cannot tell us which property to adopt if individual preferences are discounted utilitarian, how should we make this choice? In order to demonstrate what is at stake when choosing to model social preferences as time invariant it is helpful to consider the following stylized scenario: Ada and Bertha are identical twins; both are childless, single, and have identical wealth. On their 60th birthday they receive news that a distant cousin they've never met has died, and they will inherit his fortune. Their cousin died intestate, so there is no will to specify how the bequest should be divided between them. Ada is naturally impatient, while Bertha is patient. How should a utilitarian social planner allocate the bequest between them? Suppose that the planner is time consistent. In this case first best intertemporal allocations can be decentralized by establishing property rights over the bequest, and then allowing each sister to consume her share as she pleases. Suppose that the planner's ethics dictate that an equal share of the bequest be given to each sister. The planner makes this allocation, and walks away happy in the knowledge that she need never revisit this decision. Now consider a time invariant planner, and suppose that she also initially allocates equal shares to the sisters, before departing on other business, never intending to return. To her surprise however, the planner is asked to revisit the sisters 10 years later and assess their well-being. Ada has led a wild life in the intervening decade, consuming her inheritance much faster than Bertha. But the planner asks herself: what has changed since the sisters were 60? Although Ada has had a better life than Bertha, this is in the past and can't be changed. Since the circumstances are the same, except that the sisters happen to be 10 years older, the planner seizes some of Bertha's carefully saved cash, and reallocates it to Ada, so that their holdings of what remains of the bequest are again equal. Ada immediately spends her new wealth on fast cars and fancy footwear, while Bertha continues to put something away for a rainy day. Are there situations where time invariance might seem a more appropriate assumption than time consistency? We believe that modeling social preferences as time invariant might be more plausible if we are attempting to describe conflicts between successive groups of different people, as might occur in e.g. intergenerational decision-making. To illustrate, consider the following example: A nation discovers an enormous oil deposit in its territorial waters. The deposit is large enough to generate rents for generations to come, which will be invested in public goods that benefit all citizens equally. Citizens are born in different time periods, and have finite lives.8 Suppose that each citizen is either Avaricious (i.e. favours a high intergenerational discount rate), or Benevolent (i.e. favours a low intergenerational discount rate). In each period the currently living citizens must decide how much oil to extract. How is this likely to be done? The crucial difference between the ‘twins’ and ‘oil’ scenarios we have sketched is that in the case of the twins we were concerned with intertemporal distribution within a fixed group of people over time, whereas in the oil case we were concerned with intertemporal distribution between different groups of people. In the former case all agents who are affected by a decision are present at all points of time, whereas in the latter only a subset of agents affected by a decision are present at each point in time. Time consistency seems a much more plausible descriptive property of social preferences than time invariance in the former case, since the ahistorical nature of time invariant preferences would be seen as deeply unfair by current agents. In the case of between-group choice however, although time invariance may do violence to the preferences of previous or successive groups, these individuals have no voice in current decision processes. As a descriptive matter, it seems more likely that the interests of past and future groups are only relevant insofar as current decision-makers account for them. To our minds this makes time invariance a more plausible descriptive modeling assumption in this case. Of course, this argument makes no claims about the normative status of time invariance as a property of social welfare criteria.","While we find time invariance to be a not implausible descriptive property in between- group choice, the empirical literature focusses exclusively on within-group choice. Our hypothesis that time consistent, but non-invariant, social preferences are likely to prevail in this case receives support from the innovative study of household consumption behaviour in Adams et al. (2014). Although they strongly reject the hypothesis that household behaviour can be described by time invariant and time consistent preferences of the form (11), they find considerable support for the hypothesis that household consumption behaviour can be rationalized by time consistent social preferences with time varying welfare weights (8). Adams et al. (2014) refer to this as the ‘full efficiency’ model, and claim that this identifies individual discount rate heterogeneity as a source of household time inconsistency. Yet the appropriate interpretation of these findings is that collective choices can often be rationalized by a time consistent utilitarian model, provided that household preferences account for the cumulative lifetime utility experienced by each household member."],["In the last century, U.S. diets were transformed, including the addition of sugars to industrially-processed foods. While excess sugar has often been implicated in the dramatic increase in U.S. adult obesity over the past 30 years, an unexplained question is why the increase in obesity took place many years after the increases in U.S. sugar consumption. To address this, here we explain adult obesity increase as the cumulative effect of increased sugar calories consumed over time. In our model, which uses annual data on U.S. sugar consumption as the input variable, each age cohort inherits the obesity rate in the previous year plus a simple function of the mean excess sugar consumed in the current year. This simple model replicates three aspects of the data: (a) the delayed timing and magnitude of the increase in average U.S. adult obesity (from about 15% in 1970 to almost 40% by 2015); (b) the increase of obesity rates by age group (reaching 47% obesity by age 50) for the year 2015 in a well-documented U.S. state; and (c) the pre-adult increase of obesity rates by several percent from 1988 to the mid-2000s, and subsequent modest decline in obesity rates among younger children since the mid-2000s. Under this model, the sharp rise in adult obesity after 1990 reflects the delayed effects of added sugar calories consumed among children of the 1970s and 1980s. --------------------------------------------------------------------------------","In approximately two generations, obesity has become an epidemic across the Developed world (Goryakin et al., 2017). Worldwide, obesity nearly tripled between 1975 and 2016; by 2016, more than 650 million adults were obese, about 13% of the world's adult population (World Health Organization, 2018). Among children and adolescents aged 5–19, the global obesity rate rose from 1% in 1975 to 6% of girls and 8% of boys in 2016 (World Health Organization, 2018). Although the rise in U.S. obesity dates to the mid 20th century (Sobal and Stunkard, 1989), the most substantial and rapid increase in adult obesity has occurred over the past 30 years (Cook et al., 2017; Dwyer-Lindgren et al., 2013; Flegal et al., 2016; Kranjac and Wagmiller, 2016; Ogden et al., 2014). From 1990 to 2016 the national adult obesity rate almost doubled; in certain U.S. states (WV, MS, AK, LA, AL, KY, SC) it nearly tripled, from about an eighth of the population to more than a third. There are clearly many contributors to the obesity crisis, many of which have been studied over past decades through randomized control trials and through population-scale statistical analyses (Hruby and Hu, 2015; Rippe, 2013). Among the major factors in this recent increase, one is age. In the U.S. in 2015–16, obesity prevalence was about 36% among adults aged 20–39 years, 43% for 40–59 years, and 41% for 60 and older (Centers for Disease Control & Prevention, 2017a). In 2015–2016, nearly 20% of 6–19 year-olds were obese (Hales et al., 2017). Data from the state of Wisconsin (Wisconsin Health Atlas, 2019), which ranks near the middle of U.S. states in obesity, document the age profiles in more detail for the year 2015, showing the continuous rise in obesity from age two years to middle age, where the obesity rate plateaus near 47%, before falling off among those aged over 75 years (Fig. 1a). The increase in obesity through childhood is a recent phenomenon, however; in the U.S. in 1970s there was an average decrease in obesity rate from ages two to nineteen (Wang and Beydoun, 2002). Another major factor is socioeconomic status (Hruschka, 2012; Smith, 2017; Sobal and Stunkard, 1989; Tafreschi, 2015; Salmasi and Celidon, 2017; Villar and Quintana-Domeque, 2009). In the U.S., the increase in obesity rate between 1990 and 2015 has correlated inversely with median household income (Bentley et al., 2018). Looking at county-scale obesity rates from 2004 to 2013 (Centers for Disease Control & Prevention, 2017b), the average obesity among the 100 poorest counties in the U.S. increased from about 29.9% to 36.5%, but only from 21.2% to 24.6% among the 100 richest counties. As obesity rates have increased faster in poor communities, a likely driver has been the lowering cost of food calories (Mattson et al., 2014; Mullan et al., 2017), including sugar-sweetened beverages (Basu et al., 2013a; Hu and Malik, 2010; Johnson et al., 2007, 2016; Malik et al., 2006; Shang et al., 2012; Song et al., 2012; Wang et al., 2008). Since 1970, when high-fructose corn syrup (HCFS) was introduced at commercial scale into processed foods and sugar-sweetened beverages (SSBs), consumption of HFCS increased from virtually zero in 1970 to over 60 pounds per capita annually in 2000 (Gerrior et al., 2004). Overall, the consumption of added sugars increased from about 70 pounds per person annually in 1970 (340 cal/day) to almost 90 pounds per capita (440 cal/day) by the late 1990s (Fig. 1b). By the 2000s, the top 20% of consumers of caloric sweeteners were ingesting over 300 kcal from HFCS per day (Bray et al., 2004). Much of the rise in total added sugar consumption between 1970 and 2000 is attributable to HFCS (Fig. 1b), often in SSBs that contribute additional calories without suppressing appetite for the intake of other foods (Almiron-Roig et al., 2013; Mattes and Campbell, 2009). The correlation between the increased consumption of added sugars in U.S. foods and the rise of obesity (Fig. 1b) has been recognized (Bray et al., 2004; Basu et al., 2013b; Lustig et al., 2015), even by the sweetener industry (Corn Refiners Association, 2006). Consequently, sugar consumption has been hypothesized to be a primary driver of the recent obesity increase (Basu et al., 2013a,b; Bray et al., 2004; Bray, 2007; Lustig et al., 2015; Nielsen et al., 2002; Nielsen and Popkin, 2004; Popkin and Hawkes, 2016). Time series data from 75 countries, for example, indicate that a 1% rise in soft drink consumption in a country predicts about a 2% increase in obesity (Basu et al., 2013b). Critics of the sugar-obesity theory, however, point out the substantial time lag between changes in sugar consumption and change in obesity rate. After sugar consumption had begun to decline (Welsh et al., 2011), van Buul et al. (2014, p. 119) noted that “over the last 5 years, the global annual consumption of carbonated soft drinks has remained constant or even has declined “Obesity rates, however, seem to have increased independently of these shifts.” In 2014, for example, obesity rates were still rising, despite the 25% decrease in added sugar consumption among U.S. residents between 1999 and 2008 (van Buul et al., 2014; Welsh et al., 2011). By 2017, obesity rates had only just started to level off or decrease, in some but not all U.S. states. The delay is such that annual adult obesity rates from 2004 to 2013 (Centers for Disease Control & Prevention, 2017b) correlate quite well with annual per capita sugar consumption (U.S. Department of Agriculture, Economic Research Service, 2019) twelve years before, i.e. from 1992 to 2001 respectively, with r2 = 0.933 among the 100 poorest counties in the U.S. and r2 = 0.908 among the 100 richest counties. Although readily observed, this time lag remains unexplained. While most population health studies explore the effect of environmental, nutritional or behavioral factors on obesity levels, few have explicitly explored the temporal delay between cause and effect. A specific study helps illustrate why we would expect a substantial delay between increased sugar consumption and national obesity rates. In a 19-month study controlling for other variables, Ludwig et al. (2001) found that for each additional SSB consumed per day, BMI in children increased by a mean of 0.24 ± 0.14. This suggests an increase on the order of about 1.6 per in BMI decade per extra daily SSB. This “back-of-the-envelope” line of thinking motivates our specific modeling approach. Here we propose that obesity increase can be explained as the accumulated effect of childhood exposure to excess sugar calories together with continued consumption. Since childhood obesity tends to predict obesity in adulthood, the rising obesity rates of adults the 1990s and 2000s ought to reflect their diets as children of the 1970s and 1980s. In 1990, for example, the 18-year-olds entering the adult CDC dataset would have been born in or about 1972. Current adult obesity rates would have only just begun to register a decline in sugar consumption that began in 1998. With a model that can capture these different dynamics, with per-capita sugar consumption as the single input variable, we present obesity as the cumulative result of excess sugar consumption since 1970. In formulating a parsimonious, sugar-driven model, our goal is to replicate three different aspects of the obesity trends in the U.S.: (a) the overall increase of U.S. adult obesity since 1970; (b) the profile of obesity by age group for a recent year; and (c) the change in obesity rates by pre-adult age group since 1990.","As shown in Fig. 2, the model generates a matrix of the 46 calendar years, t, from 1971 to 2016, versus life ages, y, from age 2 through age 75. In each cell of this array, obesity in year t for age y is calculated as a function of obesity for the same cohort in the previous year, t − 1, when the cohort was one year younger, age y − 1. The model replicates the generational time lag through a simple stochastic process of superfluous sugar calories causing obesity rates to increase over the lifespan of each birth-year cohort. In this multiplicative process, early excesses are compounded over the long-term, such that excess sugar consumption in childhood register years later as rising adult obesity rates. The two parameters, β and α, are adjusted for the model to fit two aspects of the data: the timeline of increase in average obesity rate, and the profile of age versus obesity rate. The parameter α is the magnitude of the effect, whereas β reflects the log odds of non-obese individuals in a cohort becoming obese in the next year, as a function of their excess daily calories (expressed as fraction of a 2000 calories/day diet). For example, if α = −8.5 and β = 20, then for each age cohort we expect about 0.05% more obesity in the next year for every 100 extra daily calories (5% of a 2000 calorie diet). The logit model is non-linear, so with these same α and β values, but with 300 excess calories per day (15% of a 2000 calorie diet), we would expect an additional 0.4% of the non-obese portion of the cohort to become obese in the next year. The external input variable, ρt, is the caloric equivalent of the average excess sugar consumption per capita in the U.S., expressed for convenience as a fraction of a 2000-calorie diet, for each year since 1970 (U.S. Department of Agriculture, Economic Research Service, 2019). We used historical data (Wang and Beydoun, 2002) to approximate the age-obesity profile in 1970, which serves as the initial year of the model. The sugar data, recorded by the USDA (U.S. Department of Agriculture, Economic Research Service, 2019), are annual estimates of mean sugar consumed (cal/day) per capita in the U.S. since 1970. The data include four categories: refined sugar, HFCS, fruit juices and “other” sugars, and here we use the sum of all four of the categories as our estimate of daily added sugar calories for each year. To test the model versus change in U.S. obesity prevalence, national level obesity rates for adults adult (20 and over) in the United States were obtained from the CDC, for selected years 1988–1994 through 2015–2016 (Centers for Disease Control & Prevention, 2019). Two additional estimates of U.S. adult obesity from the 1970s were added; these are from the National Health and Nutrition Examination Surveys of 1971–1974 (NHANES I) and 1976–1980 NHANES II (Flegal et al., 2016). For obesity among pre-adult cohorts, we used data reported by Ogden et al. (Ogden et al., 2014) from 40,780 children and adolescents between 1988–1994 and 2013–2014.","Using two tunable parameters and the record of mean sugar consumption from 1970 to 2016, the model replicates the average U.S. obesity increase through calendar years and across age profiles. Using β = 19.9 and α = −8.5, the model simultaneously replicates, firstly, the growth in average U.S. adult obesity since 1990 (Fig. 3a). Secondly, using these same parameters, it simultaneously replicates a typical U.S. state's age-obesity curve for 2015 (Fig. 3b). The only aspect of the age profile not replicated is obesity for ages 75 and over in 2015 (Fig. 3b), which is the cohort who grew up before the rapid increase in added sugars in U.S. diets. The model also replicates aspects of the change in obesity rates by age group. Fig. 4 shows how the same model used in Fig. 3 (β = 19.9, α = −8.5) also yields similar timelines from 1990 to 2014 for ages 2 to 5, 6 to 11, and 12 to 19. The comparison in Fig. 4 shows that the model is predicting a decline in obesity rate first among ages 2 to 5 beginning in the late 1990s, then among ages 6–11 beginning in the early 2000s, and subsequently among teenagers after about 2006. In the model, these decreases in obesity rates are delayed responses to the decline in excess sugars in the late 1990s. The data, though noisier, reflect these declines with about the same timing for the two younger age cohorts, but not the teenagers (Fig. 4). For ages 6 to 11, the actual decline was less than predicted by the model, and for ages 12 to 19, there was no actual decline at all.","Here we have modeled the recent increase of U.S. adult obesity as being driven by added sugar calories over the lifespan of each birth-year cohort. With annual USDA sugar consumption figures as the input variable, the two-parameter model simultaneously replicates three different phenomena: the generational lag between sugar and obesity, the magnitude of the national rise in obesity, and a recent age-profile of obesity rates. Our results indicate that excess U.S. sugar consumption is at least sufficient to explain the timing and magnitude of adult obesity change in the past 30 years, even as other factors (Gentile and Weir, 2018; Muscogiuri et al., 2018; Smits et al., 2017) could be factored into future models. In doing so, the model addresses a key critique of the sugar-obesity hypothesis (van Buul et al., 2014; Soenen and Westerterp-Plantenga, 2007) by introducing the mechanism to explain the years of delay in cause versus effect. The sugar-obesity hypothesis is also supported by physiological, historical and economic evidence. Physiologically, sugar consumption elevates lipid levels in the blood (DiNicolantonio et al., 2015; Lustig et al., 2015; Grande, 1967; Grande et al., 1965; Kaufmann et al., 1966; Kuo and Bassett, 1965; Macdonald and Braithwaite, 1964; Teff et al., 2004). Although its specific role in the obesity epidemic is debated (Rippe, 2013; van Buul et al., 2014), fructose specifically does not stimulate insulin secretion or the production of the hormone (leptin) regulating long-term food energy balance, such that this sugar tends not to satisfy appetite (Bantle et al., 2000; Curry, 1989; Havel, 2005; Luo et al., 2015; Figlewicz and Benoit, 2009; Stanhope et al., 2009) and affects glucose metabolism, lipid profile and insulin resistance (Beyer et al., 2005; Bocarsly et al., 2010; Bray et al., 2004; Johnson et al., 2016; Jürgens et al., 2005; Pereira et al., 2017). Consumed from a young age, sugar consumption appears to have long-lasting effects, not just habitually but also physiologically, in ways that could explain a generational delay between U.S. sugar consumption and subsequent obesity rates. Infant and toddler obesity is correlated with high-sugar infant foods (Koo et al., 2018) and even fructose in breast milk (Goran et al., 2017). Sugar consumption during pregnancy leads to increases in recruitment of pre- adipocytes to adipocytes in utero, such that children are born with an increase in fat cells that will accumulate more fat during life (Goran et al., 2013). By the mid 1970s, children 0–2 years old were consuming about 6 grams of added sugar per kg of body weight, about three times that of adults at the time (Fomon, 1975; Life Sciences Research Office, 1976). From the 1970s to the 1990s, soft drinks were increasingly sweetened with HFCS, an inexpensive, domestically produced liquid sweetener (Corn Refiners Association, 2002). Between 1977 and 2001, sweetened beverage consumption increased by 135% across all age groups, equivalent to about 278 additional calories per day (Nielsen et al., 2002; Nielsen and Popkin, 2004). In the 2000s and 2010s, total added sugar intake in the US declined (Popkin and Hawkes, 2016), and by 2016, obesity rates in some U.S. states were leveling off. Because 75-year-olds experienced childhood before the large-scale increase of sugar in processed foods, they may have developed less lifelong preference for added sugars in foods, but also it may be that they never laid down excess adipose tissue in during gestation. These explanations, potentially complementary to each other, may be addressed by detailed research on this elderly age group. The age-stratified obesity data we used run through 2013–14, and in the future it will be revealing to see whether teenage obesity has declined since 2014 as the sugar-driven model predicts. It may be that for high school ages in particular, the effect socioeconomic status on food choice (Campbell et al., 2019) is obscured by aggregated obesity statistics. Data from the Centers for Disease Control & Prevention (2018) from 1999 to 2017, for example, suggest that obesity among high schoolers may have peaked in 2011 and 2013 among Asians and whites, respectively, but that obesity had continued to increase among blacks and Hispanics. We speculate that poverty is a driver of sugar consumption (Alvarado, 2016; Hernandez, 2015; Shrewsbury and Wardle, 2012; Anderson, 2012; Datar, 2017; Hughes et al., 2010; Kowaleski-Jones et al., 2017; Mata et al., 2017; Moreno et al., 2016; Rosinger et al., 2017; Bentley et al., 2018). Economically, sugar is an inexpensive source of calories, and sweetened beverages have been a substantial portion of expenditures for low-income households (Garasky et al., 2016). Childhood obesity decreased after the 2009 changes in the US Special Supplemental Nutrition Program for Women, Infants, and Children (Daepp et al., 2019; Pan et al., 2019); we believe cutting the juice allowance by half helps explain why. If our model is correct, the effect of this 2009 change will follow these children into adulthood. In the future, our model should be tested against different socioeconomic and demographic subpopulations, as well as for nations where there exist historical data on sugar consumption and obesity rates. For example, our delayed-effect model could offer a plausible explanation for the “Australian Paradox” (Barclay and Brand-Miller, 2011), that sugar consumption decreased in Australia at the same time that obesity increased.","In summary, we have modeled the recent increase of U.S. adult obesity rates since the 1990s as a legacy of increased consumption of excess sugars among children of the 1970s and 1980s. Our model proposes, for each age cohort, that the current obesity rate will be the obesity rate in the previous year plus a simple function of the mean excess sugar consumed in the current year. With just these inputs, the model can replicate the timing and magnitude of the national rise in obesity, as well as the profile of obesity rates by age group, and the different patterns of change in obesity among children and adolescent age group, where reduction in obesity registered first among young children in the late 1990s. This supports the perspective that the rise in U.S. adult obesity after 1990 was a generation-delayed effect of the increase in excess sugar calories consumed among children of the 1970s and 1980s."],["The aim of this paper is to better understand one of the mechanisms underlying the income-obesity relationship so that effective policy interventions can be developed. Our approach involves analysing data on approximately 9000 overweight British adults from between 1997 and 2002. We estimate the effect of income on the probability that an overweight individual correctly recognises their overweight status and the effect of income on the probability that an overweight individual attempts to lose weight. The results suggest that high income individuals are more likely to recognise their unhealthy weight status, and conditional on this correct weight perception, more likely to attempt weight loss. For example, it is estimated that overweight high income males are 15 percentage-points more likely to recognise their overweight status than overweight low income males, and overweight high income males are 10 percentage-points more likely to be trying to lose weight. An implication of these results is that more public education on what constitutes overweight and the dangers associated with being overweight is needed, especially in low income neighbourhoods. © 2013 Published by Elsevier B.V. --------------------------------------------------------------------------------","Being overweight or obese is known to be bad for your health, yet the prevalence of obesity is increasing worldwide (Lobstein and Jackson Leach, 2007). Coined as the “most prevalent nutritional problem in the world” (Lau et al., 2007), the epidemic is most prevalent in developed countries. For example, in Canada, U.S., France and Australia, 23% (Linder et al., 2010), 33% (Dorsey et al., 2009), 17% (International Obesity Task Force, 2011) and 25% (International Obesity Task Force, 2011) of the population are classified as obese (BMI of 30 kilograms per squared metre or greater “for obese (inclusive of 30)), respectively. While obesity rates are similar for males and females, there is a divergence between genders with respect to being overweight. For example, in Canada 42.8% of males and 23.7% of females are overweight (body mass index (BMI) of 25 kilograms per squared metre or greater” (inclusive of 25)) or obese. The equivalent figures for the U.S., France and Australia are 40.1% and 28.6%, 41.0% and 23.8%, and 42.1% and 30.9%, respectively (International Obesity Task Force, 2011). In England, over 40% of men and 30% of women are overweight or obese (International Obesity Task Force, 2011), with predictions that without action, 60% of men, 50% of women and 25% of children will be obese or obese by 2050 (Butland et al., 2007). Action is being taken, however, with £75 million of public health funds and £200 million of external funds earmarked for a public health campaign called ‘Change4Life’ in 2009 (The Lancet, 2009). This campaign was launched in response to the extraordinary economic costs associated with the overweight population – approximately £7 billion per year in England (National Institute of Health and Clinical Excellence, 2006). These estimates include medical costs; being overweight is associated with an increased risk of type 2 diabetes, heart disease, stroke, high blood pressure, certain cancers (colon, breast, endometrial and gallbladder), and high cholesterol. However, the campaign as yet has not produced any visible signals that it is defeating the obesity epidemic. Therefore, given that the obesity epidemic is not waning either in England or in other developed countries, there is scope to investigate further its underlying causes. In this paper, we investigate the obesity-income gradient by estimating the impact of income on weight perception and weight control in a sample of overweight British adults. While those of high income may have a lower weight because they can afford a healthier lifestyle, it is also plausible that they have a more narrowly defined standard for acceptable body size and adjust their behaviour accordingly. This would suggest an income gradient with respect to weight perceptions and a subsequent role for weight perceptions in determining a person's propensity to pursue weight control. An independent income gradient–weight control relationship is also likely to exist owing to the higher opportunity costs associated with weight control for poorer people. Our work is related to two main strands of the obesity literature. The first of these is the literature that attempts to estimate the impact of income on the propensity to be overweight or obese. So far, many studies have found that higher socioeconomic status is related to a lower risk of obesity (Costa-Font et al., 2008; Wamala et al., 1997; Zhang and Wang, 2007). However, the endogeneity of income in a weight regression complicates these studies interpretation. That is, income may cause a person to be overweight, being overweight may cause lower income or common factors may affect both income and overweight status. These factors include individual heterogeneity such as self-discipline and impulsivity (Cutler et al., 2003), along with weight misperceptions, which we explore in this work. Attempts have been made to establish a causal relationship between BMI and income with mixed results. For example, Quintana-Domeque (2005) utilise the European Community Household Panel (ECHP), and exploit exogenous variation in household income owing to inheritance, gifts, or lottery winnings of €2000 or more to instrument for income in an obesity regression. They explore this relationship for nine countries and find a relationship between income and obesity only for women in both Denmark and Italy, and men in Finland. Notably, this work suffers from a weak instrument problem. In the U.S. context, Cawley et al. (2008) exploit exogenous variation in the social security policy but are unable to identify any statistically significant relationship between additional social security income and BMI in the elderly. Schmeiser (2009) examine the effect of family income changes on BMI and obesity using data from the National Longitudinal Survey of Youth 1979 cohort. They find that income significantly raises the BMI and probability of being obese for women only. Finally, using a longitudinal Swedish panel Ljungvall and Gerdtham (2010) estimate the impact of mean income, positive deviation from mean income and negative deviation from mean income on weight status using questionable instruments. They find income to be negatively related to obesity in general. The second strand of literature that our work relates to concerns itself with the relationship between actual body size and body size perception. Self-perception of body size is a factor that can influence whether weight loss is a concern. Clearly, if a person is unaware they are overweight they cannot fully internalise the costs associated with the health risks of their weight status. This is in line with research suggesting accurately perceiving oneself as overweight or obese results in a greater motivation to engage in healthy lifestyle behaviours (Baranowski et al., 2003 and Rhee et al., 2005). Given that misperceptions of a normal weight among the overweight and obese have been highlighted in the general literature (Collins et al., 1987; Kuchler and Variyam, 2003; Maximova et al., 2008; Paeratakul et al., 2002; Viner et al., 2006) as well as in the literature specific to the UK (Wardle, 2002; Johnson et al., 2008) the problem of a failure to internalise is one that may contribute to the obesity epidemic. This work aims to explore the role of an income gradient on weight perceptions. Specifically we focus on individuals who are the targets of obesity campaigns in England. That is, we focus on the overweight and obese. The potential for income to be associated with weight perceptions is linked to it being usual for poor individuals to have poor friends (Tigges et al., 1998; Wacquant and Wilson, 1989) and the likelihood that poorer people are more likely to be overweight or obese. Therefore, peer effects may imply an increased propensity for poorer people to perceive being overweight as a ‘healthy’ weight, which may reflect ideals of body weight among that group (Kemper et al., 1994). This arises because people's behaviour is likely to be influenced by the norms in their social environment. Thus, when overweight becomes the norm within a peer group, it is likely that the negative social stigma associated with being overweight is reduced. The idea that your social circle can affect your weight is supported by recent research. Christakis and Fowler (2007) find that weight gain spreads through a population like a contagious disease owed to individuals being influenced by their friends and relatives; though, Cohen-Cole and Fletcher (2008) re-estimate these effects and find them greatly reduced and not significant once a more thorough econometric methodology is utilised. Elsewhere, Maximova et al. (2008) have shown that young people's perceptions of weight is dependent on the weight of their parents and friends. Similarly, Blanchflower et al. (2008) describe a ‘keeping up with the Jones weight effect’ where weight perceptions and dieting are influenced by the individuals that surround us. Overall they suggest that individuals have different comparison groups, with the highly educated holding themselves to a ‘thinner’ standard. Oswald and Powdthavee (2007) argue that people have a utility function defined on relative weight and hence choose their weight with reference to the weight of their peers. Given the higher rates of obesity amongst the poor, this peer effect is likely to create an income gradient in weight perception and weight control, which further reinforces the obesity-income gradient. In addition, weight misperceptions among people of lower income may be explained by lower levels of health knowledge. Alternatively, those with higher levels of education may simply be more capable of processing health information available to them about the type of behaviours that yield them good health (Gottfredson and Deary, 2004).1 Thus far the role of the income gradient on misperceptions is under explored, however, Wardle and Griffith (2001) have examined the effects of socioeconomic status – defined as occupational social class. They find using a sample of British adults that higher SES people have higher levels of perceived overweight, more closely monitor their weight, and are more likely to state they are trying to lose weight. Understanding weight misperceptions is important given that those who are satisfied with being overweight are less likely to do anything about it. Conversely, those who are aware that they have an elevated BMI are more likely to take action. It is noteworthy that feeling overweight does not in itself motivate attempts at weight loss, however the majority of those who feel this way do try to lose weight (approximately 60%) according to some received studies (Horm and Anderson, 1993; Wardle and Johnson, 2002) and the literature generally points to a positive correlation between self perceived weight status and weight control (Crawford and Campbell, 1999, Forman et al., 1986 and Riley et al., 1998). Even once weight misperceptions are accounted for, given the higher opportunity cost of weight control for those of lower income it is likely that an independent income- weight control relationship will exist. For example, this greater opportunity cost arises because the neighbourhoods in which poorer people live have characteristics that are positively correlated with obesity such as poor walkability (Sallis et al., 2009), a lack of healthy food options (Zick et al., 2009), a higher presence of unhealthy food outlets (Harrison et al., 2011) and greater disorder (Burdette and Hill, 2008). Additionally, the literature has identified a relationship between income and healthy lifestyle choices including the propensity to exercise and eat well (Pampel et al., 2010) and higher rates of dieting (French et al., 1994; Jeffrey and French, 1996).","Our data source is the annual Health Survey for England (HSE), which is a household level survey that collects information through an interview, self-completion questionnaire and medical examination. We pool data from the 1997, 1998 and 2002 surveys and consider prime working age (25–60) respondents who, according to BMI measurements collected by a nurse, are of an unhealthy weight: defined either by BMI ≥ 25 or BMI ≥ 30. The individuals are unaware that they have been classified as ‘overweight’. The survey year, age and BMI ≥ 25 restrictions, as well as a restriction of non-missing income information, leaves us with an estimation sample of 9089. Data from 1997, 1998 and 2002 are used because in these years adult respondents were asked questions regarding their weight perceptions and weight goals. Specifically, individuals were asked: Given your age and height, would you say that you are: about the right weight, too heavy or too light? At the present time are you trying to lose weight, tying to gain weight or are you not trying to change your weight? The responses are used to define two binary variables. The first represents weight perception and equals one if the individual believes they are too heavy. Given that only those who are classified as overweight (or obese) are included, this variable also measures weight misperceptions. The second key variable represents weight control, and equals one if the overweight (or obese) individual is trying to lose weight. Approximately 75% of overweight (BMI ≥ 25) respondents feel too heavy and approximately 60% are trying to lose weight (equivalent percentages for the obese sample are 95% and 73%). In other words, 25% of respondents incorrectly perceive themselves as the right weight, and 40% are not trying to change their weight (very few overweight respondents feel they are “too light” or are “trying to gain weight”). However, these sample averages mask heterogeneity. For example, mean values of weight perception and control are 64% and 47% for men, and 87% and 75% for women. This suggests that women are more likely to recognise their overweight status and more concerned with their weight. Similarly, the raw propensities depend upon income; high income respondents are more likely to recognise their overweight status. Our empirical strategy is to sequentially estimate richer variants of Eqs. (1) and (2) in order to test whether the income effect can be ‘explained’ by mediating variables. The purpose of this exercise is to gauge which covariates are the potential pathways between income and our outcome (weight misperception/weight control). First we add a set of baseline controls, which represent demographic information that is personal to the individual. Therefore, model (1) includes gender, age, age-squared, married, divorced, number of children, black Caribbean or African, Asian, year 1997 and year 1998. Second, given the link between obesity and environment our second set of variables (model (2)) pertains to area of residence information: rural versus metropolitan and North-East, North-West, Yorkshire, West-Midlands, East-Midlands, South-East, and South-West. Next, model (3) adds general health indicators: long-standing illness and limiting long-standing illness. Given that income is essentially one dimension of socio-economic status that is correlated with other dimensions, our next step is to add some of these dimensions. Therefore, model (4) adds highest educational attainment and employment status: degree, vocational qualification, A levels, O levels, and employed. In the context of this data, those who have O and A levels stay in secondary education until the ages of 16 and 18 respectively. Finally, model (5) adds occupation categories: professional, associate professional and technical, administrative and secretarial, skilled trades, personal service, sales and customer service, plant and machine operatives, and elementary.2 Importantly, all sets of control variables (1–5) include BMI since it is a significant predictor of weight perceptions and weight control even amongst samples of overweight and obese respondents. Note that if we did not control for BMI the estimated income coefficient would be downward biased – BMI is negatively correlated with income and positively correlated with our dependent variables.3 Table 1 includes descriptive statistics for some of the included covariates. Given Eqs. (1) and (2) are estimated using a non-random subset of the population, the income coefficients may suffer from sample selection bias. The direction of any bias is likely to be negative because the negative income–obesity relationship implies that high individuals in the sample (i.e. overweight) care relatively little about their weight. Therefore, the true income effects are likely to be larger. To test this proposition we estimated probit sample selection models and found that the estimated income effects were indeed larger than those from our probit regression models. However, these models were identified solely through the assumption of jointly normal disturbance terms, as our data does not contain a defendable exclusion restriction. For this reason we prefer estimates from probit regression models.","The upper panel in Table 2 presents estimates from the weight perception probit regressions for overweight samples (BMI ≥ 25), and the lower panel presents estimates for obese samples (BMI ≥ 30). The reported standard errors are clustered at the household level are reported to allow for correlation between weight perceptions and weight control of individuals living in the same household. The figures represent the percentage-point change in the probability of feeling too heavy for a 1 unit change in log income (i.e. marginal effects). Note that moving from the 5th percentile to the 95th percentile of the income distribution (i.e. from impoverished to wealthy) has the effect of increasing log income by roughly 2.5. Thus, the first estimate in column 1 – 0.046 – implies that moving from a low to a high income increases the probability of (correctly) feeling too heavy by around 12 percentage points. The equivalent effect for overweight men is roughly 15 percentage-points (relative to a sample mean of 64%, equalling a 23% increase). Three key findings are gained from Table 2. First, regardless of the sample – male, female, overweight or obese – high income respondents are significantly more likely than low income respondents to recognise they are ‘too heavy’. Second, income effects are larger for men than women: using the baseline set of control variables, the male effect is roughly 2 times larger than the female effect in both the overweight and obese samples. A potential explanation is that low income men are more likely to view larger body size as an indicator of prowess and dominance than high income men, thus creating an income effect in body size perception (McLaren and Kuh, 2004. Third, controlling for the respondent's area and their health has little effect on the income estimates. One potential explanation for significant income effects is that low income regions tend to have insufficient health services, and therefore, residents of these regions receive less information regarding the thresholds for overweight and its dangers. However, the similarity of the estimates in rows (1) and (2) suggest that this is not the case. It appears that part of the income effect – but not all – can be explained by higher income individuals having greater education and working in different occupation types. For example, the income effect for overweight males drops from 0.066 in model (3), to 0.044 in model (4) with education controls, and then to 0.028 in model (5) with occupation controls. Having a university degree is estimated in model (4) to increase correct weight perception (relative to no qualifications) by 5.2 percentage-points, while having a managerial level occupation is estimated in model (5) to increase correct weight perception (relative to an unskilled, elementary occupation) by 8.4 percentage-points. An alternative estimation approach, which can aid interpretation, is to replace the continuous log income with income categorical variables. If we take this approach and include dummy variables indicating the quintile of the income distribution, we find that individuals in the top quintile (richest 20%) are 8 percentage points more likely to feel too heavy than individuals in the bottom quintile (poorest 20%) – estimated results available upon request. Equivalent effects for the female and male samples are 5 percentage points and 11percentage points, respectively. We have also considered whether there exists nonlinear relationships between log income and misperceptions, but all higher order polynomial terms were insignificant for all subsamples and covariate sets used in the analysis.4 Our overall interpretation of the results in Table 2 is that income is an important predictor of weight perceptions given that income remains a significant predictor of perceptions even after controlling for a very large set of covariates that are correlated with income. Importantly, this result is not being driven by all high income people, regardless of their weight status, feeling fat – perhaps driven by a propensity to seek some idealised body image. We also examined whether income increases the propensity for an individual who is of normal weight to incorrectly perceive themselves as overweight. In this regression the estimate of the log of income is not significant (p = 0.472). Thus it appears that the mechanism is truly that income promotes correct self-assessment. Table 3 presents similar estimated income effects from the weight control models. We again find income to be a significant determinant of whether an individual is trying to lose weight. Importantly, these models are estimated with only those respondents who feel too heavy, and thus, income is having an effect on weight control even after controlling for the effect of income on weight perceptions. We again find the income effect is larger for men, at least in the overweight sample. For example, in the model that controls for demographics, area of residence, illness, education, employment and occupation, it is estimated that a rich overweight male is 12 percentage-points more likely than a poor overweight male to be trying to lose weight. Unlike the weight perception results in Table 2, occupation and education do not appear to be modifying the relationship between income and weight control.5 Finally, we investigate whether the income relationships in Tables 2 and 3 hold equally for younger (<40) and older (≥40) sub-samples. The estimates in Table 4 are from probit models estimated with the baseline set of controls and samples of overweight respondents. They suggest that the effect of income on the probability of correctly perceiving yourself as ‘too heavy’ is larger for older respondents. For example, the effect for older female respondents is twice as large as the effect for younger female respondents (0.039 versus 0.020), while the difference between older and younger male respondents is 2 percentage-points (0.066 versus 0.046). In contrast, the estimation results from the weight control models suggest that the estimated income effect does not differ by age.","This work investigates explanations for the strong relationship between SES and obesity using a large survey of overweight British adults. The aim is to better understand why the poor are more likely to have elevated BMIs, so that effective policy interventions can be developed. Our work finds that overweight low income individuals are more likely to incorrectly believe they are a healthy weight, and conditional on weight misperceptions, less likely to attempt weight loss. Both of these effects are larger for males than females. Further research is required to order to tease out the differences in these gender effects and establish casual effects. A suggestion is that they may be driven by peer group effects, whereby males’ peer group composition is more sensitive to income than females. Our two main findings feed into very different policy options. Firstly, for those who incorrectly believe they are a healthy weight, further research is needed to investigate the underlying drivers. People often rely on comparison with peers to make assessments of their weight status, rather than relying upon medical advice. Given that obesity has become the norm within low income groups, the existence of such effects implies that people with lower incomes tend to be less concerned with being overweight, reinforcing the obesity–income relationship. This problem may arise because of mixed messages in the media concerning optimal body weight size. Deciphering these mixed messages is more likely to be achieved by those of higher socioeconomic status. The implication of this reasoning is that more public education on what constitutes overweight and the dangers associated with being overweight may be needed, especially in low income neighbourhoods. This is however not the end of the story, as our results also highlight that there are many who realise they are overweight but are not attempting weight loss. Again, the cause for this may lie with peer effects models. That is, within a peer group, friends may realise they are overweight but reinforce bad eating and exercise habits. Therefore whilst it is not that peer group effects cause the SES/obesity gradient per se, they do contribute to the growing disparity once a threshold number of individuals with low SES are overweight. Furthermore, it may be more difficult for those of lower socio- economic status to lose weight given that their home environment often lacks the necessary inputs such as an availability of healthy foods and exercise opportunities. The latter extends from lack of gyms through to safe areas for walking. To remedy such environmental level factors would involve policy changes that go beyond health-specific policies. It should be noted that some commentators argue that any policy to address the obesity epidemic is paternalistic and should be avoided. That is, we should not intervene as individuals rationally choose their own weight (by consuming and expending a certain number of calories). It is unlikely that individuals can weigh up the costs and benefits, both future and present, of this choice. It is also unlikely that the overweight weigh up the costs that fall on the health service owing to the obesity epidemic. As discussed, these costs are expected to rise to £10 billion per year by 2050 with no government action (Butland et al., 2007). Equally they are unlikely to consider the wider costs to society and business, such as decreased tax revenue and loss of productivity due to related illnesses, which are estimated to reach £49.9 billion per year (2007 prices) if the obesity epidemic is allowed to continue its current increasing trend (Butland et al., 2007). Therefore, we argue that policy makers must take some action, and from our work, additional education on what constitutes a healthy weight and adopting a healthy life style for low income households could be beneficial, without being regressive given that our work highlights that income is a strong predictor of weight misperceptions and control. Perhaps these could be piloted in a subset of low income neighbourhoods initially so cost effectiveness can be gauged. A bigger challenge lies with addressing the environmental factors that may inhibit individuals losing weight. While the literature has done well in highlighting that various environmental factors do indeed influence a person's health status, the next challenge is to identify the main factors that could do the ‘heavy lifting’ with respect to addressing the obesity epidemic."],["A growing literature indicates that effects of early-life health on adult economic outcomes could be substantial in developing countries, but the magnitude of this effect is debated. We document a robust gradient between the early-life mortality environment to which men in India were locally exposed in their district and year of birth and the wages that they earn as adults. A 1 percentage point reduction in infant mortality (or 10 point reduction in IMR) in an infant's district and year of birth is associated with an approximately 2 percent increase in his subsequent adult wages. Consistent with theories and evidence in the literature, we find that the level of schooling chosen for a child does not mediate this association. Because of its consequences for subsequent wages, early-life health could also have considerable fiscal externalities; if so, public health investments could come at very low net present cost. --------------------------------------------------------------------------------","A growing literature documents that workers exposed to better early-life health and less disease in early life have higher human capital as adults. Although this literature has largely focused on developed countries, economists have hypothesized that an effect of early-life health and disease externalities could be importantly larger in developing countries, where disease insults are worse and more varied (Currie and Vogl, 2013; Spears, 2012b). If early-life health indeed importantly limits human capital in developing countries, wages could be a mechanism through which health has important effects on developing economies and, because of income and consumption tax revenue, on the government's budget. However, the magnitude of these effects is a topic of current debate in the development economics literature (Acemoglu and Johnson, 2007; Bleakley, 2010a; Hansen, 2014). It is therefore important to understand and quantify relationships between early-life health and subsequent wages in developing countries (Vogl, 2014). Our article documents a robust gradient between the health environment to which today's workers in India were exposed as infants in past decades, and the wages which they now earn. We match male workers in nationally representative survey data on wages in 2005 to district-level estimates of infant mortality in their year of birth, which we use as a measure of early- life health and disease, following Acemoglu and Johnson's investigation of mortality and GDP. We then use a double fixed effects (place and time) identification strategy to compare workers in the same district labour market today who were exposed to different mortality regimes when they were born. Our results suggest that being born in a district- year with a higher infant mortality rate (IMR) – and corresponding worse health and disease environment – is associated with a significant but plausible reduction in earnings decades later. A 1 percentage point reduction in infant death (or a 10 point reduction in IMR) in the environment to which an infant was exposed in his year of birth is associated with an approximately 2 percent increase in his subsequent wages as an adult. Because important threats to early life health remain widespread in India and other developing countries, these estimates are of continuing economic and policy importance. Bleakley (2010a) reviews theory and evidence that disease – and especially child health – could have important effects on adult human capital and income in developing countries. Although substantial effects would be consistent with current theory and recent empirical findings, there is debate about the quantitative importance of disease to economic development. Bleakley further presents a model demonstrating that improved early-life health is likely to increase both the returns to schooling and the opportunity costs of schooling (in the form of higher earning ability for children and young adults); as a result, improvements in early-life health may not lead to large increases in the optimally-chosen quantity of formal schooling. Our setting allows us to test this prediction, and indeed we find links between improvements in early-life mortality and adult wages, but no association between IMR and subsequent schooling. Infectious disease is well-known to have negative externalities on neighbors. Because of its consequences for subsequent wages, the disease environment also can have fiscal externalities, which are of importance for government cost-benefit analysis in the context of a developing country with limited fiscal capacity and many competing potential expenditures. We apply our estimates to compute considerable consequences for government tax revenues of reductions in income and consumption due to early-life health and disease. We show that relatively modest effects of early-life health on individual economic outcomes can add up to quantitatively important overall economic and fiscal effects, in the context of a developing country with a large burden of disease externalities. One consequence is that public action to improve the disease environment faced by infants could come at very low net present cost to governments (Alderman and Behrman, 2006). The rest of the article proceeds as follows. Section 2 provides a brief overview of relevant literatures and the Indian context, and Section 3 details our empirical strategy. Section 4 then presents our main empirical results, and Section 5 demonstrates the robustness of our results in several directions and explores the mechanisms driving our findings. Section 6 translates the main empirical results into implied consequences for government revenues and a simple measure of welfare. Section 7 concludes.","We study adult male workers in a representative survey of India. Children in India today are exposed to a considerable disease burden in early life, which was even greater at the time when today's adults were children. In 1970 – the year before our data begins, when 35 year old workers in 2005 were born – nearly 20 percent of children died before their fifth birthday, and 13 percent of infants died in their first year of life (Unicef, 2012). Infant mortality in India has fallen to about 4.1 percent today, but this still substantially exceeds infant mortality of 1.1 percent in China and 3.3 percent in Bangladesh, a poorer neighboring country. If early-life health and disease is an important constraint on development and income, it would be of considerable importance in India, where about one-fifth of all births occur. Effects of early-life health on adult economic circumstance ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ An active literature in economics documents that healthier babies are more likely to become healthier and more productive children and adults (Currie, 2009). Early-life health matters because the first few years are a critical developmental period; children who have better health and net nutrition in early life are more likely to reach their physical growth potentials and are more likely to reach their cognitive potentials (Case and Paxson, 2008). Indeed, much of this literature, unable to match adult wages to early-life conditions, has used adult height as a proxy for early-life health (e.g. Vogl, 2014). Most of this literature has focused on developed countries (e.g. Deaton and Arora, 2009). However, the impact of early-life health in developing countries may be an even more important part of labour market outcomes than in developed countries, relative to heterogeneity in genetic potentials: disease conditions are worse, care and remediation may be less available, and heterogeneity in health insults is likely to be larger than in developed countries. Economists have linked various measures in the chain from early-life human capital accumulation to its long-run consequences: childhood and adult height, childhood and adult cognitive achievement, and adult wages. For example, Bozzoli et al. (2009) show that people are shorter in countries with higher infant mortality, whereas de Oliveira and Quintana-Domeque (2014) find that GDP per capita in year of birth is the main correlate of height in late-20th-century Brazil; Alderman et al. (2009), Alderman et al. (2006), and Glewwe et al. (2001) show that taller children have greater cognitive achievement; Vogl (2014) shows that taller adults in Mexico earn more money; and many papers in labour and development economics document economic returns to cognitive achievement. Among the few studies that have been able to directly connect variations in disease environment to gains in wages or consumption are two recent papers by Bleakley (2010b) and Cutler et al. (2010), who estimate effects of early-life exposure to malaria on adult wages in the Americas and on consumption in India, respectively.2 Our article expands this literature beyond malaria and links the early-life mortality environment directly to adult wages in a developing country; we then apply our estimates to compute consequences for government revenue, in the context of a large developing economy where our estimates are of continuing relevance. Endogenous education and the envelope theorem ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ It is common in the literature linking early-life health to adult wages to verify an intermediate effect on schooling (e.g. Vogl, 2014). Bleakley (2010a) presents a standard model in which children's education is optimally chosen to maximize lifetime income. The benefits and costs of education both depend on the child's health, and therefore on the disease environment. Bleakley notes that better health almost certainly increases the returns to schooling, but better health is quite likely to also increase the opportunity cost of schooling in an economy where children can also engage in productive work, as would have been common in India several decades ago. Therefore, it is not clear that the effect of health on education should be positive. Moreover, Bleakley (2010a) shows that a straightforward implication of the envelope theorem is that the chosen schooling level is unlikely to be an important mechanism of the effect of health on adult earnings: if schooling is chosen to maximize lifetime income, for example, individuals will attain schooling up to the point at which the marginal gain from schooling in terms of added lifetime earnings (accounting for foregone earnings while in school) is zero. Therefore, even if changes in early-life health impact cognitive development in such a way as to increase the quantity of schooling chosen, this change in schooling should have no first- order impact on lifetime earnings; rather, any effect of early-life health on adult earnings should primarily accrue through improvements in human capital independent of the level of schooling. Because we observe an adult's early-life mortality environment, his wages, and his level of schooling, we are able to test these two theoretical predictions. We find that exposure to better early-life health is indeed associated with earning higher adult wages. However, we find no similar association with education levels, and flexibly controlling for a detailed education vector does not influence the magnitude of the gradient we document between early-life health and adult earnings, as predicted by Bleakley (2010a).","In this section, we outline a strategy to quantify the gradient between the early-life mortality environment and adult wages. Although we write about “identification,” we do not interpret our results literally as an effect of early-life mortality outcomes on wages; rather, adult economic outcomes and infant mortality are both shaped by an early life health and disease environment, which includes sanitation and other dimensions of public health. We use a double fixed effects (place and time) identification strategy to compare workers who compete with one another within the same labour market today, but were exposed to better or worse disease environments and mortality regimes when they were born. Our identification strategy exploits two facts about a cross-section of workers of different ages who live near one another: First, their wages today are determined, in part, by a common labour market. Insofar as labour is substitutable across workers of different ages, those workers are offering to supply their labour to the same, shared demand side of the market.3 Second, workers of different age cohorts in a cross-section implicitly form a synthetic panel: workers of different ages today represent the effects of early-life health at the different points in history when they were born. We exploit district-by-time variation to investigate the association between adult wages in 2005 and early-life health in districts throughout India in the 1970s and 1980s. In particular, we match historical, district-level census data on early-life mortality with cross-sectional survey data on adult wages. Fig. 1 presents our identification strategy graphically. The graph plots average wages, net of district fixed effects, as a linear function of age. The most visible feature of the graph is the upward slope: within essentially all Indian districts, older workers are paid more than younger workers, on average. Our identification is found in the difference between the two slopes. Early-life health improved over time in essentially all districts. However, these improvements were not uniform across districts. In districts where the mortality environment improved more quickly over this period, younger workers would be expected to have relatively better early-life human capital than older workers in the same district, in comparison with the difference between older and younger workers in other districts where the health environment improved more slowly. Therefore, our identification strategy asks whether the positive gradient between ages and wages is less steep in districts where the mortality environment improved more quickly. An initial answer is visible in the difference in the two slopes in Fig. 1: the age profile of wages is less steep in the half of districts with above-average declines in infant mortality. Sources of historical and contemporary data ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We match data from two different sources. For our present-day dependent variables and control variables, we use data on individual adult males from the India Human Development Survey (IHDS), a nationally representative 2005 cross-sectional survey of 40,000 households (Desai et al., 2007). We study men born between 1971 and 1989, who were therefore between 16 and 34 years of age in 2005, leaving us with 12,783 observations, as can be seen in Table 1. These observations are drawn from 277 districts across 17 states, including the 13 most-populated states as of the 2011 Census which comprise over 85% of the total population; the full list of states covered by our data can be found in Online Appendix B.2.4 Our primary dependent variable is the log of hourly wages in rupees, as computed by the IHDS.5 As a robustness check and an input to our welfare computations, we also estimate the gradient between the early-life mortality environment and household consumption per capita. For our independent variable, we use historical infant mortality rates, a standard variable in the economic history and economic demography literatures. Infant mortality rates (IMR) are scaled as the number of deaths in the first year of life per 1,000 live births. Economic historians have long used infant mortality as a measure of the disease environment. Because of consequences of disease for net nutrition, early-life infant mortality rates are increasingly well understood to be an important determinant of height in developing countries today (Bozzoli et al., 2009), and historically in now-rich European countries (Hatton, 2014).6 Of course, we do not literally estimate effects of infant mortality: we would not expect average wages in a district to rise as a result of an emergency medical intervention that barely prevented the deaths of the marginally last infants to die. Instead, we interpret the gradient that we document between mortality and subsequent wages to reflect the influences of the health and disease environment on both. In the 1970s and 1980s, infant mortality would have been importantly shaped by infectious disease and maternal nutrition, rather than by perinatal medical care; the first recorded data on the percent of births attended by any skilled medical staff in the World Bank World Development Indicators is 34.2 percent in 1993. We match district-level historical IMR data from various rounds of the Census of India to the IHDS. Census data is available only at 10 year intervals, specifically in 1981, 1991 and 2001; such ten-year intercensal periods are standard in demographic data, and infant mortality data at the district level does not exist for any Census prior to 1981. We use these three Census rounds to estimate a long difference in IMR by district: we run a district-specific linear regression of census IMR on year for each district, and use regression coefficients to predict IMR for each year from 1971 to 1989, thus matching the adults we study with the predicted IMR in their district, in the year of their birth. Results are robust to the alternative use of log-linear regressions, or instrumenting for the linearly-predicted IMR with the log- linear prediction to deal with potential measurement error in infant mortality, as we will show.7 Given the 10-year interval between each Census round, we do not project IMR estimates any further into the past than 10 years prior to the 1981 Census, which determines the start date of our data as 1971. This procedure provides the best possible estimates of infant mortality rates by district over the 1971–89 period, but our estimates are not sensitive to the start date; a robustness check in Online Appendix D shows that our point estimate is numerically very similar – and even slightly larger in absolute value – if we omit all years prior to 1981. In a subsequent robustness and plausibility check, we use district-level sanitation rates, operationalized as the percent of households in a district who own a toilet or latrine, rather than defecating in the open. Early-life exposure to open defecation in India has recently been shown to be a significant predictor of infant mortality, childhood height (Spears, 2013) and childhood cognitive achievement (Spears and Lamba, 2016).8 We also take this historical data from the Indian Census from 1981, 1991 and 2001, where in this case data is available separately for the urban and rural portions of each district (when a district contains both). Rural open defecation is not observed in the 1981 Census, but the World Health Organization estimated that as of 1980, only 1% of India's rural population had access to any sanitation facilities, so we assign 100% rural open defecation in 1981,9 and predict sanitation rates for each observation by performing separate district-specific regressions for rural and urban sections of each district. As a result, we use a smaller, younger sample in the sanitation robustness check; rural open defecation was almost universal before 1981, and therefore there is no improvement to be studied. Table 1 presents summary statistics for the IHDS and census data that we use. Our average workers are poor – earning about 10 rupees an hour, or about $0.23 in 2004 dollars (roughly $1 at purchasing price parity) – and were exposed to threatening early-life health environments, with infant mortality of 113 deaths per 1000 births and only about 17 percent sanitation coverage, on average. The maximum estimated IMR in our sample is 281, and the minimum is 29.7; the average individual lives in a district in which the IMR declined by 2.8 points per year, but there is considerable variation in this measure, ranging from a 10 point reduction per year to a 1.8 point increase. The workers we study are also relatively young, with an average age of about 26 years. Empirical specification ~~~~~~~~~~~~~~~~~~~~~~~ It is important to note that any coefficient on IMRdt that we observe can only be consistent with a factor that changed over time within districts in parallel with improvements in early-life health, and which, in a contemporary cross-section, differentially impacts people who were born a few years apart; that is, people who would have been subjected in similar ways to changes in village infrastructure, education, or cultural norms. Because these wages are observed in a cross-section of workers of different ages, any difference across local labor markets would be absorbed by district fixed effects. Our specification thus rules out many forms of spurious correlation driven by factors other than early-life health. To further demonstrate the robustness of our strategy, as well as the stability of our coefficient estimate, we add controls Xidt in stages: state-specific linear time trends: identifies the effect of early-life health from the extent to which the district time trend in IMR differs from the state-wide time trend, to rule out any spurious state-level omitted variables; state × urban fixed effects: controls for a separate rural–urban difference in each state; state × social group indicators: state-specific indicators for eight caste and religion groups; female literacy: district-level female literacy in the district and year of the man's birth, matched from census data in the same way as IMR. The control for female literacy – another indicator of human development and an important determinant of early-life human capital – helps ensure that we are identifying off of variation in early-life health and the disease environment rather than improvements in other district facilities and outcomes, and also verifies that no spurious correlation is mechanically introduced by our district-level matching process. We will additionally add a further set of covariates which may entail overcontrolling, relative to a properly specified model, but which will allow us to further rule out omitted variable bias while investigating possible mechanisms of the effect we document. In particular, we will add indicators for the worker's membership in seven job categories,10 which rules out spurious structural differences in district labor markets. Finally, we will add detailed indicators for years-of-school interacted with literacy. These additions would be overcontrolling if education investments were partially caused by early-life health. However, as we have discussed, models of optimal investment in education suggest that improvements in early-life health could increase adult wages without having a large effect on schooling decisions. Moreover, schooling may not be an important mechanism of the translation of early-life health into adult wages. In a test of this theoretical prediction, we will show that our empirical strategy finds no effect of early-life IMR on education, and no impact on our main coefficient of interest when we include education variables in the regression. Our strategy implicitly assumes that the young adult men in our sample were born in the same district in which they lived at the time of our data. This is a reasonable assumption, because permanent migration for adult males in India is relatively uncommon (Rosenzweig and Stark, 1989); in contrast, women often migrate at the time of marriage to join their husbands’ households.11 We will demonstrate that our results are not affected by migration: permanent migration is observed in the IHDS, and excluding the small fraction of men who have ever moved residences does not meaningfully change our coefficient estimates.","Are men who were exposed to a better early-life health environment, as measured by infant mortality, subsequently paid higher wages as adults? As an initial answer to motivate our main result, Fig. 2 verifies that men who were born in district-years with worse infant mortality and sanitation earned lower wages as adults in the IHDS in 2005. Panels A and B plot locally weighted kernel regressions depicting a clear downward trend. Panels C and D plot residuals of wages against residuals of our health measures, in both cases after controlling for year-of-birth and state-times-urban fixed effects, with the means added back in to make the range of the figures comparable to A and B. The visible downward trend remains, in an initial suggestion of the gradient we will estimate. Main result: adult wages and early-life mortality rates ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 2 presents our main empirical result: men who were born in district-years with higher infant mortality have lower adult wages, on average. The regression coefficients imply that a 1 percentage point increase in the infant mortality rate (that is, 10 more infant deaths per 1000 live births) would be associated with a decline in adult wages of almost 2 percent.12 The mortality-wage gradient is notably stable across regression specifications. In particular, similar results are found if IMR is projected linearly across survey rounds, as in Panel A, or if linear prediction is instrumented for with log- linear projection, as in Panel B, to reduce measurement error. Adding a long vector of regression controls fails to importantly change the coefficient estimate. In particular, column 2 controls for state-specific linear year-of-birth trends (age gradients), separate urban indicators for each state, and separate religious and caste indicators for each state; the coefficient remains essentially identical, and if anything slightly increases. Column 3 goes further by including indicators for job categories, which may well be in part a consequence of early-life health and human capital accumulation (Vogl, 2014). This does not change the coefficient on early-life IMR exposure. This suggests that job categories are not omitted variables that are spuriously responsible for our main result and that early-life health does not appear to be related to wages through the mechanism of sorting into these categories.13 Column 4 tests Bleakley's envelope theorem observation: chosen schooling should not mediate the relationship between early-life health and adult wages. Controlling flexibly for education, measured as a vector of school-grade indicators interacted with literacy, has no effect on the coefficient on birth-year infant mortality. Rather, the link operates through improvements in human capital at the same level of schooling. Note that this is not merely because the schooling variables are too noisily measured to have a signal; in the regression of column 4 of panel A, the vector of education indicators is highly statistically significant with a test statistic of F31,276 = 19.69, p < 0.00001. Finally, column 5 controls for census female literacy rates in the district and year of the man's birth; although the sample decreases slightly because this cannot be matched to all district-year combinations, the stability of the coefficient suggests that the gradient we observe is due to the early-life health environment, rather than historical human development more broadly.14 A robustness check in Online Appendix D further demonstrates that our coefficient estimate is stable across different start dates for our data; in fact, the association between IMR and subsequent wages even becomes a bit stronger when data on individuals born in the 1970s is discarded. A robustness check in Online Appendix D further demonstrates that our coefficient estimate is stable across different start dates for our data; in fact, the association between IMR and subsequent wages even becomes a bit stronger when data on individuals born in the 1970s is discarded. Bleakley's optimization result: no effect on education ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Improved early-life health could increase adult human capital by increasing physical strength or cognitive ability directly, or by increasing attained schooling, which would in turn increase wages. However, Bleakley (2010a) predicts that the optimally chosen quantity of schooling should not be an important mechanism linking early-life health to adult wages, and that schooling may not even increase in response to an increase in early- life health, because health increases both the benefits and costs of a child attending school. We have seen evidence for the first claim; in Table 3 we test the second. Neither the early-life environment proxied by IMR nor the level of sanitation (which we will consider as a robustness and mechanism check in subsequent sections) are associated with a difference in schooling levels. To emphasize, this is not because the schooling variables are unreliable noise: they are quite significant predictors of adult wages.","In this section, we present three further empirical tests of the robustness of our main result and of the plausibility of a link between the early-life health environment and adult wages. First, we show that early-life exposure to open defecation – one important determinant of infant mortality in India – is similarly associated with adult wages. Then, returning to the concern that we observe only a man's district of residence and not of birth, we rule out migration as a cause of our result by showing that the result is similar when migrants are excluded. Finally, we document a similar association between the early-life health environment and household consumption, but only for households in which the man we study is the main earner. This will be an input into our welfare calculations, and is an important plausibility check for the consistency of our result with economic mechanisms, rather than it reflecting a spurious correlation. Early-life exposure to open defecation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Among potential health insults, exposure to poor sanitation is particularly likely to have quantitatively important consequences for adult economic outcomes because its effects on early-life health are so large. Water and sanitation are known to be important determinants of health outcomes, especially infant mortality (for recent examples in economics, see Cutler and Miller, 2005; Galiani et al., 2005; and Watson, 2006). Nonetheless, the associations between human capital and sanitation and disease have received relatively less research attention than human capital gradients in income and education. However, a recently active literature has shown that sanitation can be particularly important in human capital accumulation in developing countries, especially in India, where open defecation – without using a toilet or latrine – is particularly widespread (Spears, 2013). Open defecation matters for health because it releases fecal germs into the environment which cause disease in growing children. According to Unicef and WHO (2012) statistics, over a billion people worldwide defecate in the open; most of these live in India, and most people who live in India defecate in the open. As a large medical and epidemiological literature documents, ingestion of fecal pathogens as a result of living near poor sanitation is well-known to cause diarrhea (Esrey et al., 1992). Checkley et al. (2008) use detailed, high-frequency longitudinal data from five countries to demonstrate effects of childhood diarrhea on subsequent height. In addition to the obvious threat of diarrheal disease, open defecation can cause net nutritional insults through worm or other parasitic infections, or by increased energy consumption fighting disease. Most recently documented in detail in the medical literature, but perhaps very important, is the possibility of widespread chronic but subclinical environmental enteric dysfunction (Humphrey, 2009). Economists have further identified effects of childhood exposure to sanitation-related disease on human capital (Bleakley, 2007; Baird et al., 2011). Substituting sanitation coverage for infant mortality as the independent variable in our regressions allows us to provide further evidence of a link between the early-life disease environment and adult wages. Open defecation is only one of many important causes of historical and contemporary infant mortality in India, and as such sanitation and IMR are far from perfectly negatively correlated: among individuals in our baseline wage regression for whom both IMR and sanitation data are available, the correlation is −0.196. In fact, the correlation between the changes in IMR and sanitation that we use for identification is essentially zero: in fact, it is slightly positive at 0.0637, because districts that saw the greatest reductions in IMR tended to be the districts with the worst disease environment, and thus with the highest starting IMR, whereas sanitation in those districts tended to improve more slowly.15 Therefore, the regressions below make use of a second, and quite different, source of variation in early-life health environment, making it less likely that both are simply capturing some other unobserved variable. Thus, insofar as results using sanitation are broadly similar to results found using IMR, we interpret this concordance as indicative that the mortality results are plausibly consequences of the disease and health environment. Table 4 reports regression results with sanitation coverage as the independent variable. The sample is smaller in Table 4 than in the IMR analysis because we only use data on individuals born in 1981 and after, as rural open defecation was almost universal before this period, meaning there was no improvement to study; this smaller sample will decrease the precision of coefficient estimates. As the table shows, we indeed find a gradient of important but plausible magnitude between early-life sanitation and adult wages. The main independent variable is the percent of households owning a toilet or latrine, rather than defecating in the open. A 10 percentage point decrease in open defecation translates into an approximately 2–3 percent increase in wages. Thus, we find that in districts where sanitation has improved more quickly over time, the within-district wage profile is less steeply increasing in age. As before, our result is quantitatively stable as a long vector of controls is added, including for state-specific time trends, caste and religious groups, and state-specific urban residence. Also as before, adding job categories and education indicators – although these are, themselves, predictive of wages – does not change the coefficient on early-life sanitation.16 Crossing these produces 15 predictions of the coefficient for wages regressed on early-life sanitation coverage. These range from 0.0006 to 0.0052, with a median of 0.0017 and a 75th percentile of 0.0021. This is exactly the neighborhood of our estimates in Table 4: 0.0018–0.0032. Although our estimates are slightly above the median of these estimates, there is reason to suspect our estimate would be larger: a higher fraction of the variation in height in India reflects early-life health than in the U.S. or likely even than in Mexico. Indeed, Spears (2012b) finds that the gradient between height and cognitive achievement is much steeper for Indian children than for U.S. children. Therefore, we conclude that our estimates are quantitatively consistent with predictions from estimates in the literature. Results are not driven by migration ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Because we observe the district in which men live, not the district in which they were born, certain patterns of endogenous migration could, in principle, bias our estimates.18 Fortunately, the IHDS includes data on migration, so we can assess the importance of this concern. The data report whether a particular individual moved homes since their birth, though we cannot see where they migrated from, so we are likely to overstate migration as many moves would be within the same district. We run the regressions of wages on IMR and sanitation only for stayers, by omitting all individuals who report ever having moved (even those who move within districts). The results are presented in Table 5, and the coefficients are nearly unchanged. Further analysis in Online Appendix C finds no evidence that selective migration along the dimension of improvements in early-life health occurred in any case. Therefore, selective migration does not appear to be responsible for our results. Effects on consumption ~~~~~~~~~~~~~~~~~~~~~~ In this section, we extend our analysis to check for a gradient between early-life mortality and household consumption per capita. In part, this is a robustness check of our main results: we would expect an increase in income to increase household consumption, especially if the adult male we study is an important source of household income. Additionally, these estimates will be used in our fiscal and welfare computations: consumption taxes – such as value added tax – are a larger fraction of government revenue in India than income tax, so if we are concerned about the fiscal impacts of the early- life health environment, it is important to confirm that consumption is affected as well. Table 6 documents the association between early-life IMR and the log of household monthly consumption per capita.19 Column 1 repeats the estimate of the gradient between early-life IMR and wages from Table 2. Column 2 shows that the association with household consumption is of similar magnitude, although slightly smaller. Importantly, however, the adult men whom we are able to study are relatively young, and only some of them will be significant earners for their households. One common family structure in India is a joint household where adult men and their spouses live with the man's parents; such households could have multiple brothers and a father earning income. Column 3 restricts the sample to the approximately two-thirds of the men we study who earn the most money of all people in their household; the effect is quantitatively similar to the effect on wages in column 1. Column 4 presents results for men who are not main earners; their early-life health environment has no detectable effect on their households’ consumption. This non-finding is important because it is consistent with what the economic demography of the Indian context would predict; this therefore suggests that our finding is not merely a spurious reflection of correlation between some aspect of households’ socioeconomic status today and the health environment in their districts in past decades.20 For example, our results are not driven by factors that impact the entire family: if IMR or sanitation in birth year was simply capturing general improvements in village infrastructure, the impact should be felt by all earners today, not just those families where the current primary earner was an infant at the time of the improvement.","Our analysis so far has produced coherent evidence that improvements in the mortality environment are associated with higher subsequent wages and consumption. As a result, we expect that such improvements would have positive consequences for the tax revenues collected by the Indian government. Importantly, this suggests that investments in improved early-life health, such as investments that lead to increased use of improved sanitation, could come at a low net fiscal cost to the government of a country such as India. It is also likely that higher income and consumption will translate into increased welfare for Indian households. However, this is not certain in a context of non-unitary households; Indian households are often large and complex, and it is beyond the scope of our analysis to evaluate who receives the increase in consumption within a household, and how that might affect intra-family relations or bargaining power. The IHDS does not observe person-level consumption, only household-level. Thus, we can evaluate the impact of improvements in the early-life disease environment on household consumption, and aggregate the gains up to an economy-wide level, and for simplicity we will refer to these as welfare gains; but it should be understood that we do not claim that increases in household consumption can be monotonically translated into gains in actual person-level welfare. Our estimated consumption gains could more accurately be interpreted as increases in potential household well-being, while the actual welfare gains could, in principle, be larger or smaller than what we estimate. Therefore, in this section, we translate the empirical estimates from the previous sections into fiscal and welfare terms, to provide an illustration of the aggregate impact of early-life health on the Indian economy. For example, if a 1% point reduction in IMR is associated with 1.74% higher wages, we use details of the Indian tax system to estimate the associated increase in future tax revenues, as well as the increase in after-tax income and thus consumption. In each case that we study, the gains in tax revenue and consumption at an aggregate level are large, at least $10 billion in present-value terms; these results demonstrate that improvements in early-life health are associated with substantial gains to the government and to the Indian population. Independent of the accuracy of our main empirical estimates, this section is important in the context of Acemoglu and Johnson (2007), Bleakley (2010a), and others for computing what moderate microeconomic relationships of the sort that we estimate could add up to in a large economy. Fiscal externalities of early-life health ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We begin with an analysis of the association between early-life mortality environment and tax revenues, using the result from column 1 of Table 2 as our baseline estimate: a 1% point increase in IMR is associated with 1.74% lower wages. We then assume that the impact of IMR on tax revenues is also 1.74%; thus we conservatively assume that the income elasticity of tax revenue is 1, even though studies of both developed and developing countries tend to find that increases in income lead to proportionately greater increases in tax revenues.21 We limit our attention to the income tax and excise and service taxes, as these are taxes which depend directly or in a close indirect way on income and consumption; we ignore customs duties as well as the corporate tax, even though one might expect more productive workers to lead to larger corporate profits. The revenue from these taxes amounted to about 5.11 trillion rupees in 2012–13, or $93.64 billion US;22 we assume that a normal working life is 40 years, from 18 to 57, and assume that each year-of-birth cohort produces an equal share of the tax revenue, or $2.34 billion per cohort, prior the change in IMR being considered.23 Because these revenue gains occur in the future, starting when the year-of-birth cohort born today enters the labour market, and gradually phasing in after that as more “treated” cohorts enter, we can use a 3.81% discount rate24 to add up these future gains and express the revenue gains as a present-value equivalent. This simple procedure allows us to calculate the expected effect on tax revenues from a 1% point reduction in IMR starting today and for each of the next 100 years. Reducing IMR by 1% point produces fiscal gains starting 18 years from now when the first treated cohort enters the labour market, and ending 157 years from now when the final cohort exits the labour market; the details of the calculations are relegated to Online Appendix E.1, but simply adding up the revenue gains from each cohort and discounting, we find that the sum of the revenue increases is equivalent to $11.70 billion in present value terms.25 As a sensitivity analysis, we have also evaluated the revenue gains for the highest and lowest estimates in panel A of Table 2, and for a range of values for the discount rate; the results are displayed in panel A of Fig. 3. The revenue gain depends on both the association between IMR and wages and the discount rate, but especially on the latter, ranging from about $6–9 billion with a rate of 5% to as much as $47 billion with a 2% discount rate. The black dotted line in the figure shows the baseline discount rate of 3.81%. As a further illustration of the potential fiscal gains from investments in improving early-life health, we also consider the gain in tax revenues associated with a more specific public health objective: the elimination of open defecation today, using our estimate of the association between sanitation and adult wages in Table 4. It was estimated that 53.1% of Indian households defecated in the open in 2011, but that number had been declining at an average rate of 1.05 percentage points per year over the previous decade. Therefore, when considering the elimination of open defecation, the appropriate counterfactual is one in which open defecation continues to decline over time; we assume a continued decline at the same linear rate, so that absent any intervention open defecation would be eliminated in about 50 years. Then, using the result from column 1 of Table 4 as our baseline estimate, where a 1% point increase in sanitation coverage is associated with 0.296% higher wages, we perform a calculation similar to the one above and find a total present-value revenue gain of $60.48 billion; details can again be found in Online Appendix E.1. As a further illustration of the potential fiscal gains from investments in improving early-life health, we also consider the gain in tax revenues associated with a more specific public health objective: the elimination of open defecation today, using our estimate of the association between sanitation and adult wages in Table 4. It was estimated that 53.1% of Indian households defecated in the open in 2011, but that number had been declining at an average rate of 1.05 percentage points per year over the previous decade. Therefore, when considering the elimination of open defecation, the appropriate counterfactual is one in which open defecation continues to decline over time; we assume a continued decline at the same linear rate, so that absent any intervention open defecation would be eliminated in about 50 years. Then, using the result from column 1 of Table 4 as our baseline estimate, where a 1% point increase in sanitation coverage is associated with 0.296% higher wages, we perform a calculation similar to the one above and find a total present-value revenue gain of $60.48 billion; details can again be found in Online Appendix E.1. The purpose of these calculations is to illustrate the potential quantitative economic importance of early-life health. If we take our results literally as capturing the causal effect of sanitation coverage on wages, then this implies that if there existed an investment capable of eliminating open defecation today at a cost of $60.48 billion or less, there would be no net cost to the Indian government, as those expenditures would be made up in future tax revenues, even after those future gains were discounted at a 3.81% rate. To provide an estimate of the revenue gain per unit of investment (households induced to use latrines), we divide this total by the number of households currently estimated to defecate in the open, which is approximately 131 million, and find that the revenue increase is $462 per household that is induced to stop defecating in the open. In panels B and C of Fig. 3, the range of values generated by trying different discount rates and using the upper and lower bounds from Table 4 are displayed; the gains per household are around $200 at the low end and over $1100 at the top end. These revenue gains associated with improvements in the early-life public health environment are substantial, and on top of the numerous conservative assumptions made earlier, we have ignored other potential sources of fiscal gains, such as reduced public health care expenditures and calorie requirements if sanitation investments lead to improvements in health among the affected population. Additionally, in Online Appendix F we present the results of quantile regressions, which show that the association between the early-life health environment and wages tends to be larger towards the upper end of the income distribution; as a result, in Online Appendix G.1, we show that the estimated fiscal benefits using the quantile regression results are even larger than those presented here. These revenue gains associated with improvements in the early-life public health environment are substantial, and on top of the numerous conservative assumptions made earlier, we have ignored other potential sources of fiscal gains, such as reduced public health care expenditures and calorie requirements if sanitation investments lead to improvements in health among the affected population. Additionally, in Online Appendix F we present the results of quantile regressions, which show that the association between the early-life health environment and wages tends to be larger towards the upper end of the income distribution; as a result, in Online Appendix G.1, we show that the estimated fiscal benefits using the quantile regression results are even larger than those presented here. Consequences for household economic well-being ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Increases in wages associated with improved early-life health not only raise tax revenues; they should also lead to higher after-tax income and consumption. In this subsection, we attempt to quantify these gains in household economic well-being, where, as stated at the beginning of this section, we loosely interpret increases in household consumption as gains in welfare. As throughout the article, we abstract from all benefits of better health and more physically and mentally capable citizens, as we are unable to measure them, and focus only on the gains from higher consumption. We consider the increase in consumption associated with the same changes as in the fiscal calculations: a 1% reduction in IMR today and for the next 100 years, and the elimination of open defecation. Per- capita GDP was estimated to be $1219 in 2010–11, and given that the 2011 Indian Census finds that about 40% of the overall population are workers,26 this implies an average income of $3063 for employed individuals. Tax revenue was estimated to be 10.39% of GDP in 2011, according to the World Bank, so we use a net-of-tax rate of 0.8961, implying after- tax income of $2745 for the average employed individual, and we continue to use 3.81% as the annual discount rate, although now this should be understood as either a personal or social rate of time preference. In our baseline estimates, each 1% point reduction in IMR is associated with an increase in wages of 1.74%, implying a $47.76 per year increase in household consumption per average worker. The details of the calculations can be found in Online Appendix E.2, and the present-value increase in welfare is equivalent to $165.15 billion of consumption; to put this number in context, it is equivalent in welfare terms to a $6.06 billion (about 0.3% of GDP) increase in annual consumption now and for every year in the future. Panel A of Fig. 4 displays the robustness of this result to varying estimates and discount rates, confirming a significant welfare gain that ranges from about $86 billion to as much as $669 billion. In our baseline estimates, each 1% point reduction in IMR is associated with an increase in wages of 1.74%, implying a $47.76 per year increase in household consumption per average worker. The details of the calculations can be found in Online Appendix E.2, and the present-value increase in welfare is equivalent to $165.15 billion of consumption; to put this number in context, it is equivalent in welfare terms to a $6.06 billion (about 0.3% of GDP) increase in annual consumption now and for every year in the future. Panel A of Fig. 4 displays the robustness of this result to varying estimates and discount rates, confirming a significant welfare gain that ranges from about $86 billion to as much as $669 billion. As an alternative robustness check, we can also use our estimates of the association between IMR and consumption directly. Average per capita consumption was 1430 and 2630 INR per month in rural and urban areas in the 2011 NSS; since 68.84% of India is rural, this implies average consumption of 1803.92 INR per month, or $458.41 per person per year. Each 1% point decrease in infant mortality is associated with an increase in average consumption by 1.73% for everyone in a household with an affected main earner, or $7.93 per person per year. Adding up these gains as described in Online Appendix E.2, we find a total welfare gain of $68.91 billion in present value terms associated with 10 fewer infant deaths per 1000 births. This value is smaller than the one calculated from the wage regressions, which should not be surprising as the estimates of per capita consumption in the NSS are considerably smaller than the after-tax value of per-capita GDP; however, a gain of this magnitude is still economically very significant. As an alternative robustness check, we can also use our estimates of the association between IMR and consumption directly. Average per capita consumption was 1430 and 2630 INR per month in rural and urban areas in the 2011 NSS; since 68.84% of India is rural, this implies average consumption of 1803.92 INR per month, or $458.41 per person per year. Each 1% point decrease in infant mortality is associated with an increase in average consumption by 1.73% for everyone in a household with an affected main earner, or $7.93 per person per year. Adding up these gains as described in Online Appendix E.2, we find a total welfare gain of $68.91 billion in present value terms associated with 10 fewer infant deaths per 1000 births. This value is smaller than the one calculated from the wage regressions, which should not be surprising as the estimates of per capita consumption in the NSS are considerably smaller than the after-tax value of per-capita GDP; however, a gain of this magnitude is still economically very significant. Meanwhile, a 1% point reduction in open defecation is associated with a 0.296% increase in wages, which translates into a $9.07 per year increase in family consumption for an average worker. This implies that the elimination of open defecation would be associated with total discounted gains of $4653 for each worker born today; this is equivalent to nearly four times the current GDP per capita, or about 71.7 years of maximal annual earnings from NREGA, a large government workfare program. Over the entire workforce, the total present- value after-tax income gains over the next 100 years or so are $853.99 billion; as before, this can be expressed as a yearly increase in consumption now and every year in the future, and in those terms it amounts to $31.3 billion per year, or about a 1.7% increase in the current GDP of India. Panels B and C of Fig. 4 displays the robustness of these results; the gains are large at all combinations of parameters, and reach as high as about $2.1 trillion in total, or $9197 per individual born today. And to emphasize again, all of these estimates of “utility” impacts refer only to utility from increased consumption, and not from any other benefits of improved health, cognitive achievement, or mortality; of course, we also abstract from the complications of calculating welfare in a non-unitary household. Finally, Online Appendix G.2 evaluates the welfare gains associated with improvements in IMR and sanitation using the results of the quantile regressions from Online Appendix F. The quantile regressions suggest that the effect is stronger at high incomes, and if this is true then the welfare gains should be smaller than in the baseline analysis if marginal utility diminishes as income increases. Accordingly, using log utility, we find that the welfare gains are all smaller than those discussed above, but always highly economically significant, with present-value gains of $33–$74 billion from a 1% point IMR reduction and $412 billion from the elimination of open defecation. Finally, Online Appendix G.2 evaluates the welfare gains associated with improvements in IMR and sanitation using the results of the quantile regressions from Online Appendix F. The quantile regressions suggest that the effect is stronger at high incomes, and if this is true then the welfare gains should be smaller than in the baseline analysis if marginal utility diminishes as income increases. Accordingly, using log utility, we find that the welfare gains are all smaller than those discussed above, but always highly economically significant, with present-value gains of $33–$74 billion from a 1% point IMR reduction and $412 billion from the elimination of open defecation. Whichever set of estimates or procedure for calculating welfare is used, the estimated gains associated with improvements in early-life health are potentially very large.","This article documents a robust gradient between the early-life health environment and adult wages, decades later. Exploiting heterogeneity across Indian districts in the time- paths of improvement in infant mortality, we find that men exposed to a better early-life health environment earned significantly but plausibly higher wages as adults. The estimated gains are similar across a wide range of specifications and sets of fixed effects; are replicated in an analysis of historical changes in sanitation; and are consistent with finding improvements in consumption of similar magnitude, precisely when the man for whom we have data is the household's main earner. These results are not spurious consequences of selective migration, which we can observe in our data. Our results are not driven by changes in education; this is consistent with Bleakley's Envelope Theorem prediction and with evidence from Cutler et al. (2010) on the effects of early-life malaria exposure on subsequent consumption in India. Moreover, the apparent unimportance of education to the health-wages relationship may make certain omitted variable threats less likely, such as coincidental other improvements in education or other human capital facilities in the same districts. These findings suggest that early- life exposure to infectious disease could have appreciable consequences for economic outcomes in developing countries, especially in contexts such as India's, where relatively high early-life mortality rates and exposure to high levels of open defecation both continue today. Relatively modest effects of early-life health on wages – of the magnitude of the gradients that we estimate – could add up to important fiscal consequences. Because improving health raises wages and consumption when children become adults, reductions in infant disease today causes positive fiscal externalities in the future. Wage gains occur decades after improvements in the early-life health environment, so the present-day benefit depends on the interest rate. Our results indicate that public investments to improve the early-life health and disease environment – potentially including efforts to reduce exposure to open defecation – could improve well-being at a low net present cost to the government. Because such public investments are under active public debate in India – where the Prime Minister has announced an ambitious plan to eliminate open defecation by 2019 – these results are of clear policy importance. Nevertheless, we must acknowledge some important limitations of our analysis. First, because we do not observe the mechanism assigning different Indian districts to improvements in infant mortality and sanitation at different times, and because data limitations force us to use long-term trends in these improvements, we cannot fully confirm the exogeneity of these changes with respect to the outcomes we study. Second, because we are matching census data with survey data decades later, and do not longitudinally track individual children as they become adult workers, we cannot observe mechanisms in a lifetime of health and human capital measurements. However, Spears and Lamba (2016) have recently documented an effect of early-life exposure to open defecation in India on later-childhood cognitive achievement, while Spears (2012b) demonstrates that Indian children who are taller (due in part to better early-life health and net nutrition) also perform better on learning tests. These prior findings suggest that our results are plausible, and that cognitive development is one important mechanism in translating better early-life health into future achievement. Finally, our welfare analysis of quantitative policy implications must assume a unitary household model, because our data source does not allow us to observe the consumption of individual members of the household. Despite these limitations, our results suggest an important role for early-life health in adult economic outcomes in India."],["We examine the effects of employee and employer social security contributions (SSCs) on labor cost, hours of work, and labor cost per hour, using a long running panel dataset that allows us to exploit 35 years of policy reforms in the United Kingdom. We find that reductions in marginal rates of employee – but not employer – SSCs have positive effects on labor cost that operate through hours of work, while labor cost falls much more when average employer SSCs rates are reduced than when average employee SSCs rates are reduced, with most of this differential effect coming through reductions in hourly labor cost. We interpret this as evidence that employees change their hours in response to SSCs, but that in the short- to medium-run at least, the formal incidence of SSCs can matter for their behavioral impacts and economic incidence. --------------------------------------------------------------------------------","Social security contributions (SSCs) make up a large part of the labor tax wedge in advanced economies, typically raising more revenue than the personal income tax.1 And, unlike income taxes, they are typically levied on both employers and employees, and only on labor earnings. Yet the large literature on the elasticity of taxable income (ETI) that emerged from the seminal papers of Feldstein (1995, 1999) focuses mainly on the response of overall taxable income to income tax changes, rather than of labor earnings to SSCs, and it generally takes as given that the economic incidence is on the individual taxpayer. Meanwhile a parallel literature has developed investigating the incidence of SSCs which generally ignores the possibility of non-hours responses (e.g. effort) at the heart of the ETI approach, and sometimes does not account for any behavioral response at all, instead interpreting changes in earnings following SSCs reforms as evidence of incidence. Moreover, surprisingly few studies directly address the question of whether the effects of employer and employee contributions differ. In this paper, we exploit 35 years of reforms to SSCs in the United Kingdom and a long-running panel dataset on the earnings and hours of employees to provide evidence on the effects of employer and employee SSCs, considering both taxpayer responses and incidence. Specifically, these data and policy reforms allow us to utilise a panel regression approach to contribute to the literature in three interrelated ways. First, we contribute to the ETI literature by estimating the effects of SSCs on individuals' labor earnings, which is relatively under-studied. Second, we bring together the ETI and incidence literatures by examining what can be said about both incidence and behavioral responses, aided by being able to separate out effects of both marginal and average SSC rates on both hours of work and hourly wages. And third, we assess whether the economic incidence of, and behavioral responses to, SSCs are affected by their formal incidence, exploiting reforms that provide independent variation in employee and employer SSC rates. Our results suggest that the effects of employee and employer SSCs differ significantly, at least for the first year or so. Marginal rates matter for employee but not employer SSCs: we estimate a compensated elasticity of labor cost (that is, earnings + employer SSCs) with respect to the marginal employee net-of-SSC rate of around 0.2–0.3 – a response that operates mainly through hours of work – whereas marginal rates of employer SSCs have no significant effects. In contrast, average rates of employer SSCs affect labor costs much more than average rates of employee SSCs, and most of this effect comes via hourly labor costs rather than hours of work. We interpret this as evidence that employees change their hours in response to SSCs, but that – at least in the short-to-medium run – the formal incidence of SSCs affects their behavioral impact as well as their economic incidence. Within the ETI literature, a small number of papers do examine the responsiveness of specifically labor earnings to income taxation, finding lower elasticities than are typically estimated for overall taxable income.2 But few studies have attempted to estimate the elasticity of taxable earnings with respect to the SSC rate. Two recent working papers, Tazhitdinova (2015) and Adam et al. (2017), find levels of bunching of employees' earnings at SSC thresholds in the UK that are consistent with small behavioral elasticities for SSCs too.3 Closer to the present analysis, Lehmann et al. (2013) estimate the response of labor cost to changes in marginal and average rates of income tax and employer SSCs in France using panel data. They find an elasticity with respect to the marginal net-of-income-tax rate of 0.2, but virtually no effect of the marginal employer SSC rate. They find an elasticity with respect to the average net-of- income-tax rate of −0.44 (although this is statistically insignificant), and an elasticity with respect to the average net-of-employer-SSC rate of −0.866. Our results for labor cost show a broadly similar pattern. Lehmann et al.'s preferred interpretation of their findings is that wage stickiness prevents employer SSCs from being shifted to employees in the short term – statutory incidence matters4 – and that employees' labor supply is reduced by higher marginal rates of income tax. Yet responses to average net-of-tax rates can pick up income effects as well as incidence. And if wages are determined by a bargain between employers and employees, changes in marginal rates of tax, holding average rates fixed, can lead to changes in negotiated wage rates, and hence labor cost, even if there is no change in labor supply (Lockwood and Manning (1993); Pissarides (1998)). Thus responses to marginal tax rates may pick up (changes in) the incidence of SSCs as well as behavioral responses. Our data allow us to shed further light on the relative roles of behavioral responses and incidence. We decompose labor cost into hours of work and labor cost per hour, and perform separate regressions for each. Estimates from these hours and hourly labor costs regressions show that the effects of marginal rates of employee SSCs operate through hours of work, while most of the effects of average employer SSCs rates on labor cost come through hourly labor cost. We consider this stronger evidence that employees respond (on the intensive margin) to employee SSCs, and that the formal incidence of SSCs matters for their behavioral impacts and incidence in the short-to- medium-run time frames examined by us and Lehmann et al. Understanding incidence is key for the interpretation of ETIs. An implicit maintained assumption in most of the ETI literature that the incidence of taxes is on the individual in question is what allows estimated coefficients on marginal and average net-of-tax rates to be interpreted as individual substitution and income effects, respectively. Following Diamond and Mirrlees (1971), Saez et al. (2012a) point out that shifting of the burden of taxation between employees and employers does not affect its overall efficiency cost, nor optimal tax rates. But the efficiency cost implied by employees' earnings responses to taxation does depend on whether the effect on earnings reflects incidence or underlying behavioral change. If a change in earnings reflects shifting of the tax burden rather than a change in the amount of work being done, then the change in the employee's earnings will be balanced by an opposite change in their employer's income (or the income of whoever the burden is ultimately shifted to/from), which will usually also be taxed. In that case the change in employees' taxable earnings does not give the full picture of the change in all tax bases, or the overall fiscal externality, which is relevant for assessing the economic efficiency effects of tax policy.5 While the ETI literature has mostly sidestepped questions of incidence, a separate literature on the incidence of labor taxation has evolved in parallel which mostly sidesteps the issue of non-hours-of-work labor supply responses that are at the heart of the ETI literature, and often ignores hours responses as well.6 As a result, there is a risk that behavioral responses confound estimates of tax incidence. For example, Gruber (1997) and Anderson and Meyer (1997, 2000) find that firms seeing bigger reductions in employers' SSCs (as a result of reforms in Chile and the United States respectively) increased their employees' earnings roughly one-for-one, while employment in the firms was unaffected. They interpret this as evidence that the incidence of employer SSCs in these cases is largely, or fully, on workers.7 But it could be the case that part of the change in observed earnings instead reflects increases in effort or hours induced by the SSC rate changes. Putting aside this issue, on average studies find two-thirds of the incidence of labor taxes to be on employees, according to the meta- analysis in Melguizo and Gonzalez-Paramo (2013),8 but there is a very wide range of estimates.9 This literature, however, tends to focus on the effects of changes in employer or combined SSC rates; relatively few studies directly examine whether the incidence of employer and employee taxes are the same. This probably reflects the fact that many countries' reforms to SSCs affect employee and employer contributions simultaneously and highly co-linearly, making identification of separate effects difficult. By utilising a long time period in which there were separate reforms to employee and employer SSCs, we can examine whether responses to them differ – one of the first studies to investigate this using micro-data. The few existing studies that address this question are mainly based on cross-country regressions (such as OECD (1990) and Arpaia and Carone (2004)); they find – as we do – evidence that the incidence differs in the short-term, perhaps reflecting short-term stickiness of nominal wages. Whether the incidence converges in the longer run – as standard theory predicts – is less clear, in part because of weak statistical power in the long-run analyses that examine this issue (CPB et al., 2015). As with most of the incidence literature, we use ‘incident on employers’ as a shorthand to mean ‘incident on someone other than the employee whose tax rate changed’. Of course, the ultimate burden of a tax must always fall on real people, not the businesses or organisations employing them; it may be passed on to the employers' owners, customers and/or suppliers, and thence perhaps more widely via general equilibrium responses. One important possibility is that a tax that is not incident on the employee whose tax rate changes may nevertheless be incident on a broader group of workers: the nature of market responses may be such that a tax change affecting one small group of employees results not in a large change in wages for those employees but in a small change in wages for all employees in the firm, or in an infinitesimal change in equilibrium wages in the wider market. All we attempt to discern in this paper is how an individual's wage is affected by the tax rate applied to that individual's earnings; insofar as their net wage is not reduced one-for-one then we refer to the tax as being at least partly incident ‘on the employer’, even though the burden may be felt by a wider group of employees rather than by (say) the employer's shareholders. The paper proceeds as follows. Section 2 describes the UK's SSC regime, the reforms we use to identify its effects, and the data. Section 3 discusses the identification of behavioral responses and incidence, while Section 4 sets out our econometric approach. Descriptive statistics and estimation results are presented and discussed in Section 5. Section 6 concludes. Overview of the UK's system of SSCs ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Like SSC systems in most countries, the UK's system of National Insurance contributions (NICs) consists of both employer and employee contributions, which are functions of the employee's gross earnings.10 The NICs rate schedule changed markedly during the period we analyse (1982 to 2015), and it is those changes that we exploit in this paper. Throughout this period, earnings below a threshold were exempt from NICs. Before 1985, once earnings exceeded this threshold (then called the Lower Earnings Limit, LEL), employer and employee NICs were levied at constant percentage rates on all earnings (including earnings below the threshold) up to a ceiling called the Upper Earnings Limit (UEL). This created a jump, or notch, in contributions at the LEL: both marginal and average rates of NICs increased from zero to the rate payable above the LEL, which at the start of 1985 stood at 9% for employees and 10.45% for employers. A reform in October 1985 replaced this single large notch with a series of smaller notches: the jump in marginal and average rates at the exemption threshold was reduced to 5% each for employees and employers, and a number of graduated steps were introduced, where higher (marginal and average) rates applied to the entirety of earnings once earnings exceeded higher thresholds. At the same time, the cap on employer contributions was abolished so that the highest rate of employer NICs applied to earnings above the UEL. The effect of this reform on the combined headline employee and employer NICs schedule can be seen in the black and solid grey lines in Fig. 1. In October 1989, the system of graduated employee contributions was replaced by a single small notch at the lowest threshold (equivalent to 2% of the threshold) and a single 9% marginal rate of employee NICs that applied to earnings between the threshold and the UEL. However, the graduated system of employer NICs with four notches remained in place at that stage until April 1999, when the remaining notch in the employee NICs schedule, and all the notches in the employer schedule, were removed. There was also a series of other – generally smaller – changes in other years.11 The combination of changes in both thresholds and rates, the move from a notch-based system to a kink-based system via a series of smaller notches, and the extension of NICs (particularly employer NICs) above the UEL, provides a rich source of variation across the earnings distribution. Sometimes employer and employee NICs rates changed together, but other times they changed differentially. Some reforms affected individuals' marginal and average NICs in similar ways, but other reforms affected them very differently. This variation allows us to separately identify responses to changes in both marginal and average rates of both employee and employer NICs. Two additional features of the NICs system are relevant for our analysis. The first is that, unlike in many social security systems, the link between NICs paid and benefit entitlements was limited during our period of analysis. This means that NICs can be thought of largely as a straightforward earnings tax rather than ‘true’ social insurance. However, there was one strongly contributory element of the National Insurance system: the earnings-related component of the state pension. Individuals contributing to a private pension could choose to forgo this future entitlement and in return pay lower rates of employee and employer NICs on their earnings: a process known as ‘contracting out’.12 Prior to 1997 our data do not record whether people were contracted in or out of the earnings-related state pension, but for the purposes of our analysis we assume that everyone was contracted out throughout. This is partly because, over the period as a whole – and especially in the years that provide most of our identifying variation – the majority of employees were contracted out. But more importantly, even for those employees who were contracted in, the contracted-out rates are arguably a better measure of the ‘tax wedge’ associated with NICs: the additional NICs associated with ‘contracting in’ generated additional future pension entitlements on a roughly actuarially fair basis, so resemble retirement saving more than a tax insofar as people value these future entitlements. In any case, the reforms that generate our variation in NICs rates apply to both contracted-in and contracted-out rates. And crucially, none of the changes in contribution rates over our period were associated with a corresponding change in benefit entitlements. When an individual in our data sees their NICs rate change, it simply changes their current budget constraint, just like a tax; they do not see changes to the implicit savings or insurance which they might value. The only reason we might expect people to respond to the NICs changes we use differently from any other tax on earnings is if they (incorrectly) perceive it differently. The second feature is a reduced rate of employee NICs payable by married women: 2% at the start of our period, rising to 5.85% by the end. This was (and remains) available, in exchange for reduced benefit entitlements, to married women who have been claiming it almost continuously since May 1977. In 1979–80, 4.2 million women were eligible for this rate, falling to 560,000 in 1995–96 (of whom 340,000 were making use of it) and fewer than 80,000 by 2000–01. Since our data do not allow us to identify women eligible for this rate, we assume all women face the standard (contracted-out) rate of NICs. For this reason, comparisons of our estimates between men and women should be treated with caution. However, our main findings hold when we restrict our sample to men only – who are not affected by this issue. Income tax, means-testing and the National Minimum Wage ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In addition to NICs, employees' earnings are also subject to income tax, which until 1990 was usually assessed on joint income for married couples. Our data do not capture whether an employee is married, or the income of an individual's spouse, meaning we cannot account for income tax in our analysis.13 Similarly, means-tested benefits and tax credits may also contribute to the effective tax wedge on an individual's earnings, but such entitlements always depend on (married or unmarried) couples' joint income and other characteristics (such as housing costs and the presence and number of children) that are not observed in our data. If changes to these other elements of the tax wedge were correlated with changes to the NICs schedule, they might confound estimates of the effect of SSCs on labor cost or hours. Likewise, the introduction of a National Minimum Wage (NMW) in October 1999 and subsequent increases in it may confound estimates of the effects of NICs rate changes if the two are correlated. The NMW was initially set at around 45% of median hourly pay and directly affected 3.4% of employees, rising to 52% and 5.2%, respectively, by April 2015 (Low Pay Commission, 2016). The NMW may also affect the incidence of NICs for low-paid workers by putting a floor below which employers are legally not allowed to reduce hourly wages. In light of these factors, we test the sensitivity of our estimates to the exclusion of years following the introduction of the NMW: 2000 and later. This period saw a significant expansion of in-work means-tested benefits as well, so excluding these years will also limit any impact of such benefits in our estimates. The New Earnings Survey Panel Dataset ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The data used in this paper are taken from the New Earnings Survey Panel Dataset (NESPD), a mandatory survey of employers' payroll records, which collects data on employees' earnings and hours worked. The target sample frame of the NESPD is civilian employees in Great Britain whose National Insurance (NI) number ends with a specific pair of digits. Since the last digits of NI numbers are allocated randomly to all adults and the NESPD sample uses the same pair of digits each year, in principle this should deliver a random 1% panel sample of employees, from the late 1970s onwards. In practice, despite the survey supposedly being mandatory, non-response reduces the sample to around 0.7% of employees on average over the period. Nevertheless, at around 165,000 individuals per year the NESPD contains a much larger sample than is available in other UK datasets of hours and earnings (such as the Labor Force Survey and the Living Costs and Food survey), and it does not suffer from the same degree of measurement error as responses are provided by employers with reference to their payroll and employment records.14 We do not observe people when they are not employed, and cannot distinguish whether someone who is absent from a particular year of data was not working, was self-employed, or was working for an employer who failed to respond to the survey. As a result, the analysis in this paper focuses on those employed in consecutive waves of the NESPD, and estimates therefore do not account for any extensive margin responses to NICs. The NESPD records weekly hours and gross earnings – including overtime, commission, performance-related pay, etc., but excluding benefits in kind and employer pension contributions – for the pay period that includes a particular date in April each year (termed the survey reference date). This corresponds closely to the tax base for NICs, which is levied on a very similar definition of earnings, separately by pay period.15 However, the collection of these data in April is not ideal in the context of our study. This is because the UK's fiscal year runs from April 6th of one year to April 5th of the next, and changes to NICs often take effect at the start of the fiscal year. Furthermore, the NICs due on the earnings in the pay period captured by the NESPD depends upon the NICs system in place on the date those earnings were actually paid. Thus if payment for the pay period in question was made on or after April 6th the NICs system of the ‘new’ fiscal year would have applied. This will be the case for a large majority of our sample, including those paid on a calendar month basis at the end or middle of the month and those paid on a weekly basis unless the reference date is very near the start of April. For such employees, the earnings and hours captured by the NESPD will thus relate to the pay period immediately following any change to NICs taking effect at the start of the fiscal year. However, for those paid on April 5th or earlier, the NICs system of the ‘old’ fiscal year would apply. This includes those paid on a calendar month basis at the start of the month, and for the two years when the survey reference date was at the start of April, could include some of those paid on a weekly basis. For such employees, the earnings and hours captured by the NESPD will actually relate to the pay period immediately before any change to NICs taking effect at the start of the fiscal year. Unfortunately, the NESPD does not record either the date an employee is paid or the length of their pay period. In what follows we proceed by assuming the earnings we observe are subject to the NICs schedule of the fiscal year just beginning, as will be the case for the vast majority of our sample. This means we are typically estimating very short-run responses to changes in NICs rates; the effect of reforms implemented earlier in the same month. However, two of the biggest reforms to the NICs schedule (in 1985 and 1989) were implemented in October, not April, so a significant part of our identifying variation comes from reforms implemented around six months before the earnings and hours we observe. In addition, changes to the NICs schedule are invariably announced at least a few months in advance (so that payroll software can be ready in time to operate it, among other reasons), so our estimates will capture the effects of reforms announced some time beforehand. In those rare cases where the earnings we observe are in fact subject to the NICs schedule of the tax year just ending, our estimates will capture only earnings responses in anticipation of a reform's implementation, and the reform's implementation will instead be reflected in the subsequent year's earnings. Effects on labor cost ~~~~~~~~~~~~~~~~~~~~~ We now turn to our discussion of the identification of behavioral effects and incidence. We define labor cost (Z) as the sum of gross earnings (W) and employer NICs, and net-of- NICs earnings (C) as gross earnings minus employee NICs. Gross earnings are taken directly from the NESPD data, and both employee and employer NICs calculated using the NICs schedule. Such an interpretation relies on labor demand being perfectly elastic, and therefore the incidence of a tax – in this case NICs – being purely on the employee. If one allows for less than perfectly elastic labor demand – and therefore the prospect of at least partial incidence of NICs on employers – the β coefficients will pick up a combination of labor supply responses, labor demand responses, and shifting of incidence. With information on labor cost alone, it is difficult to disentangle these effects. Consider first the coefficients on average net-of-NICs rates (βZ,ρR and βZ,ρE). As well as capturing income effects, they will also reflect the incidence of NICs. For example, if labor cost increases in response to a rise in the average rate of employee NICs, that could reflect employees' working harder to make up the loss of income (an income effect), some shifting of the additional tax burden onto employers (incidence), or some combination of the two effects. In order to identify either income effects or incidence, one needs to assume the other: for example, assuming zero income effects allows one to interpret the coefficients in terms of incidence, while assuming full incidence on employees allows one to interpret the coefficients in terms of income effects. Changes in the extent to which NICs are shifted also mean that positive coefficients on marginal net-of-NICs rates (βZ,τR and βZ,τE) cannot be interpreted purely as compensated labor supply and demand responses either. In wage bargaining models, reductions in marginal tax rates, holding average rates fixed, can lead to higher wages, and hence higher labor costs, even absent changes in effort or hours of work (Lockwood and Manning, 1993; Pissarides, 1998).20 This is because lower marginal rates increase the marginal benefit to employees of bargaining for higher wages and reduce the cost to employers of acceding to these demands. Again, interpretation of these effects as behavioral responses requires making an assumption about the effects of marginal rates on incidence (and vice versa). Effects on hours of work and hourly labor cost ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The availability of hours of work in our data allows us to decompose changes in labor cost into changes in hours of work (H) and changes in labor cost per hour (Z/H). This can shed further light on the roles of incidence and behavioral responses in explaining changes in overall labor cost. Meanwhile, if leisure is a normal good, income effects would mean an employee would want to reduce their hours of work when their average net-of-NICs rate increases (holding the marginal net-of-NICs rate fixed). Potential responses to average rates by employers are more complicated. A higher average net-of-NICs rate, holding the marginal net-of-NICs rate constant, can be viewed as providing a lump sum transfer per employee. As well as inducing employers to hire more people and produce more output (both responses which our analysis cannot pick up), this might also prompt employers to change the hours of existing employees, for example by producing a given level of output with more workers but fewer hours per worker. Changes in labor cost per hour could also reflect behavioral response. For instance, when their net-of-NICs rate changes, employees may adjust their effort per hour worked, potentially resulting in a change in gross earnings (and hence labor cost) per hour. As with hours of work, there would be positive substitution effects for changes in marginal net-of-NICs rates and negative income effects for changes in average net-of-NICs rates. Employers may also adjust the way in which they compensate employees, such as making more or less use of benefits-in-kind or pension contributions: thus changes in the monetary labor cost per hour that we measure may not correspond to changes in overall labor cost per hour. However, we cannot empirically distinguish between such effects and changes in labor cost per hour that reflect the economic incidence of NICs. For instance, if NICs are partly incident on employers, an increase in the average net-of-NICs rate would lead to a reduction in labor cost per hour. This is observationally equivalent to an income effect on employees' effort of the kind mentioned above. Thus we cannot draw clear conclusions about the incidence of NICs from reduced-form estimates of their effect on hourly labor cost unless we rule out – or make precise assumptions about the size of – non-hours behavioral responses of the kind that originally motivated the ETI literature: changes in effort per hour worked and shifts in compensation to forms not liable for NICs. But if, for instance, one is willing to rule out such responses, changes in hourly labor costs will reflect changes in underlying wage rates, therefore capturing incidence effects that remain ambiguous when studying labor cost alone. In particular, the coefficients on average net-of-NICs rates in an hourly labor cost equation (βZ/H,ρR and βZ/H,ρE) would capture the average incidence of NICs. Coefficients of 0 (labor cost unaffected by the net-of-NICs rate) would indicate that NICs were fully incident on employees, while coefficients of −1 (labor cost moving one-for-one with changes in NICs) would indicate incidence fully on employers. However, as already highlighted, in a labor market characterised by wage bargaining compensated changes in marginal net-of-NICs rates can have effects on underlying hourly wage rates and hence hourly labor costs. In particular, an increase in the marginal net-of-NICs rates would increase the marginal benefit to employees of bargaining for higher wages and reduce the cost to employers of acceding to those demands, raising the hourly labor cost that results. Such an effect on hourly labor cost would be observationally indistinguishable from a substitution effect operating via the effort or compensation-form margin. One can therefore draw inferences about the effect of marginal net-of-NICs rates on incidence (and from this infer something about the wage determination process) only if one is willing to rule out non-hours behavioral responses, and vice versa. Differential effects of employee and employer NICs ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ If we can reject (any of) these conditions then we can conclude that this ‘invariance of incidence proposition’ does not hold within the timeframe we measure changes in hours and labor costs. If we believed we were picking up the long-run impact of NICs then we would have to move beyond these simple benchmark models of the labor market towards, for instance, a model where collective wage bargains are struck over groups of heterogeneous workers, or where bargaining is based on gross wages (as opposed to labor costs and net wages).21 It is not clear that sticky wages implies that βZ,τR = 0: employers could respond to compensated changes in the marginal NICs they pay on their employees' wages by changing the hours they offer individual employees. These hours changes would affect labor cost per employee. This test would be valid only if hours of work were also sticky. In that case, though, we would not expect hours to respond to changes in employee NICs rates either. Basic specification ~~~~~~~~~~~~~~~~~~~ Xi,t is a vector of controls, including year dummies (to pick up, for instance, the effect of inflation), a cubic of the employee's age, their sex, and whether they are in the same job in year t as year t − 1, as well as controls for differential trends in labor costs in different parts of the labor costs or hours distributions (discussed below). εi,t,z is an error term that captures unobserved and time-varying heterogeneity. It is well known from the labor supply and ETI literatures that various econometric challenges arise with estimation of such equations (Saez et al., 2012a). The first is a potential simultaneity bias. Because of the nonlinearity of the employee and employer NICs schedules, the marginal and average net-of-NICs rates τi,tE and τi,tR are themselves functions of the left-hand-side variable. Instruments are therefore required to identify the effect of NICs. The long-established standard approach to this problem, proposed by Auten and Caroll (1999), uses changes in the log net-of-tax rates holding earnings at their year t − 1 level as an instrument for actual changes in the net-of-tax rates. This does not address two further econometric issues: that mean-reversion in income processes and differential secular trends in income can confound estimates of the β coefficients, especially when tax rate changes affect different parts of the income distribution differently. Again following Auten and Caroll (1999), the literature has traditionally tried to deal with these issues by including functions of income/earnings in year t − 1 and/or year t − 2 in the regressions. However, Weber (2014) shows that instruments based on t − 1 income/earnings cannot be exogenous if there is mean reversion even if such controls are used. Instead, she proposes instrumenting changes in the net-of-tax rates using changes in these tax rates calculated holding earnings fixed at its level in an earlier year t − k (rather than t − 1) where k ≥ 2.22 The lag k should be chosen so that it is far enough before the year in question that earnings in year t − k are unaffected by any transitory shocks affecting earnings in years t − 1 and t, but not so far that the instrument is a weak predictor of the actual change in tax rate observed. It is possible to test whether a particular instrument is exogenous, conditional upon the assumption that other excluded instruments (for example, based on longer lags) are exogenous using a Difference-in-Sargan test. As we show in Table E3 of Appendix E, we, like Weber, reject the exogeneity of instruments based on earnings in year t − 1 (the p-values of the tests are 0.000 for each of labor cost, hours and hourly labor cost). But Table E2 shows we cannot reject the exogeneity of instruments based on earnings in year t − 2 (p-values of 0.16, 0.74 and 0.24, respectively).23 We therefore use instruments based on earnings in year t − 2 in our main specification.24 While our use of Weber-type instruments should deal with the problem of mean reversion, it will not deal with the problem of longer-term differential earnings trends. For instance, during the 1980s, earnings inequality was increasing significantly in the UK (Blundell and Etheridge, 2010). At the same time, changes in NICs (most notably in 1985) reduced average NICs rates at the bottom of the earnings distribution, and increased them at the top. The risk is that one would inappropriately attribute secular changes in relative labor costs to the tax reform, biasing estimated coefficients. To control for differential trends in earnings and hours at different parts of the distribution, we include controls based on lnZi,t−2. Following Eq. (3.2)., the changes in average net-of-NICs rates included in our regressions (∆ ln ρi,tR∣ and ∆ ln ρi,tE∣) are calculated holding earnings fixed at year t − 1. Lehmann et al. (2013) show that using the actual change in average net-of-tax rates or virtual income between years t − 1 and t would lead to inconsistent estimates, even if instrumented. As with the changes in marginal net-of-NICs rates, we instrument these using changes in average net-of-NICs rates calculated holding earnings fixed at year t − 2 (and in robustness checks, t − 3) levels. Including lagged changes in NICs ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Examining changes in labor cost, hours and labor cost per hour between years t − 1 and t, and relating this to changes in tax rates between year t − 1 and t may pick up only very short run behavioral effects and incidence shifting, especially given the timing of our data collection (discussed in Section 2.3). The traditional approach to capturing longer- run responses is to use panel lengths of longer than one year: for instance, calculating ∆ ln Zi,t = ln Zi,t − ln Zi,t−2 or ∆ ln Zi,t = ln Zi,t − ln Zi,t−3 rather than calculating ∆ ln Zi,t = ln Zi,t − ln Zi,t−1. This is the approach taken by Gruber and Saez (2002), among others. However, such an approach will not accurately measure the long- run effect, but rather (at best) an average of responses over the extended period. And if there are multiple reforms, estimation may be confounded, with estimated elasticities not representing even an average of shorter- and longer-run responses.25 βZ,τR,0, for instance, picks up the immediate effect of reforms: that is, the effect of changes in the marginal net-of-employer-NICs rate between years t − 1 and t on the change in labor cost, Z, between years t − 1 and t. βZ,τR,1 picks up the (additional) effect in the subsequent year: that is, the effect of changes in the marginal net-of-employer-NICs rate between years t − 2 and t − 1 on the change in labor cost, Z, between years t − 1 and t. The overall effects of a change in NICs taking effect between t − 2 and t − 1 on labor cost in year t can be calculated by adding the coefficients on the contemporaneous and lagged changes in tax rates (e.g. βZ,τR,0 + βZ,τR,1). In Table E1 of Appendix E we also show estimates for regressions that include only lagged changes in NICs rates. As with contemporaneous changes in tax rates, it is important to instrument Δlnτi,t−1 and Δlnρi,t−1∣ appropriately. We do this using instruments based on earnings held fixed at year t − 3 levels as in our basic specification examining short-run responses. We again test the robustness of results to using instruments based on earnings held fixed at year t − 4 levels.","Our estimation sample consists of approximately 1.7 million observations.27 Men make up just under 6-in-10, those aged 50 or over make up 3-in-10 and those working in the public sector in both years t − 1 and t make up 1-in-3 of the estimation sample.28 This sample is somewhat more male, older and more likely to be in the public sector than the full NESPD sample. This reflects the fact that individuals with these characteristics are more likely to be observed for the requisite number of years. Men, for instance, are less likely to move in and out of work, and to have the very low levels of earnings that mean they may not be sampled even if working. Identifying variation in NICs rates ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As discussed in Section 2.1, there were particularly significant reforms to NICs in October 1985, October 1989 and April 1999. Fig. 2a to f show the (log) changes in the average and marginal employee and employer net-of-NICs rates between 1985 and 1986 (Fig. 2a and b), between 1989 and 1990 (Fig. 2c and d) and between 1998 and 1999 (Fig. 2e and f). Overlaid on these NICs-change schedules are histograms showing the density of the earnings distribution. Fig. 2a, for instance, shows that that changes in employee NICs rates in 1985 were largest towards the bottom of the earnings distribution, although sufficiently far up for there to be significant sample sizes. It also shows that marginal and average employee NICs rates changed in almost exactly the same way, with the exception of an increase in marginal NICs rates for a narrow range of earnings around £265 per week, associated with the UEL increasing slightly faster than average earnings growth. Fig. 2b shows a similar picture for employer NICs for the bottom part of the earnings distribution – although changes in these extended further up the earnings distribution. The extension of full employer NICs above the UEL led to a large discrete fall in the marginal net-of- NICs rate, and to rises in the average net-of-NICs rate that smoothly increased above this point. The 1985 changes were substantial: increasing mean net-of-NICs rates by up to about 5% in some cases and reducing them by up to about 10% in others. Fig. 2c and d likewise show the effect of the 1989 reform, which provides significant differential variation in average and marginal rates of employee NICs, while Fig. 2e and f illustrate the variation in average and marginal rates of employer NICs from the 1999 reform. Each of these reforms can thus make a significant contribution to disentangling the effects of different elements of the NICs regime. The figures also show clearly that the 1985, 1989 and 1999 reforms to NICs generated non-convex changes in the budget sets facing employees and employers. In such circumstances, changes in NICs at levels of earnings higher than existing earnings (and so which affect neither marginal nor average NICs rates at existing earnings) can lead to discrete jumps in earnings in response. Such effects are not accounted for in the log-linear specifications we estimate in this paper. Our estimates therefore come with the caveat that they may not pick up all intensive margin responses (recall that our data do not allow us to pick up extensive margin responses at all). Since the 1985 reforms affected employee and employer average and marginal NICs rates very similarly (except towards the very top of the earnings distribution), and since the 1989 and 1998 reforms each significantly affected only one of employee or employer NICs, it is not possible to use a single reform to identify the separate effects of changes in both average and marginal rates of both employee and employer NICs. In addition, as shown in Appendix C, mean reversion in earnings is a significant issue when estimating the impact of reforms that predominantly affect those with low earnings, making reliance on a single reform problematic. The use of instruments based on earnings in year t − 2 (or t − 3), and the inclusion of functions of year t − 2 earnings as control variables, is designed to minimise the impact of such mean reversion. However, as discussed in Kleven and Schultz (2014), having reforms that both increase and decrease rates of a tax on a given part of the income/earnings distribution is also beneficial (as the effects of mean reversion work in opposite directions for tax increases and decreases). We therefore turn to examining variation in tax rates using the full panel, covering reforms between 1982 and 2015. Table 1 shows the mean and standard deviation of (100 multiplied by) the changes in log net-of- NICs rates, by position in the earnings distribution. It shows that marginal net-of- employee-NICs rates fell over the period as a whole, especially in the lower-middle and middle part of the earnings distribution. This reflects increases in the statutory rates of NICs. Larger reductions in the marginal net-of-employer-NICs rates at the top of the earnings distribution reflect, in large part, the uncapping of employer NICs in October 1985. The abolition of the notches at the LEL means that average net-of-NICs rates increased towards the bottom of the earnings distribution for employees, and a long way up the earnings distribution for employers (though the uncapping of employer NICs means average net-of-employer-NICs rates still fell at the very top of the distribution). Changes in marginal and average NICs rates over the last 30 or so years therefore differ significantly, in principle allowing identification of the effects of both. But our identification does not come solely from variation across the earnings distribution. The standard deviations of changes in log net-of-NICs rates for given parts of the earnings distributions are much larger than the mean changes, reflecting a pattern of increases and decreases in NICs rates at specific parts of the distribution in different years. Fig. 3 shows significant variation in changes in log net-of-NICs rates both across and within years.29 Panels (a) and (b) show that the largest changes to mean NICs rates were in years when the main NICs rate changed: for instance, 1994 for employees only and 2003 and 2011 for both employees and employers. However it is the variation in NICs rate changes within years that matters most for identification (given that we include year dummies in our regressions). Panels (c) and (d) show that in many years there was relatively little variation but the major reforms of October 1985, October 1989 and April 1999 are associated with clear ‘spikes’ in the standard deviation of NICs rate changes (in 1986, 1990 and 1999). There was also significant variation in 1997, 2000 and 2011, associated with further policy changes.30 Estimation results ~~~~~~~~~~~~~~~~~~ Before separately estimating the effects of employee and employer' NICs, we first estimate the responsiveness of labor cost to the overall rate of NICs. The results of these regressions are shown in Table 2. Column 1 shows estimates where our set of controls Xi,t includes a cubic of lnZi,t−2; column 2 shows estimates based on a quintic of lnZi,t−2, and; column 3 shows estimates based on a 10-piece spline of lnZi,t−2. Estimated coefficients are broadly stable across specifications and later results are shown only for quintic controls. βZ,τNI,0 measures the effect of changes in the overall marginal net-of- NICs rate on labor cost: estimates are positive, albeit small and not quite statistically significant. This would imply that labor cost is relatively unresponsive to compensated changes in marginal NICs rates, suggesting a relatively small deadweight loss of the tax. On the other hand, estimates of βZ,ρNI,0 (which measures the effect of changes in the average net-of-NICs rate on labor cost) are negative, large and highly statistically significantly different from 0, but not from −1. These estimates can be interpreted in several ways. If we assume that the incidence of NICs is fully on employees, then this would imply very large income effects. Alternatively, if we assume that income effects are small, then the estimates would imply that the incidence of NICs changes was largely on employers. Finally, the results are consistent with moderate-to-large income effects and sharing of the burden of NICs between employees and employers. As noted above, though, the effects of employee and employer NICs may differ, especially in the short-to-medium term (e.g. due to wage stickiness). Column (1) of Table 3 shows coefficients estimated separately for employee and employer NICs. Our estimate of βZ,τE,0 is statistically significant and positive (0.25), while βZ,τR,0 is near zero. These results are very similar to those of Lehmann et al. (2013) for France, where they found positive compensated elasticities for marginal net-of-income tax rates and zero elasticities for employers' SSCs. βZ,ρE,0 is negative, of a moderate size, but not statistically significantly different from 0 (it is statistically significantly different from −1, however). βZ,ρR,0 is also negative, but much larger: indeed, statistically significantly less than −1. We can reject both that βZ,τE,0 = βZ,τR,0 and that βZ,ρE,0 = βZ,ρR,0: the statutory incidence of NICs matters for its (immediate) impact on labor costs. Recall that in the absence of income effects, βZ,ρR,0 = −1 would indicate 100% incidence of employer contributions on the employer. One possible explanation for a coefficient statistically significantly less than −1 is that changes in average employer NICs rates are ‘over- shifted’ to employers, as can happen in some models. It might also be natural to view βZ,ρR,0 < −1 as a combination of statutory incidence with standard income effects. However, if employers are bearing the burden of employer contributions, there is no change in earnings which can generate an income effect for the employee in question. This is an issue to which we return below. In columns (2) and (3) we decompose the changes in labor cost into changes in hours of work and changes in hourly labor cost, to shed further light on the role behavioral response and incidence play in our findings. For hours of work, we find a positive compensated elasticity for employee NICs (βH,τE,0) but a zero compensated elasticity with respect to employer NICs (βH,τR,0). In contrast, for hourly labor cost, elasticities with respect to employee NICs (βZ/H,τE,0) and employer NICs (βZ/H,τE,0) are virtually identical. This provides strong evidence that the differences in responses of overall labor cost reflect differential behavior responses to changes in marginal employee and marginal employer NICs, rather than differential effects of NICs progressivity on hourly wage rates (and thus hourly labor costs). Our results contrast with those of Blomquist and Selin (2010), who find that increases in earnings in response to a reduction in marginal income tax rates are largely driven by increases in hourly earnings rather than hours of work, at least for men.31 Our results also differ somewhat from Lehmann et al. (2013), who find that labor cost responses are driven almost entirely by those working less than a full year – suggestive that weeks worked per year respond more than hours per week. However, given that NICs are levied per pay period (rather than on an annual basis), in our context it makes sense that any hours response to marginal NICs rates takes place at the ‘intensive’ rather than the ‘extensive’ margin: changes in the number of pay periods worked will not affect the amount of NICs paid in any given pay period. It is also worth noting that because of the timing of our data, these estimates may reflect short- term inter-temporal responses to changes in NICs: shifting hours of work from a pay period with a higher marginal employee rate to one with a lower marginal employee rate. This is an issue we return to when discussing estimates that include lagged changes in net-of-NICs rates. Turning to responses to average net-of-NICs rates, we find hours-of-work elasticities for employee (βH,ρE,0) and employer (βH,ρR,0) NICs that are similar to each other and both statistically significantly below 0. For hourly labor costs, the elasticity for the average net-of-employee-NICs rate is not statistically significantly different from 0, while that for the average net-of-employer-NICs rate is not statistically significantly different from −1. How should we interpret these results? βZ/H,ρE,0 = 0 and βZ/H,ρR,0 = − 1 is consistent with the economic incidence of NICs following statutory incidence: employees bear changes in the amount of employee NICs due, and their employer bears changes in the amount of employer NICs due. In this context, a negative elasticity of hours with respect to the average net-of-employee-NICs rate might naturally be interpreted as indicative of income effects: higher post-NICs pay inducing individuals to ‘buy’ more leisure by reducing their hours of work. However, income effects cannot rationalise the negative elasticity of hours with respect to the average net-of-employer- NICs rate: if employers are bearing the burden of employer contributions then there is no change in earnings which can generate an income effect for the employee in question. Why then should a reduction in the average employer NICs rate (controlling for the marginal rate) cause employers to reduce their employees' hours of work? One reason, highlighted in Section 3.2, is that it can be thought of as providing firms with an extra lump sum per worker. Given this, one response for firms is to adjust their production technology to use more workers, each working fewer hours. As we cannot capture extensive margin responses, our regressions will simply pick up the reduced hours per worker. However, while this mechanism is possible, it seems unlikely to be sufficiently large to explain the size of the estimated coefficients. There remains the possibility that the result is driven by some omitted factor (perhaps a non-NICs reform or a macroeconomic shock) with a large effect on hours of work that is correlated with the variation in NICs rates that identifies the effect of average employer NICs changes. For now, this result remains a puzzle within our broader findings. Adding lagged NICs changes Panel (B) of Table 3 adds lagged changes in net-of-NICs rates as regressors. This allows us to examine whether the effects of changes in NICs differ after an additional year, with overall effects obtained by adding the coefficients for the contemporaneous and lagged changes together (e.g. βX,τE,0 + βX,τE,1)32. Two findings stand out. First, effects for contemporaneous changes are broadly similar to those in Panel (A): adding lagged changes in NICs rates does not affect the general pattern of results for contemporaneous changes – statutory incidence continues to matter for responses to both marginal and average net-of-NICs rates. Second, while the coefficients on the lagged changes in marginal and average net- of-NICs rates are sometimes statistically significant, so that effects of NICs changes differ after a year, there is no evidence of a move towards employee and employer NICs having a more similar effect after a year. For instance, the gap between the effect of average employee (βz/h,ρE,0 + βz/h,ρE,1) and average employer (βz/h,ρR,0 + βz/h,ρR,1) net-of-NICs rates on hourly labor costs grows, if anything, after a year, rather than shrinking (statistically speaking, we cannot reject the hypothesis that the gap remains unchanged). In other words, if employer NICs begin to be shifted to employees (or vice versa), this does not appear to occur within the year or so following a NICs change. Similarly, the gap between the effects of marginal employee (βh,τE,0 + βh,τE,1) and marginal employer (βh,τE,0 + βh,τE,1) NICs rates on hours of work also grows. The statistically significant positive coefficient on the lagged change in the marginal net-of- employee-NICs rate suggests that the full behavioral response to marginal employee NICs rates does not occur immediately. Thus (temporary) inter-temporal responses to changes in marginal employee NICs rates are unlikely to be driving our results for specifications including contemporaneous changes in net-of-NICs rates only. Robustness checks In Appendix E we investigate the robustness of the baseline estimates reported in Panel (B) of Table 3 to different instruments and labor cost controls, and to restrictions on the years of data included in the estimation sample. Estimates are robust to using instruments based on year t − 3 earnings, but not instruments based on year t − 1 earnings. However, the Difference-in-Sargan statistics for the latter instruments are all highly statistically significant, suggesting that they are unlikely to be exogenous (probably due to mean reversion). Estimates are also generally robust to excluding years where the survey reference date was before or close to the start of the fiscal year, and when we restrict ourselves to years with major reforms. We also find the same qualitative differences between the responses of labor costs and hours of work to employee and employer NICs when we restrict our sample to the period up to 1999, the point at which the National Minimum Wage was introduced. In particular, the effects of changes in average net- of-employer-NICs rates on labor cost per hour are very similar to those obtained for the full sample, suggesting that the introduction of the NMW did little to affect the overall incidence of employer NICs (at least across the earnings distribution as a whole). But there are some non-trivial differences in the point estimates of the coefficients on changes in marginal net-of-employee-NICs rates in the hours-of-work regressions for this sub-period. These differences suggest that we should be cautious about drawing overly strong conclusions about the precise scale and dynamics of hours-of-work responses to employee SSCs. Demographic sub-groups Estimates for selected demographic groups can be found in Appendix F. Again, the clearest findings from analysis of the full sample hold: hours responses to changes in marginal employee NICs rates are larger than marginal employer NICs rates; and hourly labor cost responses to average employee and employer NICs rates suggest that statutory incidence matters for at least a year or so following a reform. Furthermore, as one might expect, responses differ for workers in the public versus private sector. There is evidence, for instance, of responses to changes in marginal employee NICs rates occurring more quickly in the private sector than the public sector. There is also evidence, after one year, of some shifting of the burden of employee NICs to employers in the public sector (βZ/H,ρE,1 < 0) but not the private sector (βZ/H,ρE,1 = 0). Other results are more unusual, though: Men are more responsive (for instance, βH,τE,0 = 0.317) to changes in marginal employee NICs rates than women (βH,τE,0 = 0.033), which contrasts with the typical finding in the literature.33 However, this could partly be an artefact of one particular measurement problem we face: our inability to observe whether women in our sample are subject to the married women's reduced rate of employee NICs (see Section 2.1). This rate has not changed in the same way as the standard rates of employee NICs that we model. Measurement error in NICs rates may therefore attenuate estimates of responses to employee NICs for women. While incidence is initially close to statutory, statistically significant coefficients on lagged changes in average net-of-employer-NICs rates suggest that employers more than fully bear employer NICs for women and those aged over 50 after 12–18 months. This could reflect ‘overshifting’ of employer NICs to employers, which (as noted in Section 3.1) is possible in labor markets characterised by wage bargaining. This is consistent with the fact that older workers are more likely to be members of trade unions in the UK, although historically women were less likely to be members than men (this reversed around the year 2000).34 In contrast, employees seem to more than fully bear employee NICs among the over-50s. This is more difficult to rationalise. Summary Overall, while estimates differ quantitatively to some extent across specifications and across sub-samples, we find consistent evidence that the statutory incidence of NICs matters for their economic effects. In particular: The compensated elasticity of hours with respect to the marginal net-of-employee-NICs rate is generally positive (βX,τE > 0), while that for employer NICs is generally not statistically significantly different from 0. The elasticity of hourly labor cost with respect to the average net-of-employee-NICs rate is generally close to 0, while that for the average net-of-employer-NICs is generally close to −1. Our preferred interpretation is that employees respond to NICs by adjusting their labor supply on the intensive margin, but that wage stickiness (lasting at least a year) means that the statutory incidence of NICs significantly affects their impact on behavior and their economic incidence in the short-to-medium term. Our decomposition of labor cost into hours and hourly labor costs allows us to rule out the possibility that responses to marginal net-of-employee-NICs rates reflect increases in wage rates – which is a prediction of some bargaining models of the labor market – rather than changes in labor utilisation.","In this paper we have estimated the responses of labor cost, hours of work and hourly labor cost to marginal and average rates of employee and employer SSCs in the UK, using reforms during the 1980s, 1990s, 2000s and early 2010s as our source of identifying variation. The existing literature analysing the responsiveness of earnings to SSCs using micro-data is sparse relative to the voluminous literature on the responses of taxable income to income tax; this paper helps to fill that gap. Furthermore, by considering estimated effects in the context of both behavioral responses to SSCs and the (intimately related) issue of their economic incidence, this paper attempts to help link the ETI and incidence literatures. Our main specifications include both contemporaneous and lagged changes in net-of-NICs rates, allowing us to look at the impact of NICs over the first year or so following a reform, rather than just the immediate impacts over the first 0 to 6 months. And multiple policy reforms provide us with the power to separately identify the effects of employee and employer NICs. Our estimates suggest that responses to employee and employer NICs differ significantly. We find that reductions in marginal rates of employee NICs have positive and statistically significant effects on labor cost, operating via hours of work. We also find that labor cost falls more than one-for-one when average employer NICs rates are reduced, but by much less when employee NICs rates are reduced, with most of the effect (and nearly all of the difference) operating via hourly labor cost. These findings are robust across specifications based on different instruments, different sub-periods of analysis, and different sub-groups of employees. We interpret this as evidence that employees change their hours in response to NICs, but that in the short- to medium-run at least, the formal incidence of NICs matters for their behavioral impacts and economic incidence. If employers respond to changes in marginal employer NICs rates at all, they do so by adjusting the number of employees they hire and/or how they compensate their employees (things we cannot observe in our data), rather than by adjusting the hours offered to individual employees. Lehmann et al. (2013) draw similar conclusions from their analysis of the effects of changes in income tax and employer SSCs on labor costs in France. But our decomposition of labor cost into hours and hourly labor costs allows us rule out the possibility that the positive compensated elasticity of labor cost with respect to marginal employee NICs rate arises via the wage bargaining process rather than changes in hours of work. Our estimates of the effects of average employee NICs rates on hours of work suggest that income effects are significant for the wage earners for whom our elasticities are estimated. We also find that reductions in average employer NICs rates lead to statistically significant reductions in hours of work. This cannot be rationalised as an income effect if, as our hourly labor cost regressions imply, the short-run incidence of employer NICs is on firms. One potential explanation is that a reduction in average employer NICS rates, holding marginal rates constant, is akin to providing firms with an extra lump sum per worker: a plausible response by firms could therefore be to utilise more workers, each working fewer hours. Future work could extend this analysis in several ways. First, one could consider the effects of income tax and of means-tested benefits and tax credits. In addition to allowing an assessment of the impacts of these additional policies, it would also allow for more robust estimation of the impacts of NICs (by reducing the risk of omitted variable bias). The NESPD, on which this study is based, does not allow us to estimate the marginal and average tax rates associated with these parts of the tax and benefit system: they are assessed on broader, usually family-level, measures of income, rather than individual earnings. To do this in the UK, one would therefore need to make use of alternative datasets such as the Family Expenditure Survey which can provide such information, but which also has a smaller sample size and greater measurement error in earnings than the NESPD. Another possibility would be to use data from a country with detailed administrative panel data on both individual earnings and family incomes and composition. Second, it would be worthwhile to examine the longer-run incidence of SSCs by examining the impact of reforms beyond the first year. Estimating long-run incidence is more difficult, though, as the potential for other differential changes in policy/institutions or the economic environment to impact labor costs is greater, making it more difficult to isolate the effects of SSCs. A third interesting extension would be to examine the ultimate incidence of SSCs that are not incident on the individual employee on whose earnings they are levied. For instance, using matched employee-employer data, if there is firm-level variation in exposure to different reforms to SSC rates, one could examine whether incidence is shifted on to other employees in a firm that are not directly affected by particular reforms, or is shifted elsewhere, such as to shareholders, customers, or suppliers. Where the ultimate economic burden of SSCs lies can have important implications for not only their distributional effects, but also for the measurement of the fiscal externalities generated by them."],["We build a model of administrative barriers to trade to understand how they affect trade volumes, shipping decisions and welfare. Because administrative costs are incurred with every shipment, exporters have to decide how to break up total trade into individual shipments. Consumers value frequent shipments, because they enable them to consume close to their preferred dates. Hence per-shipment costs create a welfare loss.We derive a gravity equation in our model and show that administrative costs can be expressed as bilateral ad-valorem trade costs. We estimate the ad-valorem equivalent in Spanish shipment-level export data and find it to be large. A 50% reduction in per-shipment costs is equivalent to a 9 percentage point reduction in tariffs. Our model and estimates help explain why policy makers emphasize trade facilitation and why trade within customs unions is larger than trade within free trade areas. --------------------------------------------------------------------------------","Exporters and importers around the globe face many administrative barriers. They have to comply with complex regulations, deal with a large amount paperwork, subject their cargo to frequent inspections, and wait for lengthy customs clearance. Minimizing the burden of these procedures, “trade facilitation” has been a priority for policymakers from developed and developing countries alike. In December of 2013, all members of the World Trade Organization (WTO) have agreed to the Bali Package, the first comprehensive agreement of the Doha round of negotiations. The main component of the Bali Package is an agreement on trade facilitation, requiring WTO members to adopt a host of measures streamlining the customs process, such as pre-arrival processing of shipments, electronic documentation and payment, and the release of goods prior to the final determination of customs duties, “[w]ith a view to minimizing the incidence and complexity of import, export, and transit formalities and to decreasing and simplifying import, export, and transit documentation requirements […]”1 Why do countries rush to facilitate trade? They hope to increase trade volumes without endangering government revenues by reducing inefficiencies. In fact, studies of various trade facilitation measures find that they are associated with larger trade volumes.2 Even among countries within free trade areas (FTAs), tighter economic integration and a reduction of administrative barriers often lead to higher trade.3 We build a model of administrative barriers to trade to understand how they affect trade volumes, shipping decisions and welfare. A large part of administrative trade barriers are costs that accrue per each shipment, such as filling in customs declaration and other forms, or having the cargo inspected by health and sanitary officials. Hornok and Koren (in press) document that countries with high administrative barriers to importing (as reported in the Doing Business survey of the World Bank) receive less frequent and larger shipments. The starting point of our model is hence a tradeoff between administrative costs and shipping frequency. In the presence of per-shipment administrative costs, exporters would want to send fewer and larger shipments. However, an exporter waiting to fill a container before sending it off or choosing a slower transport mode to accommodate a larger shipment sacrifices timely delivery of goods and risks losing orders to other, more flexible (e.g., local) suppliers. With infrequent shipments a supplier of such products can compete only for a fraction of consumers in a foreign market. We first use our model to derive a modified gravity equation of trade flows, in which administrative barriers show up as an additional tax on imports. Intuitively, when administrative barriers are high and shipments are infrequent, customers suffer utility losses that can be quantified as an ad-valorem tax equivalent. They also substitute towards local products accordingly. We then show how to measure the welfare losses from administrative barriers to trade by estimating two key elasticities from the data. First, we need to know the sensitivity of consumers to timely shipments. This can be recovered from the observed shipping choices of exporters: if customers are very sensitive to timeliness, firms send many small shipments. Second, we need to estimate shippers' reaction to administrative costs. Calculating deadweight losses from the elasticity of consumer and firm behavior to prices is in the spirit of semi-structural estimation of Harberger (1964), Chetty (2009) and Arkolakis et al. (2012). In our empirical analysis, we first show that per-shipment trade costs are sizeable and important for trade flows. We use the Doing Business database to measure the cost of shipping. Across 161 countries, the average trade shipment was subject to $3000 shipping cost in 2009. High shipping costs are associated with low volumes of trade: country pairs at the 25th percentile of per-shipment costs trade 68% more than country pairs at the 75th percentile. This magnitude is comparable to the trade creating effect of sharing a common border. Administrative barriers of trade are larger in poor countries than in rich ones. Doubling the income of an importing country is associated with a 6% decrease in per-shipment costs. This pattern is consistent with the fact reported by Waugh (2010) that poor countries have higher trade barriers than rich ones, without correspondingly higher consumer prices of tradables. In our model, administrative barriers affect the convenience of imports and hence trade volumes, but not consumer prices. In addition, we find that administrative costs help explain trade flows among countries with no preferential trade agreements and even within FTAs, but not within customs unions. One potential reason for this is that customs unions are subject to much less administrative barriers than FTAs and our measured administrative costs do not apply. In fact, this could provide a new explanation for why trade within customs unions is higher than trade within FTAs.4 Traditionally, the analysis of FTAs relative to customs unions focused on tariff harmonization, rules of origin and political economy (Krueger, 1997 and Frankel et al., 1997, 1998). We then study how exporters break down trade into shipments by exploiting shipment-level data for Spain for the period 2006–2012. We find that countries facing higher administrative barriers receive fewer shipments. This is similar to the finding of Hornok and Koren (in press). Using our estimated elasticity of the number of shipments with respect to per-shipment costs, we can calibrate the welfare effect of these costs. We conduct two counterfactual trade facilitation experiments in the model. In the first one we reduce per-shipment costs by half. In the model, this is equivalent to about a 9% reduction in tariffs and results in about a 31% increase in trade volumes. The second exercise harmonizes administrative barriers so that each country matches the per-shipment cost of the average country in the top decile of GDP per capita. Because richer countries have lower trade barriers, this typically involves a reduction in per shipment costs. The tariff-equivalent effects of this policy vary substantially with development, being 13% for the lowest income decile and 2% for the highest. Our counterfactual exercises suggest large trade creating effects of trade facilitation and large distributional effects from harmonizing administrative barriers. Our emphasis on shipments as a fundamental unit of trade follows Armenter and Koren (2014), who discuss the implications of the relatively low number of shipments on empirical models of the extensive margin of trade. We relate to the recent literature that challenges the dominance of iceberg trade costs in trade theory, such as Hummels and Skiba (2004) and Irarrazabal et al. (2013) They argue that a considerable part of trade costs are per unit costs, which has important implications for trade theory. Per unit trade costs do not necessarily leave the within-market relative prices and relative demand unaltered, hence, welfare costs of per unit trade frictions can be larger than those of iceberg costs.5 The importance of per-shipment trade costs or, in other words, fixed transaction costs has recently been emphasized by Alessandria et al. (2010). They also argue that per-shipment costs lead to the lumpiness of trade transactions: firms economize on these costs by shipping products infrequently and in large shipments and maintaining large inventory holdings. Per-shipment costs cause frictions of a substantial magnitude (20% tariff equivalent) mostly due to inventory carrying expenses. We consider our paper complementary to Alessandria et al. (2010) in that we exploit the cross-country variation in administrative barriers to show that shippers indeed respond by increasing the lumpiness of trade. Relative to their work, our focus is on characterizing the welfare consequences of administrative barriers in a simple-to-calibrate framework. We can do this by leveraging the semi-structural approach. Our work is most related to Kropf and Sauré (2014), who build a heterogeneous-firm trade model to study how fixed costs per shipment affect shipment size. They characterize the size and frequency of shipments as a function of firm productivity, and also show that aggregate exports follow from firm-level trade patterns. They then recover shipment costs from the observed shipment sizes, showing that these imputed costs are large and correlate plausibly with geographic variables and trade agreements. Given our lack of firm-level data, our model is admittedly simpler, but our focus here is also different: we want to understand the welfare consequences of administrative trade costs in a tractable aggregate framework. For this purpose, we derive a standard gravity equation in our model, and show how our model can be calibrated using a limited set of aggregate moments. We also offer new evidence on the trade effects of administrative barriers within and outside customs unions and conduct various counterfactual experiments.","This section presents a model that determines the number and timing of shipments to be sent to a destination market. Sending shipments more frequently is beneficial, because consumers value timely shipments. Producers engage in monopolistic competition as consumers value the differentiated products they offer. Each producer can then send multiple shipments to better satisfy the demands of its consumers. There are J countries, each hosting an exogenous number of sellers and consumers. A seller can sell to a domestic consumer at no shipping cost. It can also sell to a foreign destination j, in which case it has to pay iceberg shipping costs as well as the administrative cost of exporting to country j. The difference between administrative barriers and other trade costs is that the former apply for every shipment. We hence model them as per-shipment costs that are pure waste. We characterize the shipping problem of sellers, and derive a gravity equation for trade flows between countries. We show that administrative costs act as an ad-valorem tax on bilateral trade. We then discuss the welfare implications of administrative costs, deriving a semi-structural formula for consumer surplus in the spirit of Harberger (1964), Chetty (2009) and Arkolakis et al. (2012). Consumers ~~~~~~~~~ Consumers are willing to consume at a date other than their preferred date, but they incur a cost doing so. In the spirit of the trade literature, we model the cost of substitution with an iceberg transaction cost.8 A consumer with preferred date t who consumes one unit of the good at date s only enjoys e− δ|t − s| effective units. The parameter δ > 0 captures the taste for timeliness.9 Consumers are more willing to purchase at dates that are closer to their preferred date and they suffer from early and late purchases symmetrically. Clearly, because of perfect substitution, the consumer will only purchase the shipment(s) with the closest shipping dates, as adjusted by price, e− δ|t − s|/ps. Exporters ~~~~~~~~~ There is a fixed Mi measure of firms producing in each country i. Because there are no entry costs, each firm exports to each destination country j.10 Exporters decide how many shipments to send at what times. Sending a shipment incurs a per-shipment cost of fij. They then decide how to price their product. Both decisions are done simultaneously by the firms. The marginal cost of production of supplier ω is constant at c(ω).11 It takes gross iceberg costs tij > 1 for goods to reach country j from country i. This involves the per- unit costs of shipping, such as freight charges and insurance (it does not include per- shipment costs.) The cost-insurance-freight value of a good in country j is hence c(ω)tij. We abstract from capacity constraints in shipping, that is, any amount can be shipped to the country at this marginal costs. Because we do not study free entry, we abstract from entry costs for both production and market access. Net revenue is markup times the quantity sold to all different types of consumers at different shipping dates. The per- shipment costs have to be incurred based on the number of shipping dates, which we denote by nj(ω).12 Equilibrium ~~~~~~~~~~~ An equilibrium of this economy is a product price pj(t, ω, s), the number of shipments per firm nj(ω), and quantity xj(t, ω, s) such that (i) consumer demand maximizes utility, (ii) prices maximize firm profits given other firms' prices, (iii) shipping frequency maximizes firm profits conditional on the shipping choices of other firms, and (iv) goods markets clear. To construct the equilibrium, we move backwards. We first solve the pricing decision of the firm at given shipping dates. We then show that shipments are going to be equally spaced throughout the year. Given the revenues the firm is collecting from n equally spaced and optimally priced shipments, we can solve for the optimal number of shipments. Pricing Price is a constant markup over the constant marginal cost. Firms may be heterogeneous in their marginal cost because of differences in productivity or factor prices in their source country. Importantly, a given firm charges the same price for each shipment date. Shipping dates Clearly, revenue (4) is concave in |s − t|, the deviation of shipping times from optimal. Because of that, the firm would like to keep shipments equally distant from all consumers. This implies that shipments will be equally spaced, s2 − s1 = s3 − s2 = … = 1/n. The date of the first shipment is indeterminate, and we assume that firms randomize across all possible dates uniformly. An equal-spaced shipping equilibrium is shown on Fig. 1. Revenue The revenue of a firm is the product of two components: one depending only on market size and relative price as in a Krugman model, the other solely a function of shipping frequency. The ad-valorem equivalent of infrequent shipping, τ(n), has the following properties. It is decreasing in n: the more shipments the firm sends the more consumers it can reach at a low utility cost. Because they appreciate the close shipping dates, they will perceive this firm as relatively cheap. At the extreme, if n → ∞, τ(n) converges to 1, and the firm sells r(ω). From the firm's point of view, the demand for timely shipping is fully captured by the function τ(n), which acts as an ad-valorem tax on the firm's product. Later we will show that this analogy also applies to welfare calculations. Number of shipments It increases in δ (less patient consumers), increases in σ (consumers willing to substitute to other firms), increases in Ej/Pj (bigger firms in equilibrium) and decreases in fij (costly shipments). The number of shipments only depends on the ratio of maximal firm size rj(ω) and per-shipment cost fij. Everything that makes the firm larger in a market (large market size, weak competition, low costs of production and shipping) increases the optimal frequency of shipments. Large firms lose more by not satisfying their customers' need for timeliness and they are willing to incur per-shipment costs more frequently. Intuitively, lower per- shipment costs also imply more frequent shipments. At the extreme, as fij tends to zero, the firm sends instantaneous shipments, nj(ω) → ∞ and τ converges to one. The left-hand side of this equation is the absolute value of the elasticity of τ with respect to n. The right-hand side is a constant markup times total shipping costs paid by the firm (nf), divided by total revenue of the firm. The last fraction can hence be thought of as the ad-valorem amount of shipping costs. The intuition for this result is that the more elastic τ is with respect to the number of shipments, the less willing is the firm to sacrifice revenue with infrequent shipments. It will hence send many small shipments, making the ad-valorem amount of shipping costs large. We can use this formula to recover the elasticity of τ from the data. Trade flows The analysis so far is conditional on firm-level unit costs. To derive aggregate trade flows, we need to take a stand on these costs. Because we do not have firm- level data, we take the simple view that firms within the same country are identical in their cost of production, ci(ω) ≡ ci. An alternative approach, pursued by Kropf and Sauré (2014) would be to assume heterogeneous firms with Pareto-distributed unit costs. Chaney (2008) and Arkolakis et al. (2012) discuss under what conditions such a heterogeneous-firm model leads to a similar gravity equation to the one we derive below. This is exactly the gravity equation in Anderson and van Wincoop (2003) and Eaton and Kortum (2002), except for the additional term τ(nij). Infrequent shipment hence acts as a bilateral trade cost between countries. We can use this insight to calculate the magnitude of trade losses from administrative barriers.","What is the welfare cost of administrative barriers? Here we calculate how welfare depends on the choice of shipping frequency. The utility of the representative consumer is a monotonic function of real income Ej/Pj. We hence need to calculate the income and the price index faced by the representative consumer. Our gravity equation satisfies Restriction 3 (CES import demand) of Arkolakis et al. (2012), but not Restriction 2 (constant profit shares) and Restriction 3 (identical elasticity of trade flows to wages and trade costs). This is because profit net of shipping costs is a nonlinear function of revenue and trade policy hence also changes the profit to cost ratio of the economy. We thus cannot use the result of Arkolakis et al. (2012) to characterize welfare across all equilibria. We can still use a loglinear approximation around the equilibrium to show how the additional inconvenience from an infinitesimal increase in shipping costs maps into welfare losses for the consumer. Because we only consider changes to fij when analyzing welfare, we can treat customer income Ej as fixed as long as j is a small country. In this case, changes in the profits of exporters in country i do not matter for consumer income in country j. We can focus on changes to the price index.13 The welfare effect of a change in per-shipment costs is a weighted average across source countries. The contribution of country i to this welfare effect is ψij. We use this result in the counterfactual exercise in Section 5 when we estimate ψij.","We study how administrative barriers affect trade flows. We first estimate a gravity equation for bilateral trade volumes, including cost of shipping as an additional bilateral trade cost. We then show how shipping costs affect the number of shipments going to a country. Data and measurement ~~~~~~~~~~~~~~~~~~~~ We identify administrative costs from the Doing Business survey of the World Bank, from 2006 to 2012 (World Bank, 2014a). Doing Business measures the costs of exporting and importing a standard containerized cargo, noting the various customs and administrative procedures, documents, and the time and money they take. Our measure of shipping costs is the total monetary cost per shipment incurred by the exporter and the importer. Although not all components of these costs are strictly administrative, these correspond to the per-shipment cost in our model.14 Table 1 reports the average shipping costs across countries, broken down by the type of procedure and the direction of trade (export or import). Taken together, the average trade transaction would be subject to a total of $3000 cost and a waiting time of 50 days. To study the covariates (though not necessarily determinants) of administrative barriers, we regress the log total per shipment costs (including both exporting and importing costs) on a host of country and country-pair observables. Table 2 reports the results. Administrative barriers tend to be lower for larger and richer countries that are closer to one another and are members of an FTA. By far the most variation in administrative barriers is due to the level of development of the exporter and the importer. Doubling the GDP per capita of an importer is associated with a 6% decline in per-shipment trade costs. Twice as rich exporters, in turn, have 4% lower shipping costs. Motivated by this observation, we will study the effects of the counterfactual policy of reducing administrative barriers to rich-country levels.15 Data on trade flows comes from the UN Comtrade database (United Nations, 2014). We use bilateral distance measures and geographical variables from CEPII (Mayer and Zignago 2011), and gross domestic product data from the World Bank (World Bank, 2014b). To estimate how shipping choices depend on administrative costs, we need information on shipments. We use the shipment-level export database of the Spanish Agencia Tributaria (Agencia Tributaria, 2014). This contains information on every single international shipment leaving Spain. It records the date of shipment, its product code, value and weight, destination and transit countries, and the specifics of shipping, such as the mode of transport, the flag of vessel and whether the cargo is containerized. Agencia Tributaria does not make firm identifiers available, so, even though each shipment is made by a single firm, we cannot conduct firm-level analysis. Table 3 reports some shipment- level statistics about Spanish export in 2009. It shows, for selected destinations, the shipment value of the median product, the number of times it is shipped in a month, and the number of months it is shipped in a year. Our first observation is that shipments are relatively large and infrequent. The average shipment size across all importers is $13,234 and the typical product only ships twice a year to the typical destination. This observation, noted before by Armenter and Koren (2014) and Hornok and Koren (in press), motivated us to model shipments as infrequently spread through time. We also find that countries with lower per-shipment costs receive smaller and more frequent shipments. In our model, a shipment can only contain a single type of product from a single firm. In practice, shipments may be consolidated. Multiproduct firms may send different products or freight forwarders may send cargo of different firms in the same shipment. To check how important this model restriction is, we explored the bundling of shipments. For this exercise, we define a shipment based on shipping characteristics alone (such as date, final and transit country, vessel, containerization), while ignoring information on the product or its value. The vast majority, close to 60%, of such shipments contains only a single product item. We hence view the single-product, single-firm approximation of our model as empirically relevant. In order to differentiate customs unions from free trade areas, we use the May 2013 version of the database created by Baier and Bergstrand (2007) to measure economic integration agreements. Because that data ends in 2005 and Doing Business starts in 2006, we use the year 2006 in our estimation of trade volumes. Trade volumes ~~~~~~~~~~~~~ We first estimate a gravity equation of bilateral imports. We are interested in the trade- creating effect of customs unions relative to free trade areas, as well as the effects of per-shipment costs. Imports from country i to country j depend on an FTA and a customs union dummy, per-shipment costs, as well as standard gravity control variables. Note that fij is the sum of export-specific costs in country i and import-specific costs in country j. The unilateral control variables include total nominal GDP, GDP per capita in PPP terms (World Bank, 2014b), an indicator for whether the country is an island, and an indicator for its continent (CEPII GeoDist16). Bilateral controls include distance, adjacency, former colonial status, indicators for common language, common currency and common legal origins (CEPII GeoDist and Gravity17) and average bilateral tariff rate (CEPII MacMap18). We also control for whether one or both country is in the European Union, because within- EU trade data is collected differently.19 Table 4 reports the results. All specifications are estimated by ordinary least squares (OLS). Columns 1 through 2 are estimated on the full sample of 178 exporter countries and 148 importer countries. The specifications vary by the degree of economic integration of the country pairs. We include an FTA and a customs union dummy, as well as our measure of shipping costs. The omitted category of economic integration includes country pairs with no trade agreement, or only preferential trade agreements short of an FTA. All standard gravity variables have the expected sign and magnitude. As column 1 shows, countries in FTAs trade much more with one another than countries outside. An FTA is associated with a more than two-fold increase in trade. We separate customs unions from FTAs. Since all customs unions are also FTAs, the estimated effect of a customs union is in addition to the effect of being in an FTA. That is, customs unions are associated with a 44% increase in trade relative to FTAs.20 This is consistent with the model and the fact that customs unions require much less administration than FTAs. Column 2 reports the elasticity of trade volumes with respect to per-shipment cost to be strongly negative at − 1.06. The interquartile range of per- shipment costs is $1900 to $3100. This implies that country pairs at the 25th percentile of shipping costs trade 68% more than country pairs at the 75th percentile.21 What is the relationship between administrative costs, FTAs and customs unions? To answer this question, we break the sample into three. Column 3 includes country pairs that are not members of an FTA. Column 4 includes country pairs that are in an FTA but not in a customs union. Column 5 includes members of customs unions. We are interested in how the effect of shipping costs varies across these samples. Since the Doing Business survey asks about a standard cargo, it does not allow for the special administrative provisions of FTAs and customs unions. We find that the negative effect of per-shipment cost is strong among countries outside of FTAs. The effect is similar among FTA members. The negative effect of shipping costs disappears for customs union members. Indeed, for these country pairs, much of the shipping costs do not apply, as there is no customs clearance and documentation needs are much reduced. Column 6 reports a regression in which we interact shipping cost with FTA and customs union indicators. This is different from the regressions on the three subsamples in that other variables are restricted to have the same coefficient. Again, we see large negative association of shipping costs with trade for non-FTA members, somewhat smaller effects for FTA members, and the effect disappears for customs unions. The differences between the three groups are highly significant. Shipments ~~~~~~~~~ We then turn to see how exporters break down total trade into shipments. We use shipment- level export data from Spain for the period 2006–2012. We identify the number of shipments nij as the total number of shipments going from Spain to country j in given year. One drawback of the Spanish data is that it contains no firm identifiers. We thus cannot calculate the number of shipments per firm, so we use the total number instead. Although admittedly a limitation, this measure is consistent with the model, where all firms are symmetric, and the total number of shipments Nij = Minij is just a constant multiple of the number of shipments per firm. Table 5 reports the estimates. Because reporting standards are different for intra-EU trade, we only include non-EU destinations. Column 1 reports a simple OLS estimate for the 131 non-EU destinations. Countries with higher per- shipment cost receive significantly fewer shipments, with an elasticity of − 1.34. The interquartile range of per-shipment costs for non-EU destinations is $2000 to $3000. A country with $2000 per-shipment costs receives 72% more shipments from Spain than a country with $3000 costs. Column 2 reports a specification with destination fixed effects. Such fixed effects can soak up any time-invariant heterogeneity across countries and their relation to Spain. (This is why the gravity variables are omitted.) The coefficient of per-shipment costs is still negative but no longer significant, with a p-value of 0.235.22 The fixed effect estimate is very noisy because there is little time-series variation in administrative costs. An analysis of variance reveals that 91% of the variation in log shipping cost is soaked by up destination country dummies. An additional 5% can be attributed to common time dummies, leaving about 4% idiosyncratic time variation. In column 3 we report a random effect specification, which allows for time-invariant heterogeneity across countries, but restricts these error terms to be orthogonal to explanatory variables and to have a normal distribution. Given these restrictions, the random effect estimator uses both cross-section and time-series variation. The estimated elasticity of the number of shipments with respect to shipping costs is − 0.764. Our preferred estimate of this elasticity is the more conservative − 0.451. We will use this estimate in the baseline counterfactual exercise, and explore sensitivity to other values. In Hornok and Koren (in press), we have estimated product-level regressions to determine the elasticity of the number of shipments and the average shipment size with respect to per-shipment costs. Countries with higher per-shipment import costs receive fewer and larger shipments from both the U.S. and Spain. The elasticity of the number of shipments is between − 0.262 and − 0.104.23 Our estimates are larger. One possible explanation is that there are many zero trade flows at the product level, which biases a log-linear estimation. Missing trade is not an issue at the country level with a large exporting country such as Spain. Table 5 of Hornok and Koren (in press) also shows that shipments are spread throughout the year: countries with high per-shipment cost receive shipments in fewer months. These empirical patterns motivated our model. We also conducted an empirical analysis of the margins through which exporters change their shipping frequency. Simply put, they may (i) send more of the same good in larger shipments, (ii) pick slower modes of transport that allow for larger shipments and (iii) send bulkier products instead of small products. We do an index-number analysis to decompose the aggregate response into these channels. The results are reported in Appendix C. The main results are that shipping frequency is negatively associated with administrative costs even after controlling for mode of shipping, and that the mode itself does not vary significantly with administrative barriers.","To quantify the effects of administrative costs in the model, we conduct two simple counterfactual exercises. In the first trade facilitation scenario, we reduce per-shipment costs fij by half. The second exercise exploits the cross-country variation in administrative costs. We change the administrative cost of each country to that of the average country in the top income decile. Because poorer countries have higher shipping costs, this scenario affects them more. The average import cost in the top income decile is $942, whereas the average export cost is $913. As seen from Propositions 2 and 3, both the trade volume and the welfare effects are as if bilateral tariffs changed. We hence only need to calculate the tariff equivalent changes, ψij. We can only do this for the trade relations of Spain, because we need shipment-level data to calibrate the ad-valorem equivalent of per-shipment costs. Because of this data limitation and because our semi- structural approach only applies locally, the counterfactual exercises below should be understood as partial equilibrium changes. More specifically, we do not study the effect of the trade facilitation reform on firm profits and the potential spillovers across countries. We can however, characterize changes in consumer surplus and changes in bilateral trade volumes as long as the trade policy changes are small. To calculate ψij, we need to know σ. Following Simonovska and Waugh (2014), we calibrate σ = 4.1. This means that a 1% increase in ad-valorem trade costs reduces trade by σ − 1 = 3.1%. It also implies a 32% markup. We also report results with the estimates of Eaton and Kortum (2002), σ = 8.2. Table 6 reports the average tariff equivalent effects of reducing fij across the ten income deciles. The first column reports the average per-shipment cost in the income decile. Consistent with the evidence presented in Table 2, poorer countries have higher shipping costs. The second column reports the effects of reducing shipping costs by 50%. For example, for the lowest income decile, such a policy would be equivalent to an 11.6% decline in tariffs, whereas for the 9th decile, this would be equivalent to a 5.2% tariff decline. The effects are heterogeneous, because the effect of a shipping cost reduction on the import price index is not linear (recall Proposition 3). Countries with large per-shipment costs enjoy a bigger gain from a 50% reduction. The third column reports the tariff equivalent effect of setting each administrative cost to the average of the top income decile. This effect has a clear tendency with income per capita. Countries in the poorest decile see an effect equivalent to a 13.4% tariff reduction, whereas the average effect for countries in the 9th decile is 1.5% (note that even importers in the top decile gain from Spain reducing its somewhat higher-than-average shipping costs.) For comparison, the last column of Table 6 shows the average bilateral tariff rate with respect to Spain. Recall that only non-EU countries are included in this exercise, hence, even for rich countries, the tariffs are substantial. There is much less variation in statutory tariff rates across income groups than in the tariff-equivalent effects of administrative barriers. Hence a trade facilitation reform equalizing administrative barriers offers a stronger force for convergence than a tariff harmonization reform. Alternative calibrations ~~~~~~~~~~~~~~~~~~~~~~~~ Table 7 reports the average tariff-equivalent effects across non-EU countries for alternative calibrations for the shipment elasticity and the elasticity of substitution. Reducing per-shipment costs by half is equivalent to reducing tariffs by 7.7 to 15.8 percentage points. A larger shipment elasticity (which implies that timeliness is more important) corresponds to a larger gain from administrative barrier reduction. The gain does not depend heavily on σ. The average effects are smaller in the scenario where we match the shipping cost of rich countries, because this corresponds to a smaller than 50% reduction for most countries. The equivalent tariff reductions range between 3.9 and 7.7%. Table 8 reports the average percentage increases in bilateral imports from Spain. Given that Spain is a small trade partner for most of the 131 countries, this exercise ignores third-country effects. Trade volumes go up dramatically after this reduction in per- shipment costs, especially for high σ. With σ = 8.2, the trade creating effect of trade facilitation reform ranges from 31.3 to 148.0%. Even with σ = 4.1, we see a 14.6 to 57.5% trade increase. These magnitudes are comparable to the trade creating effects of customs unions (Table 4). There is a wide distribution of the effects across countries, because they are subject to different per-shipment costs. Fig. 2 plots the tariff equivalent of the per-shipment cost reduction for the cross-section of non-EU countries. For the bulk of the countries, the counterfactual trade facilitation reform is equivalent to 0 to 20% age point reduction in tariffs.","We built a model of administrative barriers to trade to understand how they affect trade volumes, shipping decisions and welfare. Because administrative costs are incurred with every shipment, exporters have to decide how to break up total trade into individual shipments. Consumers value frequent shipments, because they enable them to consume close to their preferred dates. Hence per-shipment costs create a welfare loss. We derived a gravity equation in our model and showed that administrative costs can be expressed as bilateral ad-valorem trade costs. We estimated the ad-valorem equivalent in Spanish shipment-level export data and find it to be large. A 50% reduction in per-shipment costs is equivalent to a 9 percentage point reduction in tariffs. Our model and estimates help explain why policy makers emphasize trade facilitation and why trade within customs unions is larger than trade within free trade areas."],["Households hold vastly heterogeneous amounts of wealth when they reach retirement, and differences in lifetime earnings explain only part of this variation. This paper studies the role of intergenerational transmission of ability, voluntary bequest motives, and the recipiency of accidental and intended bequests (both in terms of timing and size) in generating wealth dispersion at retirement, in the context of a rich quantitative model. Modeling voluntary bequests, and realistically calibrating them, not only generates more wealth dispersion at retirement and reduces the correlation between retirement wealth and lifetime income, but also generates a skewed bequest distribution that is close to the one in the observed data. --------------------------------------------------------------------------------","Why do U.S. households reach retirement with vastly different wealth levels, even when we condition on realized lifetime income? Can we construct a model that explains this fact? What are the important features that help generate this heterogeneity? The answers to these questions help improve our understanding of the factors affecting savings and thus help inform many kinds of policy reforms, including taxation and social insurance reforms. As Heckman et al. (1998) wrote, “It is potentially very dangerous to ‘solve’ problems whose origins are not well understood.” Our main contribution is to study the role of intergenerational links, in the form of intergenerational transmission of ability and both accidental and voluntary bequests, in shaping wealth dispersion at retirement and its correlation with lifetime income. The aggregate importance of wealth transmitted across generations calls for an investigation of the impact of these transfers on individual savings, wealth inequality, and its interaction with earnings shocks over the life cycle. To do so, we study these elements in the context of a rich, quantitative, overlapping- generations model with incomplete markets, which features earnings risks and incorporates the key institutional policies providing public insurance for these risks. The interaction of risks and insurance, in fact, affects saving motives and, thus, observed wealth accumulation in an interesting and non-trivial way. Besides voluntary and accidental bequests and transmission of ability from parents to children, other factors have been shown to be important to understanding saving behavior and cross-sectional wealth inequality in the U.S. However, we know little about the effects of these elements on wealth holdings at retirement. These institutional factors include a government-provided consumption floor (Hubbard et al., 1995), defined benefit pensions, and a more realistic modeling of Social Security rules (Scholz et al., 2006). We explicitly model all of these factors, as they might have important effects directly, but might also interact with the effects of intergenerational links on savings at retirement and, if not accounted for, provide a misleading picture of the effects of intergenerational links. In our framework, after we control for lifetime earnings, wealth differences at retirement arise for the following two reasons. First, when borrowing constraints prevent households from smoothing consumption intertemporally, households that differ in the timing of earnings over the life cycle will have very different wealth levels at retirement. Second, inheritances add another source of wealth heterogeneity among households with similar lifetime earnings. Our calibrated bequest motive implies that bequests are a luxury good and thus increases the desire to leave bequests for households receiving large bequests, thereby increasing their saving rate and leading to more wealth inequality and less correlation of wealth to lifetime earnings. Our main finding is that the model with intergenerational links matches the data well. More specifically, it generates an average of the Gini coefficients for wealth, after controlling for lifetime income, of 0.53, compared with 0.52 in the data; and a correlation coefficient between lifetime earnings and retirement wealth of 0.75, which is a bit higher than in the data (0.62), but lower than in the version of the model that does not allow for intergenerational links. Adding a modest amount of measurement error in earnings and wealth further reduces this correlation to 0.71. Removing the voluntary bequest motive, even in the presence of different assumptions about the receipt of involuntary bequests, increases the gap between the model and the data. To make this point, we compare the benchmark model with all intergenerational links, including voluntary bequests, with two models without voluntary bequest motives, which differ in the timing and size of accidental inheritances receipt. First, accidental bequests of different sizes are received at age 50. Second, accidental bequests of different sizes are randomly received at realistic ages between 35 and 55. The versions of the model without voluntary bequest motives generate average Gini coefficients for wealth that, controlling for lifetime income, are 0.46 and 0.46, respectively, which are lower than the 0.52 observed in the data. They also imply a tighter relationship between lifetime earnings and retirement wealth. More specifically, the correlation coefficients between these variables are 0.81 and 0.80, respectively, compared with 0.62 in the data. In contrast, as one might have expected, the intergenerational transmission of earnings ability from parents to children tends to increase the correlation between wealth and lifetime earnings at retirement. In fact, the model with bequest motives, but without intergenerational transfer of productivity, implies a correlation of 0.73, which is lower than the 0.75 in the benchmark calibration. Without intergenerational transmission of earnings ability, the distribution of inheritance received is independent of households earnings; this implies that poorer people tend to receive comparatively larger inheritances, which helps reduce the correlation coefficient between lifetime earnings and retirement wealth. We also investigate the distributional effects of means-tested minimum consumption programs in old age, Social Security, and pensions. Eliminating the realistically calibrated consumption floor from our benchmark calibration results in the income-poor households having more of an incentive to save to insure against bad income shocks; this tends to decrease wealth inequality and generates a lower average Gini coefficient, after controlling for lifetime income, of 0.49, compared with 0.53 in the benchmark model, while the correlation of wealth at retirement and permanent income barely changes. Removing private pensions raises the average wealth of the households in the high earnings deciles, as they save much more to smooth consumption over their lifetime. As a result, the Gini coefficients decrease a lot at the higher earnings deciles, and the average of the Gini coefficients across income deciles drops to 0.49. For the same reason, the correlation coefficient between retirement wealth and lifetime income increases to 0.79. If, on top of removing private pensions, one were to switch from a history-dependent Social Security system, the wealth Gini coefficients for the lower earnings deciles would increase, while the ones for the highest earners would decrease. On net, the average of the Gini coefficients across income deciles would barely change, while the correlation coefficient between wealth and lifetime earnings would increase to 0.82. The most closely related paper to ours is the one by Hendricks (2007a), who points out the importance of accounting not only for the marginal distribution of wealth, but also for the joint distribution of wealth and earnings. His work first documents key features of these data in the Panel Study of Income Dynamics (PSID) and then studies a life-cycle model in which all households inherit at age 52 and have no information about these inheritances, which are also not correlated with earnings. We add to his contribution by introducing transmission of both physical and human capital across generations. These features, once realistically calibrated, imply that households receive an inheritance at a random time that is consistent with their parent׳s life expectancy; and the size and timing of the inheritance is consistent both with their parents׳ estate and with the distributions of estates left in the U.S. economy. In addition, due to the luxury voluntary bequest motive coming out of our calibration, people who either receive high labor incomes or large bequests hold onto more wealth due to a voluntary bequest motive, which is stronger for richer people, thus generating savings behavior in old age that is more consistent with the observed data (see De Nardi et al., 2010). The paper is organized as follows. Section 2 frames our contribution in the context of the rest of the literature. Section 3 presents the model and the calibration of the model. Section 4 shows the quantitative results of the benchmark model and investigates the role of various features of the model. Section 5 concludes.","Many authors document that households with similar characteristics, such as lifetime income, age, and family structure, hold vastly different amounts of wealth at retirement (see Hurst et al., 1998; Grafova et al., 2006). Hendricks (2007a) shows that the correlation coefficient between lifetime earnings and wealth at retirement (0.61) is much less than unity and that substantial wealth differences remain after controlling for lifetime earnings and age. In fact, the average of the Gini coefficients in wealth at retirement across lifetime earnings deciles is 0.54, compared with 0.62 in the full sample. Several economists (for example, Bernheim et al., 2001; Hendricks, 2007a) argue that these features of the data are inconsistent with most life-cycle models of consumption-saving behavior and thus constitute a challenge to such theories and their policy implications. Our goal is to see how close we can get to the observed data once we allow for a rich model with earnings risk and many features of public insurance, which also includes intergenerational transmission of physical and human capital across generations. An extensive literature, both empirical and theoretical, shows that the transmission of physical and human capital from parents to children is a very important determinant of households׳ wealth in the aggregate economy (see Kotlikoff et al., 1981; Gale and Scholz, 1994), and of both wealth and earnings ability over the household׳s life cycle (see Becker and Tomes, 1979; Hurd and Smith, 2001). As a result, it is also a prime candidate to investigate retirement savings. This paper also builds on a large literature that studies the cross-sectional wealth inequality at all ages (see Huggett, 1996; Quadrini, 2000; Castaneda et al., 2003; De Nardi, 2004, and Cagetti and De Nardi, 2006). It is also related to studies on retirement saving. Engen et al. (1999, 2005), and Scholz et al. (2006) study the adequacy of household retirement saving. Those papers abstract from the intergenerational links of bequests and earnings ability. Gokhale et al. (2001) abstract from voluntary bequest motives. We assume imperfect altruism in the form of warm glow and partial inheritability across generations, and we discuss the effects of removing or changing Social Security on inequality and the correlation of income and wealth at retirement time. Interesting and complementary work by Fuster (1999) and Fuster et al. (2003, 2007) assumes perfect altruism across generations and finds that Social Security programs have important implications for capital accumulation and inequality. Our results under imperfect altruism are broadly consistent with theirs.","The model is a discrete-time, incomplete markets, overlapping generations economy with an infinitely lived government. Demographics, preferences, and labor productivity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Agents also derive utility from the bequests transferred to their children upon death. This form of ‘impure’ bequest motive implies that an individual cares about total bequests left to his/her children, but not about the consumption of his/her children. Calibration ~~~~~~~~~~~ Unless stated otherwise, we report parameters at an annual frequency. Table 1 lists the parameters that are either taken from other studies or can be solved independently of the endogenous outcomes of the model. The tax rate on labor income and the contribution rate to the defined benefit plan, belong to the latter group because, due to the assumption of exogenous labor supply and retirement decisions, they only depend on the earnings shocks and the population demographics, which are exogenous to the model. These bend points are also expressed in terms of average earnings, while the marginal rates are chosen to match the holding of defined benefit wealth relative to Social Security wealth by lifetime earning deciles from the Health and Retirement Study (HRS) data set, as reported in Scholz et al. (2006). Appendix A reports the details of these computations. According to this formula and consistent with the data, households in the first three lifetime earnings categories barely have any defined benefit wealth. The discount factor affects saving and average wealth in the economy. The term ϕ1 measures the strength of bequest motives, thus we choose the aggregate bequest as a moment. The term ϕ2 reflects the extent to which bequests are luxury goods, affecting the bequest distribution, especially the high end of it. To better gauge the quantitative implications of the model, we begin by evaluating the lifetime earnings implications of the exogenous earnings process that we feed into the model. Table 3 first reports the percentage of total lifetime earnings earned at selected percentiles as a fraction of total lifetime earnings generated by the model and then displays the corresponding figures computed from the PSID observed data. The model earnings process produces lifetime earnings earned by each lifetime earnings decile that are very close to those from the PSID data.","We present our numerical results as follows3: The benchmark model, its mechanisms, and its comparison to the actual data. The role of the bequest motive. The role of intergenerational transmission of ability. The role of a government-provided consumption floor. The role of pensions and Social Security. The role of measurement error in both earnings and wealth in the observed data. In each run, we solve for the dynamic programming problem and impose budget balance for the government. We then simulate 100,000 households starting from age 20, drawn from the initial distribution of each model. We define retirement wealth to be the wealth at age 65 and lifetime earnings to be the total earnings from ages 20 to 60, discounted to age 65. The benchmark model ~~~~~~~~~~~~~~~~~~~ Overall, the benchmark model fits the data well. The first subsection regarding the benchmark model discusses its fit of the wealth distribution at retirement and of the wealth distribution across all ages. The second subsection focuses on its implied correlation between wealth at retirement and lifetime income. The third subsection studies bequests. The fourth subsection displays the fit of the model for assets and consumption over the life cycle. Wealth inequality in the benchmark model The top panel of Table 4 reports wealth percentiles at retirement. The first line of the top panel refers to data from the PSID4 and shows that, in the data, wealth at age 65 is highly unevenly distributed. The richest 1% of people at retirement time hold 16% of total retirement wealth, while the richest 5% hold 34% of total net worth at retirement time. The second line of the top panel reports the corresponding numbers for the benchmark model with intergenerational links and bequest motives. That model succeeds in generating a skewed retirement wealth distribution that is comparable with the data, with the exception of the top 1% of the wealth holdings. The bottom panel of Table 4 reports values of the wealth distribution for the whole economy. Wealth for the whole economy is more unevenly distributed than wealth at retirement, both in the data and in all models. This indicates that a large amount of wealth dispersion in the economy is due to differences in age. Here, too, the model fits the data well. Wealth and lifetime earnings in the benchmark model The benchmark model generates a correlation between lifetime earnings and retirement wealth of 0.75. While this is a bit higher than the 0.61 in the PSID data, the model does generate a good deal of wealth inequality even after controlling for lifetime earnings and age. Table 5 illustrates the Gini coefficients of retirement wealth for each lifetime earnings decile. We notice two important features. First, after controlling for age and lifetime earnings, there is still large wealth inequality in the benchmark model: all the Gini coefficients are above 0.32. Second, the degree of wealth inequality declines as lifetime earnings increase, as is observed in the data. Third, the model does generate significant wealth inequality conditional on permanent income. In fact, the average of the Gini coefficients across permanent income quantiles turns out to be 0.53 in the model and 0.54 in the data. To better gauge the amount of wealth dispersion at retirement generated by the benchmark model, Fig. 2 compares the retirement wealth distributions for the 2nd, 5th, and 9th lifetime earnings deciles in the model with those in the PSID. The model successfully replicates the fact that households with similar lifetime earnings hold diverse amounts of wealth. At each lifetime earnings decile, households in the lower wealth deciles hold very little wealth, while households in the higher wealth deciles hold much more wealth. Compared with the data, the model matches especially well for the households in the middle deciles of the income distribution, while it overstates the high wealth percentiles at the 2nd earnings decile and the low wealth percentiles at the 9th earnings decile. Bequests in the benchmark model In the benchmark model, retirement wealth inequality arises among households with similar lifetime earnings because households differ in the timing of earnings over the life cycle and in the amount and timing of inheritances received. Let us now turn to the effects of bequest and inheritance heterogeneity on retirement wealth heterogeneity. The model with intergenerational links of bequests and earnings ability endogenously generates differences in the timing and amount of inheritances. To see how large the variation in inheritances is, Table 6 reports values for the discounted lifetime inheritance distribution. In the PSID, inheritances are highly unevenly distributed, with a Gini coefficient of 0.89, and 50% of the households receive very little or no inheritance. The top 1% of the households receive 35% of all the inheritances. The model generates a skewed inheritance distribution that is comparable with the one in the data, with the exception of its top 1%. Treating bequests as luxury goods and modeling the transmission of earnings ability across generations are essential to match the observed skewness in the inheritance distribution. Several modeling choices and aspects of the calibration are important in generating this result. First, the marginal utility of leaving a bequest is finite at zero bequests, which helps us to generate a large fraction of households receiving no inheritances. Second, some large inheritances are transmitted across generations because of the voluntary bequest motive to hold onto assets if alive at a very old age. Given our calibration, the marginal utility of leaving zero bequests is 0.8027. The average bequest during our model period, which is five years long, is 0.0140, and the average consumption during our model period is 0.3473. Thus, the marginal utility of leaving the average bequest is 0.8017, while the marginal utility of average consumption is given by 4.8859. Hence, for the person consuming the average consumption amount, the marginal utility of consumption is much higher than the marginal utility of leaving even a small bequest. However, because the marginal utility of bequests declines more slowly than the marginal utility of consumption, the richest households have strong bequest motives to save in order to leave some assets to their children, even when they are very old. Moreover, in the presence of a positive correlation between parents׳ and children׳s earnings, children are more likely to be earnings-rich and to receive a large inheritance and thus tend to save to leave more wealth to their offspring, hence generating a skewed inheritance distribution. Table 7 shows the fraction of lifetime inheritance received by households in each lifetime earnings decile. In the PSID, there is a positive correlation between lifetime inheritance and lifetime earnings. The benchmark model generates an increasing relation between inheritance and lifetime earnings. Modeling the transmission of earnings ability across generations and a highly correlated lifetime earnings process is key in generating this pattern. The monotonicity relation is weaker in the data. This might be due to the fact that in reality, individuals differ by numbers of siblings and marital status and each married couple might receive inheritances multiple times. More inheritances at the lowest earnings decile might help the model to generate less correlation between retirement wealth and earnings. Interestingly, assuming that the parent׳s productivity and his child׳s initial productivity are perfectly correlated increases the discrepancies between the fraction of inheritances received by each lifetime earnings decile in the model and the data. The reason is that higher intergenerational persistence in productivity generates a stronger positive correlation between lifetime inheritance and lifetime earnings. As is shown in Table 7, this implies that the fraction of inheritances received by the lowest two earnings deciles is only 4.7%, compared with 14% in the data and 10% in the benchmark model. We now turn to the effects of inheritance heterogeneity on retirement wealth in the benchmark model. The upper panel in Table 8 shows extra wealth holding at selected percentiles among those who did inherit, compared with those who did not. At each earnings decile, those who never inherited hold less wealth than those who have inherited, and the difference increases as the wealth percentile increases. The reason is, with operative bequest motives, those who have inherited hold a large part of the inheritances at retirement. The top panel of this table thus documents the role of bequest receipt on savings by earnings and wealth decile. The bottom part of the table, labeled “Random inher. at age 50” changes the distribution of inheritances and studies its effects on savings; this exercise highlights that the extra wealth holdings among those who have inherited are much smaller in the random bequest model than in the benchmark model with bequest motives. Consumption and wealth accumulation over the life cycle The goal of this subsection is to show that our benchmark calibrated model can also generate the life-cycle profiles of asset holdings and consumption observed from the U.S. data. Fig. 3 displays the life-cycle profiles of average assets and consumption for both the model and actual data. Both consumption and wealth are expressed as a fraction of annual output per capita and are scaled so that the data and model coincide at age 55. The consumption data is expenditure and comes from Dotsey et al. (2012). The fit of the model is good, and in particular, the model generates a hump-shaped consumption profile. The wealth data comes from Cho (2012). Here, the model also does a good job of matching the hump-shaped profile of wealth accumulation and the relatively slow decumulation after retirement, although the model understates the retirement peak and predicts more decumulation than the data, despite the presence of a bequest motive. The role of the bequest motive ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To isolate the role of the bequest motive from the distribution of inheritances, we compute two versions of the model without bequest motives in which the inheritance distribution is exogenously taken from the data (see Appendix B for more details): All 50 year olds receive an inheritance of random size (taken from the data), whose size is uncorrelated with their earnings at age 50, as in Hendricks (2007a). The timing of inheritance receipt is random and consistent with the data. Table 9 reports several measures of inequality for the PSID data and the model-generated data in various versions of the model. The benchmark model with intergenerational links and bequest motives generates a skewed wealth distribution comparable with the one in the data. In contrast, the two models without intergenerational links and bequest motives generate much less wealth concentration than in the observed data. In the model with random bequests that are received at age 50, the correlation coefficient between lifetime earnings and retirement wealth and the mean Gini coefficient are farther away from the data (see Table 10) compared with the benchmark model, thus pointing out that a bequest motive is needed for those who have received large inheritances to hold a significant portion of them by the time retirement comes around. In the benchmark model, in fact, households hold onto more wealth due to operative bequest motives, and this generates more heterogeneity in retirement wealth for given lifetime earnings. As we have seen in Table 8, the extra wealth holdings among those who have inherited are much smaller in the random bequest model than in the benchmark model with bequest motives. Table 10 also shows that making the timing of bequest receipt more realistic further weakens the relationship between lifetime earnings and retirement wealth: a borrowing-constrained household that receives an inheritance earlier consumes more of it and holds onto less wealth at retirement than an otherwise identical household that receives an inheritance later. However, the quantitative effect of this factor is very small. The role of intergenerational transmission of ability ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To study the interaction of bequest motives and intergenerational transmission of ability, we keep the bequest motive, but we change the intergenerational persistence in earnings, which in turn endogenously changes the distribution of inheritance. In particular, if there is no intergenerational link of earnings, inheritances are evenly distributed by lifetime earnings decile. More generally, higher intergenerational persistence of earnings ability leads to more wealth accumulation across generations and increases wealth inequality, as more productive parents save more to leave bequests and have more productive children, who in turn also save more to share their wealth with their descendants. The role of the government-provided consumption floor ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 11 shows that setting the government-provided consumption floor to zero reduces the mean Gini by 0.04 to 0.49, because a means-tested minimum consumption guarantee provides strong incentives for low-income individuals not to save. Due to the persistence of the earnings process, those low-income households are more concentrated in low lifetime- earnings deciles. Table 12 shows that without a consumption floor, the Gini coefficients decrease a lot at the lowest three earnings deciles, since there are fewer poor households in those deciles. When the consumption floor is removed, the government adjusts the tax rate on labor to balance its budget. As a result of this change, the wealth Gini coefficients in the higher income deciles barely change, because the households in those deciles are too rich to qualify for a consumption-floor transfer. In addition, the change in the labor tax needed to rebalance the government budget is very small, because the government outlays due to the consumption floor are very small. The role of pensions and Social Security ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Tables 11 and 12 show that eliminating pension benefits raises the correlation coefficient between lifetime income and retirement wealth substantially, because it raises average wealth for households in the high earnings deciles. Compared with the benchmark model with pensions, the Gini coefficients decrease a lot at the highest five earnings deciles, since these households now increase saving for retirement. We then further eliminate any kind of Social Security payments (in addition to removing pension benefits). The fact that the mean Gini in Table 11 is the same in the benchmark model and in the model without Social Security and pension hides substantial changes in the Gini coefficient by lifetime earnings deciles. Table 12 shows that the Gini coefficients at the lower earnings deciles increase: wealthy households increase savings while poor households do not increase their saving, allowing them to qualify for the government-provided consumption floor. The Gini coefficients at the higher earnings deciles decrease since poor households, who are not likely to qualify for the government-provided consumption floor, increase saving relatively more than rich households. This comparison shows that Social Security programs have important implications for inequality. We provide further evidence in Table 13. Without pensions and Social Security, the Gini coefficient for the retirees at age 65 is much lower. As an additional test, after eliminating Social Security payments, we give the same Social Security payment to all retirees, in addition to removing private pensions. As a result of this change, the resulting Social Security payment to all retirees, determined from the government׳s budget balance, is large enough that no retiree qualifies for a minimum consumption-floor transfer anymore. This experiment decreases the Gini coefficients at the lower earnings deciles because rich households reduce saving while poor households, holding no wealth before the experiment, do not reduce saving. The Gini coefficients at the higher earnings deciles increase, because the poor households decrease saving relatively more than the rich households. The role of measurement error in income and wealth ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The observed log wealth at age 65 follows Adding a small amount of measurement error reduces the correlation coefficient from 0.75 in the benchmark model to 0.71. The mean Gini coefficient of retirement wealth is slightly higher (0.55) than that in the benchmark (0.53). The reason is that adding random noise in earnings affects Gini coefficients only when enough agents are regrouped into other deciles. Given the wide range of earnings in each decile, not many agents are regrouped.","Empirical studies using micro data find that there is large heterogeneity in retirement wealth among households with similar lifetime earnings and raise doubts about the ability of a standard life-cycle model of saving behavior to reproduce the observed facts. We use an incomplete-market life-cycle model with intergenerational links of bequests and earnings ability, a government-provided consumption floor, a history-dependent Social Security system, and a defined benefit pension. We show that this model with earnings heterogeneity and inheritance heterogeneity generates a substantial amount of heterogeneity in retirement wealth for given lifetime earnings. This suggests that a properly specified life-cycle model with bequest motives, consumption floor, and Social Security pensions captures the fundamental determinants of households saving and wealth accumulation. This framework might shed light on the effects of policy reforms that affect saving. We show that government-provided minimum consumption, pensions, and Social Security have very different distributional effects. In a separate paper, Yang (2013) studies the consequences of eliminating Social Security in a similar environment to the one constructed in this paper and finds that the presence of bequest motives reduces life- cycle saving and thus reduces the gains from Social Security reform. We have assumed that households are ex-ante identical. Interesting directions for future work include allowing for heterogeneity in wealth holdings by education (Hubbard et al., 1995; Cagetti, 2003) and by marital status (Cubeddu and Rios-Rull, 2003 and Guner and Knowles, 2007). Other interesting extensions include allowing for heterogeneity in preferences (Krusell and Smith, 1998; Samwick, 1998; Hendricks, 2007b) and in self-control (Ameriks et al., 2007), in number of children (Scholz and Seshadri, 2007), in rates of return (Guvenen, 2006), and in health expenditures (De Nardi et al., 2010)."],["In this paper, we investigate extreme events in high frequency, multivariate FX returns within a purposely built framework. We generalize univariate tests and concepts to multidimensional settings and employ these novel techniques for parametric and nonparametric analysis. In particular, we investigate and quantify the co-dependence of cross-sectional and intertemporal extreme events. We find evidence of the cubic law of extreme returns, their increasing and asymmetric dependence and of the scaling property of extreme risk in joint symmetric tails. © 2014 The Authors. --------------------------------------------------------------------------------","Elliptical distribution of asset returns is a convenient assumption in finance that allows for numerous applications including risk management, asset and option pricing and portfolio decisions. In particular, the standard practice with many risky assets is to assume that the density is multinormal with perhaps a time-varying covariance matrix (see Diebold et al., 1999). However, Leon et al. (2009), inter alia, find that joint normality is not supported by empirical evidence. Moreover, correlation – the standard measure of dependence in the multivariate context – is only a measure of linear dependence and suffers from a number of limitations (see Embrechts et al., 2002). For example, Patton (2004) argues that the dependence between assets is stronger during market downturns than during market upturns and here we find that the probability of extreme co-events varies over time. These deficiencies are compounded in the covariance measure – an essential input in many financial applications including hedging, portfolio selection and systematic risk. While there is a natural interest in the joint density of asset returns, in some cases a more focused approach is required. Since the density's interior characterizes small day-to-day disturbances, it may be of substantially less concern to financial institutions, managers and regulators than the tail behaviour. For this reason, attention has recently shifted to exceedance measures that account for the expected magnitude of large movements in the underlying variables of interest such as stock returns, interest or exchange rates and changes in energy prices and GDP (see, for example, Longin and Solnik, 2001; Butler and Joaquin, 2002; Bae et al., 2003). In particular, Berkowitz (2001) proposes a censored likelihood test, in which the observations not falling into the negative tail of the distribution are truncated. While in the univariate setting, the tails of financial processes have been studied at length (see, for example, Adler et al., 1998), the literature on multivariate tail analysis is still in its infancy. Employing results from the multivariate Extreme Value Theory (EVT), Hsing et al. (2004) fit specific copulas on the bivariate densities (see also Jondeau and Rockinger, 2006; Brodin and Kluppelberg, 2010; Ning, 2010). Then, employing coefficients inherent in these copulas, they examine the dependence between returns in the tails. However, recent attempts to generalise this framework to the multivariate case turned out to be technically and computationally demanding (see, for example, Aas et al., 2009; see also Diebold et al., 2000 for some commonsense caveats to uncritical use of EVT). In the theoretical part of this study, we develop novel techniques tailored to multidimensional data. The proposed techniques employ a natural generalization of the VaR concept. Multidimensional Value at Risk (MVaR) is the region of the intersection of univariate VaRs with a nominal probability mass under a given density function. MVaR is defined by a single cut-off value and a directional vector. Despite its conceptual simplicity, MVaR is a versatile framework that allows for simple testing of the tails as well as the overall accuracy of a multidimensional density forecast (MDF). Moreover, it is straightforward in this framework to examine extreme co-events and to evaluate dependence in risk. In the empirical part of this paper, we illustrate the forecast evaluation technique by investigating two of the most widely used elliptical distribution functions to approximate the empirical joint distribution of (high frequency) FX returns overall and in the tails. Using a rich set of synchronized returns, we examine further the co-dependence among FX returns in the tails. We analyse and quantify both, cross-sectional and intertemporal dependence of the extreme returns. Our empirical study covers several return frequencies providing a rich picture of FX (extreme) returns. Further, to the best of our knowledge, this is the first study to investigate and present clear evidence of the scaling law for multivariate extreme returns. This is important because the tails of a fat-tailed distribution are invariant under addition although the distribution as a whole may vary according to temporal aggregation (Feller, 1971). For instance, if daily returns are i.i.d. t-distributed, then the central limit law implies that weekly returns converge to the normal distribution. However, the tails of the weekly return distribution behave like the tails of the daily returns with the same tail index. The outline of the remainder of this paper is as follows. In Section 2, we discuss the concept of MVaR, its accuracy evaluation and dependence in risk. This section contains also a short discussion of the policy and practical significance of the MVaR framework. Section 3 summarises the high frequency FX dataset employed and presents the results of our empirical studies. Finally, Section 4 summarises the main findings and offers some concluding remarks.","In this section, we develop a formal framework for investigating distributional characteristics and co-dependence of multidimensional variables. Joint density tails ~~~~~~~~~~~~~~~~~~~ Fig. 1 illustrates a JDT in the two-dimensional space. A proof of property (4) is included in the Appendix. As we will see in what follows, the above property of projection (3) allows for a systematic accuracy evaluation of MDFs over the tails and also overall. Note that there is an infinite number of directional vectors d and hence MVaRs. Statistically, there may be no reason to prefer one over the others. Economically, however, the directional vector implicitly defines the proportion of the assets in the portfolio. For example, the symmetric directional vector d = −(1, …, 1) captures negative events for a portfolio with equally weighted assets. This vector is of main relevance for an investor holding such a portfolio. Multidimensional value at risk ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Therefore, if f is the true data generating process, we can identify observations lying in the extreme tails by simply inspecting their z-scores. Dependence in risk ~~~~~~~~~~~~~~~~~~ Poon et al. (2004) and Hartmann et al. (2010) employ tail dependence coefficients to examine extreme events in the equity and FX markets, respectively. However, these coefficients, while theoretically robust, measure only asymptotic dependence between marginals of a joint distribution (see Heffernan, 2001 for a directory of tail dependence coefficients). By contrast, the dependence measures, which we propose below as natural extensions of the MVaR framework, can capture dependence between more complicated events (e.g., between two disjoint sets of marginals) with vanishing or non-vanishing probabilities. Some applications of MVaR framework ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The recent financial crisis brought to the forefront of attention systemic risk. Due to the interconnectivity of the financial institutions, a shock faced by one institution in the form of an extreme event, increases the probability other financial institutions experiencing similar extreme events, (see Nijskens and Wagner, 2011). Recently, there has been increasing concern among researchers, practitioners and regulators over the evaluation of models of (financial) risk (see, for example, the report on Global Risks 2012 by the World Economic Forum). Moreover, while it is important to have an aggregate measure of the total risk, often it is also important to know the direct dependence on, and inter-relationships of, the specific sources of risk. These developments accentuate the need for modelling and evaluation techniques that are flexible and yet powerful (Lopez and Saidenberg, 2000). However, while the literature on aggregating the multiple sources of risk is gaining momentum (see, for example, Rosenberg and Schuermann, 2006), there appears to be little research into the joint distribution and evaluation of such risks. By focussing on the joint distribution, the MVaR framework measures not only the risk inherent to each source but also the co-dependence of these risks. Important innovations in the derivatives markets include basket and rainbow options whose payoffs depend on the value of a basket of assets. Pricing basket options is difficult as the underlying portfolio is a function of the constituent asset prices. There are basically two approaches to address this issue. The first is by modelling the (unidimensinal) distribution of the basket value (e.g., Borovkova et al., 2007). The second approach, which is perhaps more intuitive, is by focusing directly on the joint density of the basket's constituent assets. For example, Huang and Guo (2009) price basket (Bermudan) options as a function of the value of the option in each state of the basket's constituent assets times the joint probability of the assets being in that state. They estimate the joint probability of each state using copulas. By offering a simple and versatile method for an efficient evaluation of joint density estimates, MVaR framework is readily adapted to assess theoretical or empirical probabilities that can then be used for multidimensional option pricing.","In this section, we present the results of our empirical studies that employ the techniques discussed in Section 2. The data was provided by Bank of America and covers the period from 2 January 2001 to 29 December, 2006. It consists of synchronized 1-min exchange rates for three currencies, EUR, GBP and CHF against the USD (a total of 1,006,544 observations). The FX market operates continuously from 10.00pm GMT on Sunday to 10.00pm GMT on Friday, so there are a total of 1440 observations in each 24-h window while the market is open. Summary statistics for the three log returns at 8 and 64 min frequencies are reported in Table 1. For all three series, daily log returns are leptokurtic but the departure from the normal kurtosis diminishes for lower frequencies. The ARCH(4) test for up to fourth order serial correlation in squared returns shows that all three currency return series display significant volatility clustering. All mean returns are close to zero and there is a strong positive correlation between each pairs of the series that increases as the frequency decreases.4 The strong positive correlation is clearly visible in Fig. 4. In the first part of the experiment, we test the accuracy of two parametric distributions – the multinormal (MN) and the multivariate-t distribution (MT) – over multivariate tails. Both specifications are time-varying. Below we discuss the details of the dynamic estimation of the parametric distributions via a multivariate GARCH model. Multivariate GARCH model and estimation method ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To obtain forecasts of the time-varying three-dimensional covariance matrix we employ the simplified GARCH (S-GARCH) model of Harris et al. (2007). Our choice can be explained on the basis of the ability of this model to handle t-distributed residuals and its ease of estimation. However, we experimented with other, widely used multivariate GARCH models such as BEKK and DCC but the estimation did not always converge. Like many other multivariate GARCH models, the S-GARCH does not guarantee that the conditional variance- covariance matrix is positive semi-definite. However, for all three pairs, the estimated correlation coefficients were found to be between −1 and +1 for all observations. We ML- estimate the S-GARCH parameters (and the degrees of freedom parameter for the t-distributed residuals) using the entire sample. We then use these estimates to obtain the one step-ahead forecast of the covariance matrix. Empirical results ~~~~~~~~~~~~~~~~~ First, we focus on the goodness-of-fit test of the MN and MT specifications over JDTs that are defined by the selected directional vectors and nominal probabilities. To this end, we compute from the observations in the relevant tails the normalized z-scores and the corresponding p-values of the uniformity test. The results of the experiment for returns at 8 min frequency are reported in Table 2. Similar results were obtained at other frequencies, for which a sufficient number of tail observations was available. While we strongly reject multinormality for all tested JDTs, there is some support for the multivariate-t distribution in symmetric tails with probability mass 5 percent or less. As all p-values for these tails are larger than 5 percent, we would not reject the null that the corresponding tail observations were drawn from the multivariate-t at this significance level. The (cumulative) distribution of asset returns has frequently been assumed to be Gaussian, Lévy or a truncated Lévy distribution, where the tails become “approximately exponential”. In contrast, Gopikrishnan et al. (1998) find that the asset return distribution exhibit a strong cubic power-law behaviour which differs from all three previous models: unlike the Gaussian or the truncated Lévy distribution, it has diverging higher moments, and unlike the Gaussian or Lévy it is not a stable distribution. Our results in Table 2 together with our estimate of 2.75 for the degrees of freedom (i.e., relatively thick tails), can be seen as a generalization of the cubic law for extreme returns to the multivariate context. However, our tentative evidence of the cubic law is limited to symmetric tails with probability mass not exceeding 5 percent. For the intertemporal dependence for EUR/USD exchange rate, i.e., the co-dependence between two consecutive EUR/USD returns,7 the co-dependence in Table 4 appears to decrease in α and it is stronger for higher frequencies. The overall significant intertemporal dependence is a hallmark of volatility clustering and indicates that clustering becomes stronger for more extreme returns. The patterns emerging from Tables 3 and 4 have been confirmed for other bi-dimensional tails and frequencies. It appears that the relationship between the CMVaR and the risk dependence coefficient on the one hand, and alphas on the other, is relatively consistent and stable across the three tails and in both cross-section and intertemporal frameworks. The practical implications of these findings seem to be that, at least for symmetric tails, we can estimate the multidimensional risk at high frequencies and then, scale this estimate up to match the desired return time interval, which is generally daily, weekly or monthly.","We develop a formal framework for investigating the distributional characteristics of multivariate variables with a particular focus on the tails. We extend important unidimensional risk concepts to the multivariate settings and employ these to examine risk dependence for a rich set of high frequency FX data. Our investigation into the tails of high frequency multidimensional returns reveals interesting phenomena, such as asymmetry of dependence in the positive and negative tails, cubic law of extreme multidimensional returns and risk scaling in the symmetric tails. Generally, MVaR framework seems to be a versatile approach, which could be applied also in portfolio decisions and to study systemic risk. We intend to pursue these avenues in future research."],["We show that a model of an advanced small open economy with exports in commodities and search and matching frictions in the labour market can match the impulse responses from a panel vector autoregression to an identified commodity price shock. Using a minimum distance strategy, we find that international financial risk sharing is low even for advanced small open economies. Moreover, a strong real exchange rate appreciation is key for an unexpected commodity price increase to induce a tightening of labour market conditions in the model that is in line with the empirical evidence. As in the case of technology shocks discussed by Shimer (2005), proper amplification of the commodity price shock requires a high value of the outside option for unemployed agents. However, vacancies and unemployment hardly respond whenever the real exchange rate channel is mute. These findings suggest the relevance of the open economy dimension for the transmission of demand-type shocks to the labour market more generally. --------------------------------------------------------------------------------","Wealth effects play an important role in the international transmission of shocks and thus the international business cycle. Corsetti et al. (2008) demonstrate how even under a technology shock the associated wealth effect can dominate the substitution effect and induce an appreciation of the terms of trade rather than a depreciation as commonly argued. In other words, wealth effects, if strong enough, can fundamentally alter the international transmission mechanism.2 Commodity price shocks are one of the most common sources of wealth effects. We use the fact that commodity price shocks can be readily identified in a panel vector autoregression (PVAR) framework to study their effects on labour market dynamics in a model of an advanced small open economy with exports in commodities and search and matching frictions. A strong real exchange rate response is key for the commodity price shock to transmit to the labour market in line with the data. Small open commodity-exporting economies and the consequences of fluctuations in the relevant commodity prices are rarely the focus of academic studies. Instead, Mendoza (1995), Kose (2002), Schmitt-Grohé and Uribe (2018) and many others studied the effects of exogenous terms of trade shocks on business cycle fluctuations rather than the more fundamental concept of commodity prices. The assumption of exogenous terms of trade shocks might be appropriate in the context of developing economies with a small and homogeneous set of exportable goods. However, for the case of advanced commodity-exporting economies this assumption fails to account for the differences between the commodities sector and the non-commodities traded goods sector: improved terms of trade may, for example, reflect increases in commodity prices or productivity losses in the traded goods sector and may thus provide no clear signal about output, consumption, or the labour market. Distinguishing between commodity prices and the terms of trade is not only interesting in its own right. Given the small and price-insensitive share of employment in the commodity- producing sector in our sample countries— the share ranges from 2 to 6% of overall employment— commodity price shocks resemble a pure wealth transfer and are a prime example of a demand shock. Our panel includes data from Australia, Canada and New Zealand, all of which are net exporters of commodities with high quality data on unfilled vacancies, hours worked, and unemployment. In the PVAR, a shock that raises commodity prices expands all components of GDP, notably consumption and investment, leads the real exchange rate as well as the terms of trade to appreciate while the trade balance to GDP ratio turns positive. Furthermore, labour market conditions tighten as evidenced by a drop in the unemployment rate, an increase in individual hours worked, and a larger number of unfilled vacancies. Applied to the case of the Great Recession, the estimates suggest that the collapse in commodity prices accounted for up to one third of the rise in unemployment in the sample countries. The impulse response functions from the PVAR can be successfully matched by the theoretical responses from a standard small open economy model with net exports in commodities and search and matching frictions in the labour market as in Diamond (1982), Mortensen (1982), and Pissarides (1985). This exercise provides insights into the transmission of commodity price shocks, the role of international financial linkages and the empirical performance of the search and matching model. First, even advanced commodity-exporting economies are far from sharing risk efficiently through international financial markets. For our model to match the empirical size and dynamics of the response in the trade balance to GDP ratio and the real exchange rate after a commodity price shock international financial risk sharing needs to be low in the model. Moreover, matching the empirical movements in the real exchange rate is essential to obtain data-congruent movements in unemployment and vacancies. If financial markets were better used to smooth consumption, or if commodity price shocks were absorbed by institutional arrangements such a sovereign wealth fund, the real exchange rate response would be muted to such an extent that the model could not replicate the dynamics of the labour market.3 Second, commodity price shocks, as well as demand shocks in general, transmit to the labour market through their impact on the real exchange rate. We show this by decomposing labour market tightness (the ratio of unfilled vacancies and job searchers) into a labour productivity channel and an international channel that captures movements in the real exchange rate. If the economy is closed or if the real exchange rate response to commodity price shocks is limited, the response of the labour market to demand-type shocks is small in particular in the near term given that labour productivity adjusts only slowly and by a small amount to demand-type shocks.4 Third, the reason why it is the real exchange rate and not the real price of commodities that matters for the transmission of commodity price shocks lies in the low share of employment in the commodity-producing sector. A rise in commodity prices must transmit to the non-commodity sector of the economy to account for the magnitude of the total employment response in the PVAR. In the model, a positive shock to commodity prices raises domestic demand for non-commodity goods. The accompanying real appreciation provides an incentive to the producers of non- commodity goods to post additional job vacancies causing the number of job matches to go up. As a result, employment in the non-commodity sector of the economy rises, and overall unemployment falls even if employment in the commodity-producing sector of the economy is negligible. Fourth, when distinguishing between traded and non-traded goods our analysis suggests that the countries in our sample suffer from the Dutch disease effect analysed in Corden and Neary (1982). In this extension, the employment increase caused by higher commodity prices is always concentrated among the firms in the non-traded goods sector. Employment in the traded goods sector falls— the Dutch disease effect— for standard choices of the elasticity of substitution between traded and non-traded goods and the elasticity of substitution between domestic and foreign traded goods, but it could rise under less conventional choices. As the model fits the empirical impulse response functions better under standard values of the substitution elasticities, we feel comfortable concluding that employment in the domestic traded goods sector falls even absent quality data on non-traded and traded goods. Finally, commodity price shocks can be identified with less controversy than technology shocks and have robust implications for the labour market.5 Small open economies have negligible impact on world commodity prices; a recursive identification scheme with commodity prices ordered first appears to be a defensible specification that leads to a yardstick against which to assess the performance of theoretical models. The analysis of commodity price shocks identifies the amplification of labour market tightness in the model as potentially insufficient as did Shimer (2005) for the case of technology shocks. We address this issue through the aforementioned real exchange rate channel as well as a preference specification that allows for a consumption differential between employed and unemployed agents which in turn determines the value of the outside option for the unemployed agents.6","Among developed economies, net exports of commodities are significant only for a small set of countries. According to the IMF (2012), net exports of commodities account for more than 30% of total exports in Australia, Iceland, New Zealand and Norway, and around 20% in Canada. Furthermore, commodity net exports account for 5% to 10% of GDP on average. We exclude Iceland and Norway from our analysis. For Iceland, data availability is limited and for Norway, the presence of its sovereign wealth fund significantly alters the macroeconomic response to a commodity price shock compared to the other countries. Data description ~~~~~~~~~~~~~~~~ We estimate a panel vector autoregression (PVAR) using quarterly data for Australia, Canada and New Zealand spanning from 1994 Q3 to 2013 Q4.7 For each country, we include a trade-weighted real commodity price index, expressed in US dollars. Nine country-specific macroeconomic time series complete the dataset: GDP per capita, consumption per capita, investment per capita, the unemployment rate, unfilled vacancies, net exports of goods and services relative to GDP, the real effective exchange rate, the real wage deflated by consumer prices, and hours worked per capita. With the exception of the unemployment rate and net exports, the data are transformed into logs. All data are de-trended. For the baseline PVAR, we subtract a quadratic trend from the data. To assess the importance of trade in commodities, consider the most important commodity for the three countries in the sample, based on net exports as reported in the Harvard Atlas of Economic Complexity (2014 data). Iron ores and concentrates account for 33% of Australia's net exports. In Canada exports of crude oil account for 30% of net exports, while in New Zealand's 24% of net exports are accounted for by milk concentrates. The countries in our sample may be considered to be important players in selected commodity markets. For example, Australia is the world's largest exporter of iron ore. However, Australia's total production of iron ore is significantly below China's and world production of iron ores. Similarly, New Zealand's large exports of milk products (24% of net exports are accounted for by milk concentrates and another 7% by butter) pale in comparison to milk production in India, the United States, and the European Union. Economic areas with large domestic markets produce and consume a significantly larger share of commodities, but export less. As we focus on a country's role across all commodities and we use country-specifictrade-weighted commodity prices, assuming price taking behaviour in commodity markets for the countries in our sample appears justified. For all three countries, commodity prices have experienced high volatility over the sample. Relative to real GDP, commodity prices are between 5 and 21 times more volatile. Estimation strategy ~~~~~~~~~~~~~~~~~~~ The effects of commodity price shocks on the labour market are estimated using a PVAR approach. As the relevant time series are short, but the countries in our sample experience commonalities in their economic structure, combing the data across countries can improve the quality of the coefficient estimates. Furthermore, estimation of a panel provides a single benchmark for matching the impulse response functions implied by theoretical models to their empirical counterparts. The factor A(L) ≡ A0 + A1L + A2L2 + … denotes a lag polynomial where L is the lag operator. The vector ui, t summarises the mean-zero, serially uncorrelated exogenous shocks with variance-covariance matrix Σu and μi denotes the country fixed effect. The lag length is set at 2 in our baseline.8 The prices of the commodities traded by the countries in our sample are determined in the world markets. Commodity price shocks are identified through a recursive identification scheme. With commodity prices ordered first in the Cholesky decomposition, country- specific shocks are ruled out from affecting commodity prices contemporaneously. However, domestic developments in our sample countries can in principle feed back into the world market at all other horizons.9 Estimation results ~~~~~~~~~~~~~~~~~~ Fig. 1 plots the median impulse responses, the black solid lines, together with the 90% confidence intervals of the PVAR to a one-standard-deviation increase in commodity prices. The shock to commodity prices is both hump-shaped and persistent. The median response of commodity prices indicates a rise by about 6% by the second quarter. Commodity prices return to trend after 12 quarters. Rising commodity prices lead to a boom in the commodity-exporting economies. Output, consumption and investment rise on impact. Output and investment increase gradually and peak at about 0.15% and 0.75%, respectively. The increase in private consumption reaches 0.12% after two quarters. In line with the Harberger-Laursen-Metzler prediction, the net exports position improves by as much as 0.3% of GDP. The measure of the real exchange rate appreciates following an increase in commodity prices, thus increasing the international purchasing power of domestic households and firms. Labour market conditions improve on impact and continue to do so beyond the rise in commodity prices. At its peak, the median response of vacancies reaches almost 3% and the unemployment rate drops by 10 basis points. CPI deflated real wages decline on impact but recover quickly, whereas hours worked rise by somewhat less than 0.25% by the third quarter following the shock. To put these magnitudes into perspective, when applied to the case of the Great Recession, our estimates suggest that the collapse in commodity prices accounted for up to one third of the overall rise in unemployment in our sample countries. At that time, broad commodity price indices declined by 30–50% while the unemployment rates rose by 2.5 percentage points on average. In the working paper version, Bodenstein et al. (2016), we report a number of robustness checks such as the impulse responses from PVARs estimated with data transformed by a linear as well as the Hodrick-Prescott filter. The shape and the magnitude of the impulse response functions appear robust to the de-trending method. We also examined sensitivity to changes in the lag length as well as to the assumption of block-exogeneity of commodity prices. In all of these cases we obtained similar results as under the baseline specification of the PVAR. In Section 5 we also analyse the case of less aggregated employment data. The effects of an increase in commodity prices on commodity net exporters mirror those found for developed countries that are net importers of commodities. For example, Blanchard and Galí (2007) report that a shock that raises the price of oil unexpectedly leads to a contraction in economic activity with GDP falling and unemployment rising. As in other recent studies, commodity price shocks have a significant, yet quantitatively modest effect on domestic economic activity in countries that are net exporters of commodities. After adjusting the magnitude of the shock, Pieschacon (2012) finds that for Norway an 8% increase in the price of oil pushes up private consumption by 0.2%.10 The qualitative movements and overall magnitudes of the non-labour market variables are also comparable to those in Schmitt-Grohé and Uribe (2018) for shocks to the terms of trade rather than commodity prices in less developed economies. These studies, however, do not report results for labour market variables.","We are interested in understanding the economic channels through which a commodity price increase induces the persistent fall in unemployment and the lasting increases in unfilled job vacancies, consumption and investment found in the empirical analysis. To this end, we augment a standard small open economy model with net exports in commodities and search and matching frictions in the labour market as in Diamond (1982), Mortensen (1982), and Pissarides (1985)— the DMP framework. Our empirical findings have some immediate implications for a theoretical model. First, an increase in commodity prices raises the revenues from commodity exports. If the increase in revenues induces a strong (negative) wealth effect on the labour supply in the form of hours worked, then employment, investment, and non-commodity output could contract depending on the importance of capital and labour in producing commodities in the short-term. Such wealth effects on hours worked after a commodity shock are ruled out in our framework by specifying preferences as in Greenwood et al. (1988). Second, the response of the trade balance to GDP ratio suggests that the countries in our sample are limited in their capacity to share risk in international financial markets. Under a low supply elasticity for commodities, the commodity price increase (fall) constitutes a pure wealth transfer to (from) the commodity-producing country. If financial markets were complete in the sense of Arrow and Debreu (1954), these transfers would be very small and would have a negligible impact on the domestic economy. As the salient features of the model are well understood and documented in the literature, we simply summarise the model equations in Table 1 and confine our discussion to those aspects of the model germane to our analysis—the DMP framework and the commodity-producing sector. The households' first order conditions with respect to consumption, the marginal value of finding a match as well as the definition of total consumption of working and unemployed households are given in Eqs. (i) to (iv) in Table 1. The firms' first order conditions and constraints are Eqs. (v) to (x). Trade in goods and assets is defined in Eqs. (xvii) to (xxiv). The model equations pertaining to the flow of labour are summarised in Eqs. (xxv) to (xxxii). Finally, Eqs. (xi) to (xvi) relate to the firms' adjustment cost functions for investment and vacancies. A detailed model description is found in the working paper version. Commodities ~~~~~~~~~~~ The profits from commodity extraction, πtc = ptcytc, are distributed to the households. The data suggest that resource-intensive commodity extraction and endogeneity of the supply response are not of primary importance for our anlaysis. Fig. 2 panel (a) shows the share of employment in the commodity-producing sector in total employment. This share ranges from 3% and 6% for the countries in our sample. Moreover, despite large swings in commodity prices over the sample period, this share has been roughly constant in Australia and Canada, and it has been gradually falling in New Zealand. In panels (b) and (c), we examine the employment dynamics in the commodity-producing sector following a commodity price shock in alternative specifications of the PVAR. Panel (b) displays the impulse responses when we included both commodity and non-commodity employment as observables in the PVAR. Following a commodity price shock employment in both the commodity-producing and the non-commodity sectors increase, but the increase in total employment stems primarily from the non-commodity sector. Moreover, employment in the non-commodity sector rises immediately after the shock, whereas employment in the commodity-producing sector only increases after 2 quarters. If the expansion in total employment was driven by the commodity-producing sector, timing and magnitude of the responses should be opposite from what is shown in panel (b). Alternatively, Panel (c) displays the impulse responses when we included the share of employment in the commodity-producing sector in total employment as observable in the PVAR. The statistically significant increase of the share of employment in the commodity-producing sector is negligible (0.02%). Overall, the evidence portrayed in Fig. 2 suggests that employment in the commodity-producing sector is not a major source of variation in total employment after a commodity price shock. Other studies have also argued that changes in the supply of commodities in response to price changes are slow to occur. Unless sizeable excess capacity persists in the commodity-producing sector, the supply response is muted. Focussing on oil-producingNorway, Pieschacon (2012) includes oil production into a structural VAR. The estimated response of oil production after an oil price shock is small and insignificant. By contrast, the expansion in non-oil output is highly significant and about 5 to 8 times larger than the expansion in oil production depending on the horizon. Further back in history, Kindleberger (1973) documents for the interwar period that overall production of commodities did not contract significantly despite sharply declining commodity prices. To summarise, with commodity production being capital-intensive, employing only a small share of the domestic labour force, and being slow to respond to price shocks empirically, we deem it defensible to assume that commodity price shocks primarily transmit to the remainder of the economy through their impact on transfers and thus wealth. Nevertheless, we assess robustness of this modelling decision in Section 5.","We assess the empirical performance of the baseline model using a minimum distance strategy. In doing so, we distinguish between calibrated and estimated parameters. The calibrated parameters are listed in the top half of Table 2. The discount factor, β, implies a real interest rate of 4% per annum. The parameter σ governs the intertemporal elasticity of substitution and is set at 1.1. The share of capital in the production function, α, is 0.33, the depreciation rate, δ, is 2.5% per quarter. All these values are standard in the literature. The elasticity of substitution between home and foreign non- commodity goods, θ, is set at 2, which is within the relatively wide range of values commonly used in the literature. In 2013, the ratio of goods export to GDP in the OECD national accounts averaged at 20% across the economies in the sample. For 2013, the Harvard Atlas of Economic Complexity implies an average share of commodities in net exports for our sample of 85%. Estimates reported in IMF (2012) suggest a lower range of 30–40%, albeit taken over a sample dating back to 1960. We set the share of commodity exports in total exports at 65%. The share of non-commodity exports in GDP is set at 5%. With these target values in hand, the implied value for ν, the parameter measuring home- bias in consumption and investment, is 0.85. We estimate those parameters for which there is little or no direct empirical evidence: the bond holding cost parameter, ϕb, the vacancy adjustment cost parameter, ϕv, the investment adjustment cost parameter, ϕx, and the share of unemployment in the matching function, ζ. Given the choice of the replacement ratio, the dynamics of unemployment and vacancies depend on the relative consumption share of the unemployed to the employed agents, cssu/cssw, or equivalently the gap between the two, Φ = cssw − cssu. Consequently, we allow the data to determine the value of the consumption ratio. Performance of the model ~~~~~~~~~~~~~~~~~~~~~~~~ The red solid lines in Fig. 1 denote the fitted impulse responses of the baseline model. The corresponding parameter estimates and standard errors are reported in the bottom half of Table 2. Our model is able to closely replicate the dynamics of the 10-variable PVAR. For most variables, the model impulse responses capture the shape of the PVAR and lie within the error bands of its median response. Importantly, the model matches the initial impact and the approximate paths of unemployment and vacancies. To account for the slight ‘hump-shaped’ path of vacancies, the estimation yields a value for the vacancy adjustment cost parameter, ϕv, of 0.45. The accompanying standard errors are small. The model is able to reproduce the initial decline and the following gradual decrease in unemployment, albeit compared to the data, the path of unemployment is somewhat more persistent. The model cannot account for the initial decline or indeed the shape of the path of CPI-based real wages, and the baseline model somewhat underpredicts the volatility of individual hours. The commodity price rise transmits through a wealth effect to the economy as it induces a transfer to households. These transfers are used to boost consumption and domestic investment and to increase savings in the form of foreign bonds. Consumption in the model rises by around 0.2% and closely tracks the path of consumption in the PVAR. The theoretical model does rather less well in capturing the dynamics of chain-linked GDP. The data is far more volatile than the model suggests. Investment in the model closely tracks both the magnitude as well as the path of investment implied by the PVAR. The investment adjustment cost parameter, ϕx, takes the value of 0.5. The rise in consumption pushes up demand for both the domestic and the foreign non-commodity good, although the appreciation in the real exchange rate holds back demand for the former. In the short-run, the output expansion is driven by the increase in employment, while over time the gradual buildup of the capital stock also contributes to the modest rise in production of the domestic non- commodity good. The shape and the magnitude of the impulse responses depend on the real interest rate movement induced by the shock under the international financial market arrangements and on the household's decision on how to allocate the additional transfers towards savings in foreign bonds, consumption, and investment. In our framework the interest rate faced by households and firms is equal to the world interest rate adjusted by a small risk premium that declines with the country's net foreign asset holdings as in Schmitt-Grohé and Uribe (2003). The elasticity of this risk premium to the net foreign asset position of the home country, ϕb, is estimated at 0.25 indicating that international risk sharing is limited.11 In the data as in the model, the rise in domestic consumption occurs alongside a real exchange rate appreciation which further suggests that country- specificconsumption-risk from the shock cannot be effectively shared via relative price movements or trade in bonds. This feature differs from the transmission of a technology shock in the open economy. Cole and Obstfeld (1991) point out that movements in the terms of trade provide a powerful source of insurance against technology shocks independently of the financial market arrangements— with the exceptions of a low trade elasticity of substitution or permanent technology shocks stressed in Corsetti et al. (2008). An institutional arrangement that might insulate commodity-producing countries from the consequences of price fluctuations are sovereign wealth funds. The Norwegian Government Pension Fund is probably the most widely known example of such a fund. Focusing on the role of the fiscal policy regime, Pieschacon (2012) argues that the presence of this sovereign wealth fund helps in insulating the Norwegian economy from commodity-price related fluctuations.12 Results that we reported in the Appendix to the working paper version suggest a similar conclusion: after a commodity price increase Norway appears to run up its net foreign asset position by much more relative to the size of the shock than the three countries in our sample and thereby it strongly reduces the effect of the shock on consumption and investment in the near and medium term. Turning to parameters pertaining to search and matching in the labour market, we find that the share of unemployment in the matching function, ζ, is highly significant and comes out at 0.758 which is consistent with the literature reviewed in Mortensen and Nagypal (2007). Our estimates imply that the bargaining weight of households, ξ, assumes the value of 0.3 which suggests that firms have a rather higher weight in the wage bargaining process than workers. Finally, we find that unemployed members of a household enjoy about 54% of the consumption of employed agents. With the replacement ratio set at 40% of steady state wages, there is a modest, but far from complete, requirement to reduce consumption inequality between household members. We embed the discussion of the role of this parameter in influencing labour market dynamics and of the plausibility of its estimated value into the discussion of the transmission mechanism following next. Transmission mechanism ~~~~~~~~~~~~~~~~~~~~~~ The top two panels of Fig. 3 assess the relative importance of the productivity and the international channels by economic disturbances and time horizon. The top panel shows the movements of labour market tightness (with positive vacancy adjustment costs as estimated) in response to a commodity price increase along with the contributions of labour productivity and the real exchange rate to these movements. Taken together the productivity and the international channels as approximated by Eq. (12) account for the bulk of the movements in labour market tightness. As the commodity price increase triggers an immediate appreciation of the real exchange rate, but only a slow improvement in labour productivity, the international channel determines labour market tightness in the periods following directly the realisation of the shock. Following the logic in Bodenstein et al. (2011), the improvement in the commodity trade balance due to a commodity price increase must be accompanied by a worsening of the non-commodity trade balance over time as under rational expectations the absolute value of the net foreign asset position is bounded away from infinity. This worsening of the non-commodity trade balance is induced by a swift rise in the relative price of the domestically produced good which in turn translates into the observed appreciation of the real exchange rate. By contrast, labour productivity improves little on impact. Although firms are willing to increase investment in light of a higher marginal product of capital (due to the exchange rate appreciation) and a reduced real interest rate (due the improved net foreign asset position), the gradual expansion of the capital stock raises labour productivity only over time. Thus, as the international channel wanes in importance for the dynamics of labour market tightness, the productivity channel gains momentum. The ability of our model to deliver data-congruent movements of labour market tightness in response to a commodity price shock is tightly linked to the relatively low estimated degree of international risk sharing. Higher degrees of risk sharing, reduce the real exchange rate response and thus the contribution of the international channel to movements in labour market tightness. In the extreme case of setting the bond holding cost parameter ϕb very close to zero, the model converges to a simple permanent income model with an exogenously fixed real interest rate and barely any appreciation of the real exchange rate. As result, the response of labour market tightness is almost nil for this specification of the model. Under financial autarchy, the real exchange rate movement becomes too amplified and both vacancies and unemployment are too volatile relative to the data. The transmission of a commodity price shock differs significantly from the transmission of a technology shock. A positive neutral technology shock raises labour productivity persistently on impact while inducing a depreciation of the real exchange rate. The middle panel in Fig. 3 decomposes the response of labour market tightness after a neutral technology shock into the movements due to labour productivity and due to the real exchange rate. In this case, the productivity channel shapes the response of labour market tightness at all horizons. The depreciation of the real exchange rate is too small for the international channel to play a meaningful role for the dynamics of labour market tightness under the technology shock. Amplification of shocks and the labour market ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ This interplay between the parameters ru and Φ in amplifying the labour market response is not unique to the open economy context. Since Shimer (2005) argued that the textbook DMP model with Nash bargaining and CRRA preferences explains less than 10% of the volatility in U.S. unemployment and vacancies when fluctuations are driven by productivity shocks, the “correct value” of the replacement ratio has been the subject of lively discussion. Hagedorn and Manovskii (2008) and Hall (2008) argue that the flow value of unemployment ought to capture the value of leisure and home production in addition to the direct insurance payments to the unemployed. In a final assessment of the search and matching framework, we construct a prediction for labour market tightness from Eq. (12) using the estimated impulse responses of labour productivity and the real exchange rate to a commodity price shock. The bottom panel of Fig. 2 compares this approximation (dashed red line) to the empirical impulse response of labour market tightness which is depicted by the solid blue line.15 This approach predicts labour market tightness to respond similarly to its counterpart in the data after a commodity price shock. Judged by their respective peaks the magnitudes of the responses are similar, although the timing of the responses is shifted since the approximation in (12) does not account for time leads and lags.","Thus far we have abstracted from the use of resources in the extraction of commodities and the production of non-traded goods. Both these features introduce the possibility for labour and capital to reallocate into and out of the traded goods sector, the non- commodity sector in the baseline model. Resource-intensive commodity extraction ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As our discussion of Fig. 2 highlights, the share of employment in the commodity-producing sector is small for the economies in our sample and the bulk of the pickup in total employment after an increase in commodity prices occurs in the non-commodity sector. These findings also point to the existence of frictions in reallocating labour and capital from the production of goods to the production of commodities. In Fig. 4, the solid lines depict the impulse responses from an extended PVAR that includes data on employment in the non-commodity sector and in the commodity-producing sector. The lines labelled Model 2 show the fitted impulse responses of the model in which commodity extraction requires the use of labour and capital. The extended model is able to capture the empirical dynamics of employment both in the non-commodity sector as well as in the commodity-producing sector of the economy. The estimated impulse responses of the remaining variables are roughly identical to those in Fig. 1 despite differences in the data used to estimate the underlying PVARs. In response to an increase in commodity prices, employment in both sectors increases, but given the small relative size of the commodity-producing sector, most of the variation in employment is driven by the increase in employment in the non- commodity sector. To properly size the increase in the use of labour in the commodity- producing sector, the model requires sizeable adjustment costs in labour and capital. In line with the evidence discussed earlier, the implied expansion in commodity production is small. Finally, adding employment in the commodity-producing sector has only a minor effect on the previously estimated model parameters reported in Table 2. Non-traded goods ~~~~~~~~~~~~~~~~ When all domestically produced goods are fully traded, the wealth effect associated with the commodity price increase must lead to higher demand and production of the domestically produced traded good for the model to replicate the empirical responses of the labour market variables. Sufficient home bias in consumption and limited substitutability between domestic and foreign traded goods guarantee this outcome. When some domestically produced goods are not tradable, the appreciation of the terms of trade may cause the demand and production of domestically produced traded goods to decline. If the non-traded goods sector expands in response to the commodity price increase, the model can still replicate the empirical responses. Whether the economy experiences this Dutch disease effect analysed in Corden and Neary (1982) depends on the substitutability between the domestically produced and the foreign traded goods and the substitutability between traded and non-traded goods. The dashed lines in Fig. 4, labelled Model 3, shows the fitted impulse responses of a model with both resource-intensive commodity extraction and a non- traded goods sector. Overall, Model 3 provides a similar fit to the data as Model 2. Noteworthy differences in the estimated parameters reported in Table 2 are the smaller investment adjustment cost parameter (some investment goods are now non-traded, making investment less volatile) as well as the larger bargaining weight of the households. While the estimate for the parameter governing the costs of bond holdings is reduced for Model 3, risk sharing is still rather limited. In matching the impulse response functions of the model with non-traded goods to the data, we set the elasticity of substitution between the two traded goods at 2 and the elasticity of substitution between traded and non-traded goods at 0.74 as in Mendoza (1991). For this specification, the economy experiences the Dutch disease effect after an increase in the commodity price. As shown in Fig. 5 employment in the traded goods sector drops sharply, as does production (not shown). By contrast, the increase in employment in the non-traded goods sector more than compensates for this drop, so that, in accordance with the data, total employment expands (with a small contribution from the commodity-producing sector). In theory, the Dutch disease effect can be tempered, if not eliminated, by either lowering the elasticity of substitution between the two traded goods or by raising the elasticity of substitution between traded and non-traded goods. When commodity prices increase, the price of both non-traded and domestically produced traded goods goes up— in the case of integrated factor markets and identical technologies, prices go up by the same amount. Since the price of the foreign good is determined in the world market and does not change in this experiment, the terms of trade depreciate and the price of the traded good (an aggregate of the domestically produced and the foreign trade good) rises by less than the price of the non-traded good. If the traded and the non-traded goods are sufficiently substitutable, the demand for both the foreign and the domestically produced traded good rises (due to home bias and limited substitutability between traded goods). Fig. 5 shows that employment in the domestic traded goods sector expands when the elasticity of substitution between traded and non-traded goods is raised to 3 while maintaining the elasticity of substitution between traded goods at 2. When maintaining the elasticity of substitution between traded and non-traded goods at 0.74, the demand for both traded goods can also rise under a low elasticity of substitution between the two traded goods. Fig. 5 shows that employment in the domestic traded goods sector expands when the elasticity of substitution between traded goods is lowered from 2 down to 0.75. We estimated the model under alternative parameter choices for the substitution elasticities between goods. Overall, a good fit of the model to the empirical impulse responses requires the elasticity of substitution between traded goods to be above unity. Assuming that the elasticity of substitution between traded and non-traded goods is low (around 0.74), it appears therefore likely, that the countries in our sample suffer from the Dutch disease effect after an increase in commodity prices. Absent high quality data on traded and non- traded goods at quarterly frequency, this prediction of the model is rather strong.","Using a structural PVAR and a standard small open economy model augmented by commodity net exports and search and matching frictions in the labour market, we find: Even advanced commodity exporting countries are restricted in their ability to smooth out the effects of commodity price fluctuations through financial markets and relative price movements. To match the empirical dynamics of the trade balance and the real exchange rate, the theoretical model requires a low degree of international risk sharing. Matching the empirical movements in the real exchange rate is also essential to obtain data-congruent movements in unemployment and vacancies in the model. We show this using a novel decomposition of labour market tightness into a labour productivity channel and an international channel that captures movements in the real exchange rate. This decomposition suggests that the open economy dimension is important in understanding the transmission of shocks to aggregate demand components to labour market tightness. In addition to the international channel, the model also requires a preference specification that drives a wedge between the consumption of the employed and the unemployed household members for sufficient amplification of the commodity price shock to the labour market. These findings are robust to augmenting the model with resource-intensive commodity extraction and non-traded goods. In the presence of non-traded goods, the economy is highly likely to experience the Dutch disease effect after a commodity price shock."],["We study the effects of the recent economic crisis on firms' bidding behavior and markups in sealed bid auctions. Using data from Austrian construction procurements, we estimate bidders' construction costs within a private value auction model. We find that markups of all bids submitted decrease by 1.5 percentage points in the recent economic crisis, markups of winning bids decrease by 3.3 percentage points. We also find that without the government stimulus package this decrease would have been larger. These two pieces of evidence point to pro-cyclical markups. --------------------------------------------------------------------------------","The recent financial crisis provides a unique testing ground for analyzing causal economic relations. Although some economists warned of the bubbly nature of asset prices (e.g. house prices, see Shiller, 2005), the financial crisis and the ensuing economic crisis were unforeseen.1 We use the current economic crisis to shed light on how markups move in response to an exogenous shock in demand. We investigate firms׳ competitive behavior in procurement auctions before and during the crisis; and build a bridge between micro-level behavior and the effects of macro-level shocks. The recent economic crisis allows testing for (at least) two effects. First, it enables us to estimate the effects of the recent economic shock on fundamental structural characteristics of the economy. In particular, we test for the effects of the shock on the process and intensity of competition, i.e., pricing behavior and markups. Second, the effects of fiscal policy can be tested in an unprecedented manner. In normal times, the effects and determinants of fiscal policy are difficult to ascertain due to their endogenous nature. Activist fiscal policies in the crisis are clear cut and mostly targeted to specific industries like the construction sector. Thus, we can test for the effects of fiscal policy by constructing a proper counterfactual allowing us identification. We use data from before the economic crisis to predict what would have happened but for the state stimulus intervention. Our theoretical reasoning and strategy of identification are as follows. In “normal” times (in the years before the recent economic crisis) firms operate at their equilibrium capacity utilization rate. Sometimes they are unconstrained, but sometimes they hit their capacity constraint. Using data from highway procurement auctions, Jofre-Bonet and Pesendorfer (2003), for example, estimate that firms are capacity constrained in 32% of the contracts during the “normal” economic time period 1996–1999, and find that bids are 18% higher if all bidders are constrained. This implies that economic rents are higher, if firms are capacity constrained. The negative demand shock resulting from the economic crisis relaxes capacity constraints: bidders have idle capacity. Given idle capacity, they participate more often in the remaining auctions and bid more aggressively since their costs are lower. In a dynamic model, if firms are capacity constrained, they may also price into their bids the lost option value of winning today versus winning later, i.e. the (additional) cost of winning today consists also of the loss of future discounted profits due to limited capacity. Our results are based on an auction model fitted to detailed and comprehensive data from procurement auctions in the Austrian construction sector in the period 2006 to 2009. Using the start of the downturn as a quasi-experiment, the model enables us to show how firms react to the massive negative demand shock. The economic crisis has led, however, governments in many countries to implement stimulus packages and counteract these negative shocks. This has fueled interest in and discussion of the questions if, how and which government measures are effective. The Austrian government also implemented a stimulus package in response to the economic crisis. The main target was the construction sector. Applying the same logic as above, the government intervention should — ceteris paribus — increase markups. The focus of stimulus measures on the construction industry is not idiosyncratic for Austria. Many countries implemented stimulus measures, of which the construction sector plays a large part. In Austria, the share of the stimulus packages covered by construction spending was about 22% in 2009. In Germany, about 17% of the stimulus in 2009 and 2010 were directed to construction. In the US, the share was about 13% in 2009. Thus, we believe that our results are relevant for fiscal policy in general and are not confined only to Austria. For the empirical implementation, we follow and simplify Athey et al. (2011) and use a parametric version of Guerre et al. (2000) to recover the distribution of bidder costs from the observed bids. We assume that the demand shock exogenously changed bidders׳ participation.2 The distribution of bidders׳ costs is estimated based on the assumption that in equilibrium the estimates of the distribution function summarize bidders׳ beliefs and can be used to infer bidders׳ costs based on the first-order conditions of optimal auctions. Estimation of the distribution of bids is based on auction characteristics, firm characteristics as well as the level of competition to which firms are assumed to respond optimally. We utilize a detailed data set that includes a rich set of variables. Our data set covers nearly the population of all public procurement auctions in the construction sector in Austria during the years 2006 to 2009. It includes auction specific variables such as the engineers׳ cost estimates of one firm for each project and the identities and bids of participating bidders. By matching to a firm-level database and combining it with data on travel distances from the firm׳s address to the location of the project, several bidder-specific variables are added. We are thus able to control for bidder heterogeneity in the econometric model. We show that the crisis indeed had the expected effects on competition: the negative demand shock led to more bidders in the remaining auctions and these bid more aggressively. We find a significant decrease of the average markup in the crisis period of about 1.5 percentage points relative to a pre-crisis mean markup of 12%. The winning bids are 3.3 percentage points lower relative to 22.9% pre-crisis. Extending our analysis to the dynamic model of Jofre- Bonet and Pesendorfer (2003) shows that the main crisis effect is robust to the incorporation of the option value effect of winning. Markups increase by about 2 percentage points overall in the dynamic model. The crisis drop is about 2 percentage points for all bids and about 4 percentage points for winning bids. We attribute the decrease in markups mainly to two characteristics that are affected simultaneously by the negative demand shock: lower backlogs of firms and an increase in the number of bidders. Because of the unforeseeable nature of the recent financial crisis, firms did not reduce their capacities to the lowered demand level. Exit did not immediately follow. Roughly the same number of firms bid for a reduced number of projects and consequently more firms bid for each project. More competition and reduced marginal costs lead to lower markups and lower prices. A counterfactual analysis provides an estimate of the effect of the stimulus on markups. In this counterfactual, we reduce the backlogs of bidders by the amount the government spent on stimulus in the construction sector. Our estimate indicates that the drop in markups would have been about a third higher — increasing from 1.5 to 2 percentage points — without government stimulus. While the macroeconomic literature compares markups over the business cycle across industries,3 we follow the microeconomic literature and estimate bidders׳ cost distribution from submitted bids in construction procurement auctions.4 By focusing on one industry and imposing the structure of a game-theoretic model we are able to back out firm specific cost and do not have to rely on data averaged over firms. We have to assume, however, that bidders behave according to the theoretical model.5 In the literature, construction procurement auctions have been extensively studied. Jofre-Bonet and Pesendorfer (2003) show non-parametric identification in repeated auction games and using a parametric model for the estimations, they find evidence of capacity constraints in Californian highway procurement auctions. Krasnokutskaya (2011) proposes a nonparametric estimation method to recover the distribution of bidders׳ private information when unobserved auction heterogeneity is present and finds that private information is estimated to account for 24% of the variation in bidders׳ costs in Michigan highway procurement auctions. Balat (2012) looks at road construction in California and investigates the effect of the American Recovery and Reinvestment Act on equilibrium prices paid by the U.S. government. He extends the models by Jofre-Bonet and Pesendorfer (2003) and Krasnokutskaya (2011) and allows for unobserved project heterogeneity and endogenous participation. His results show that prices rather than quantities increase as a consequence of the implemented stimulus. Groeger (2014) studies participation and bidding decisions in repeated highway procurement auctions and the estimated parameters suggest that participation synergies exist. Further studies accounting for endogenous participation are Krasnokutskaya and Seim (2011) and Athey et al. (2011, 2013). The organization of the paper is as follows. In the next section, we describe the construction sector in Austria and provide insights into the organization of procurement auctions. Section 3 presents the theoretical background and the econometric model of the distribution of bidders׳ construction costs. We provide a discussion of the specific timing of the crisis as it hit the Austrian construction sector and describe the likely consequences of the crisis on competition. The section also contains details on the data as well as summary statistics. Section 4 presents and discusses the estimation results. Based on the estimates, we discuss the effects of the recent economic crisis and the associated fiscal stimulus on firms׳ costs and markups. In addition, we report robustness checks as well as counterfactual effects of the model, together with more detailed results of markups in the stimulus and non-stimulus sector and for specific stimulus projects. The relation to other findings and competing models as well as possible policy conclusions are discussed in Section 5. In Section 6, we conclude.","This section provides information on Austria׳s construction sector and how the Austrian government implemented the stimulus packages to combat the economic crisis following the breakdown of Lehman Brothers. A main part of the stimulus package concerned the construction sector and thus affected procurement auctions. Finally, we describe the organizational background of these auctions. Construction sector and stimulus packages in Austria ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In Austria, the construction sector accounts for 7% of GDP on average between 2006 and 2009 (in Germany this share is 4.2%, in the USA 4.4%; OECD STAN data set). As can be seen from Fig. 1, total construction value added steadily increased from around 10 billion euros in 1999 to 2002 to 16.2 billion euros in 2008 in Austria in nominal terms. In 2009, when nearly all countries went into recession due to the financial crisis, total construction value added fell to 15.5 billion euros, which is a drop of 5% in real terms (the inflation rate was 0.5%) compared to a drop in real GDP of 3.8%. Building (corresponding to 2-digit-SIC code 15) and heavy construction (SIC-2: 16) are the two main subsectors. While building construction remained nearly flat in real terms, heavy construction value added fell by around 10% in real terms. Like other countries, the Austrian government responded to the global economic crisis and did this with similar measures and at similar dates. In October 2008, the Financial Market Stability Act and the Interbank Market Support Act became effective to support the interbank market, allow the government to grant guarantees, assume bank liabilities and acquire bank equity. Also in October 2008, the parliament voted for the first of two stimulus packages. In December 2008, the council of ministers voted for the second stimulus package. The major parts of the stimulus packages consisted of credit and financing for private companies, a bonus depreciation plan, additional funds for R&D, subsidies for thermal rehabilitation of buildings, a car scrappage program (“cash-for-clunkers”) and increased public expenditures on infrastructure. These infrastructure expenditures were made by advancing investment on public infrastructure, originally scheduled for years later than 2009 and 2010. In 2009 — the year when our data end and the first year of stimulus — the stimulus directed at the construction sector consisted of 363.3 million euros. It was invested into roads, railways and (public) buildings. Within these categories, it was spread over all sorts of construction projects: maintenance, rehabilitation, rebuilding and new building of railway stations, roads, schools, universities, judicial buildings and so on. Organization of procurement auctions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Austria׳s public authorities (federal and regional government, social security institutions, and the like) are subject to the Federal Public Procurement Law (PPL, “Bundesvergabegesetz”). Private companies have to follow these rules only if (1) they are active in a “sectorial activity” (provision of water, mail services, energy, or traffic), and (2) their activity is regulated (e.g. entry regulations). In principle, there are no lower limits for the applicability of the PPL. Upper limits are in effect, but only to enforce EU-wide announcement rules for larger projects (in construction, currently 4.845 million euros). The main purpose of the PPL is to stimulate competition and safeguard equal treatment of all potential bidders. The law stipulates some restrictions on bidders, but they do not, in our judgment, affect competition, but are intended to forestall opportunistic behavior (e.g. the requirement of a 5% bid bond). Except for very small projects the public authority must choose between an “open procedure” or a “restricted procedure with publication of a contract notice”. Such projects have to be made public and documents must be freely accessible. In our sample, 85% of the auctions are “open”, and only 2% are of the “restricted” type. The remaining 13% were not publicly announced. Instead, selected qualified companies were invited to submit bids. All contracts are awarded by first-price sealed-bid auctions. There is no explicit reserve price – the seller has the right to withdraw the auction when the winning bid is “contra bonos mores” (offending against good morals). In court, the seller would proof this by providing a cost estimate that is based on standard commercial and professional principles and has to show that the winning bid is far higher than this estimate. Each project is described in a procurement catalogue. The seller defines the components and quantities needed for the construction project. The bidder provides a unit price for each component. The final price in the auction is then the sum of prices for the various components multiplied with the quantity. Bidders who want to participate in an auction have to proof their commercial and professional abilities before being admitted to the bidding process. On the letting day, all bids are unsealed, ranked, and the low bidder wins the auction. After the lowest bid has been identified all bidders are informed. If no bidder objects the decision within 10 days, it is final.","This section describes the theoretical model, the econometric model of the distribution of bidders׳ construction costs and the hypotheses for our empirical analysis. We also describe our data and give summary statistics. Our theoretical reasoning and strategy of identification are based on the assumption that in “normal” times, i.e., in the years before the Great Recession, firms operate at their equilibrium capacity utilization rate. Once firms hit their capacity constraints, their costs increase. We then expect economic rents to be higher. The negative demand shock due to the economic crisis implies that the capacity constraints are relaxed: bidders have idle capacity in the crisis. Given idle capacity they participate more often in the remaining auctions and bid more aggressively since their costs are lower. To estimate the distribution of bidders׳ costs, we implement a static first-price auction with an exogenously given number of bidders as our main specification. As robustness checks, we also implement a dynamic first-price auction that accounts for the strategic effect of capacity constraints following Jofre-Bonet and Pesendorfer (2003), and a model with endogenous entry following Athey et al. (2011).6 As our main econometric model is however based on the static model, we stick to the static theoretical model in the following and provide the respective implementations of the other models in Appendix A. Estimation of bidders׳ costs ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To illustrate the econometric model and the structural estimation of the distribution of bidders׳ construction costs, we use sealed-bid data to estimate the parameters of the theoretical model as a function of auction and bidder characteristics making use of the approach developed by Guerre et al. (2000). They suggest to estimate the distribution of bids in a first step and to recover the distribution of bidders׳ costs in a second step by using the first order-condition for optimal bidding behavior. With the estimates of the distribution of bidders׳ costs at hand, we then calculate markups. Timing of the crisis and empirical hypotheses ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In our empirical analysis we want to exploit the economic crisis as a quasi-natural experiment after which firms in the construction sector start going into public procurement auctions more often. It is not obvious when the demand changes were perceived as signs of an economic crisis and consequently when firms changed their behavior. Nevertheless, we want to identify a rough date when to expect changes borne by the factors outlined above. Fig. 2 puts developments in the construction sector in a longer time perspective. Total construction (private and public contracts) displays an upward trend both in the stock of contracts and new orders — new orders are the gross inflow of new contracts that are accepted by firms in a given month — up to about 2007. Stocks of contracts reach a plateau a few months earlier than new orders. Not before September 2009 do the stock values fall short of the level seen in January 2008. Subtracting eleven months leads to the first possible impact of the trailing moving average. October 2008 then is a potential date when we could expect changes in the bidding behavior of firms if stock values of contracts are relevant for the bidders. New orders provide early information on the output of the construction sector in the months that follow. Construction firms can anticipate their backlog in the near future and may adjust bidding behavior. A relatively consistent downward movement of new orders starts in March 2009, which, after subtracting eleven months, gives April 2008 as a second relevant date. In the crisis demand shrinks, capacity constraints are relaxed, and (roughly) the same number of bidders bid in fewer auctions, implying an increase in the number of bidders per auction. Thus, due to this competition effect, bidders bid more closely to their true costs in the crisis. Fig. 3 shows a striking negative relationship between the number of bidders per auction and the new orders series in the construction sector (both public and private contracts).19 While new orders in the construction sector sharply decrease due to the crisis, the number of bidders per auction sharply increases. New orders in public procurement (also in Fig. 3), however, are larger in the last few months of 2009. Accordingly, our backlog measure — after dropping sharply in 2008 — begins to rise again in early 2009.20 We attribute this to the growing inflow of stimulus projects. To provide further evidence for our argumentation that the crisis affected the number of bidders in public procurement auctions and bidders׳ capacity utilization measured by their backlog, we show descriptive regressions in Table 2. We regress each variable on a dummy variable that is zero before and one in the crisis — starting in October 2008 — and a time trend to control for general movements in the time series. The results depicted in Panel A of Table 2 confirm our impressions from Fig. 3. We observe that the number of bidders in an auction increases in reaction to the crisis. This effect is significantly different from zero and about one and a third bidder on average. In columns (2)–(5), we present the robustness of our estimate to placebo treatments. One might be concerned that the increase in the number of bidders picks up some additional unspecified effects over time. Similar to Black et al. (2008), we introduce a placebo treatment and add hypothetical crisis dates to the specification. We use four additional dates for the crisis: two months and six months before and in the crisis. The estimated coefficient for the change in the number of bidders does not change significantly. We repeat this exercise for the backlog. We find that the estimated effect of the crisis on the backlog is negative as expected and significantly different from zero. Again we run placebo experiments and find that the estimated coefficient does not vary much over the specifications shown in columns (2)–(5) in Panel B of Table 2. In the presence of a negative demand shock, the incentive to shade bids above one׳s cost decreases. A bid is equal to the expected cost of the competitors conditional on the competitors׳ costs being less than the bid. If there is more competition in the sense that there are more actual and capacity unconstrained bidders in an auction, the markup goes down, (1) since the expected conditional cost of competitors decreases, so bids are closer to costs, and (2) capacity unconstrained bidders bid more aggressively since their costs are lower. This implies that bidders׳ expected rents go down in equilibrium and gives the first hypothesis. Hypothesis 1: The markup goes down in an economic crisis. As argued, our structural approach allows us to measure the impact of activist fiscal policy. The state intervened drastically during the crisis with the two stimulus packages, and particularly so in the construction sector. Economically, targeted fiscal stimulus shifts the demand curve in the construction sector at least partially back to where it would have been without the crisis. This gives our second hypothesis. Hypothesis 2: The markup would have decreased by more had the state not intervened in the construction sector. If validated, both hypotheses imply pro-cyclical markups. The main underlying assumption of our two hypotheses is that an increase in the number of bidders yields lower procurement prices and lower bidder revenues. This is a standard competition argument, but need not always be true. In a model with common values,21 informational asymmetries such as the “winners׳ curse” may offset the above argument. We argue that the independent private value model is valid in our context. As described in Section 2.2, construction projects consist not only of one (final) price, but are actually the sum of prices for various components expressed in fixed quantities that are known from the procurement catalogue. Firms put in the prices they are willing to offer for the various components. As the components of the project are fixed,22 firms then differ only by their private cost. However, even in an independent private values model as we apply here, bidder asymmetries23 or costly bidding24 may also lead to the effect that more bidders could lead to reduced price competition. In our case, the empirical model allows for bidder asymmetries as we control for firm specific variables such as the distance between firm location and construction site as well as firm size, and estimate bidder specific distributions. In addition, we run a robustness check modelling participation. Finally, we do not consider bidders to be risk averse, but risk neutral. If, however, bidders׳ attitude towards risk would change in the crisis, we may attribute some of the change in markups incorrectly to competition as risk aversion induces more aggressive bidding than risk neutrality (Maskin and Riley, 1984).25 While estimating a model with risk aversion would go beyond this paper, we can at least partially account for changes in risk aversion over time by splitting the sample and estimating two different cost functions as we do as one of our robustness checks.26 Data and summary statistics ~~~~~~~~~~~~~~~~~~~~~~~~~~~ One large Austrian industrial construction company provided our main data, containing bids, firm names and auction characteristics. This data set covers all auctions in both building and heavy construction in the period January 2006–December 2009, where this company took part either as the parent company itself or as one of its subsidiaries. According to the company, the database covers more than 80% of all auctions which must be conducted according to the Public Procurement Law. Thus our sample covers nearly the population of all public procurement auctions in Austria during the four years 2006–2009. Within this period, our database reflects on average 14% of Austria׳s total construction sector. In 2009, this share is at 16%, which mirrors the increased activity of the state through the stimulus. This increase is more visible in heavy construction, where our data set represents 26% of total heavy construction in 2009, up from an average of 21%. Our data set covers a relatively stable share of 10% of building construction. Of these firms, 1342 or 81%, were matched to Amadeus. Thus, 96% of the total 29,850 bids come from matched firms. The population of Austrian construction firms, as reported by the Austrian federal agency, Statistics Austria, consists of 4796 firms on average for 2006–2009 (building: 3958, heavy: 838). Therefore, about one-third of all construction firms submitted bids in the procurement auctions. Table 4 gives detailed yearly summary statistics on firms, auctions, bids and competitive structure. Bidder characteristics: We cover many relatively small firms with an average of around 150 employees (median: 54 employees) and 20 million euros in total assets (median: 4 million). Auction characteristics: Splitting the sample into the two main construction sectors reveals that two-thirds belong to building and the rest to heavy construction. On average, the technical estimate for the projects׳ cost is 2.4 million euros (at 2006 constant prices) and the average winning bid is 4% below that. Transportation costs play an important role in the construction sector. We measure these costs by the distance of a firm to a construction site. Based on the construction sites׳ and bidders׳ postal codes, we used Microsoft׳s Bing Maps to calculate the driving distances for all bidders to the project sites corresponding to the auctions. Postal codes for virtually all bids are available in our main data. 4-digit postal codes give very detailed information on location in Austria and exact street addresses can be neglected without loss of relevant information for our purpose.29 Besides driving distance, we have also retrieved travel times. While in singular cases the two may differ substantially, the correlation is very high at 0.973 in our sample. Not surprisingly, all calculations are robust to the choice between them or the inclusion of both. Mean distances equal 123 km in 2006 and 2007, reach its highest value of 139 km in 2008 and fall slightly to 135 in 2009. Thus, firms are willing to bid in projects farer away and thus to drive longer distances to the project sites in the crisis.30 Capacities: Backlog used by a firm at a point in time is calculated as in Jofre-Bonet and Pesendorfer (2003). Every project is added to the backlog when a firm wins it. Projects are linearly worked off over their construction period, which releases capacity. Every firm׳s backlog obtained is standardized by subtracting its mean and dividing by its standard deviation. We have also experimented with alternative definitions of backlog but they do not change our findings substantively. The variable captures all auctions of the company providing us with the data but does not carry information on backlogs of auctions where this firm did not participate. Although this company participates in most auctions, potential bias could be introduced. To relieve the variable from potential bias, additional data about projects won in the past come from the “Austrian Register of Tenderers”. The register gathers all public procurement tender offers and information about the winners and often the contract value of past concluded contracts. We selected those projects where at least one recorded subclassification belongs to the construction sector (Common Procurement Vocabulary 2-digit code equal to 45). By matching the winning bidder and thus the winning bid, where available, to the main data set, unbiased procurement backlog data are available for all companies. A backlog variable needs some precursory data to account for projects already in the books of the firms when our main data set starts. The additional information from these external data contains projects before 2006. We go back until 2005, one year before our estimation sample starts, in our matching. Of the 5235 winning bids in the Register, 42% come from firms that also bid in our main data. These winning bids correspond to 72.3% of all project values won in the Register. Keeping in mind that some of the projects from the Register that we retrieved are primarily other-than-construction projects and have often only one subclassification from the construction sector, this high percentage of procurement values from matched firms gives further indication that our main data is representative. To proxy firms׳ outside option, we use the inflow of new orders into the construction sector, i.e., a macro variable, in our regressions. Herewith, we can account — at least to some extent — for contracts that have not been procured. Stimulus sector and stimulus projects: Stimulus in the construction sector took place predominantly in three incorporated but government owned institutions for managing public buildings (Bundesimmobiliengesellschaft), public highways (Autobahnen- und Schnellstrassen- Finanzierungs-Aktiengesellschaft) and public railways (Österreichische Bundesbahnen). From public information we could identify which specific projects in the sample were designated by the government for the stimulus package. 17 of the 561 projects in the estimation sample in 2009 are part of the stimulus package, representing 3.8% of the total project value of 2009. These specific projects were scheduled after the year 2009 originally, but were actively brought forward by the government to counteract the crisis. We thus differentiate between two types of stimulus by the government. First, the government did not abandon projects as the private sector did. We call this sector “stimulus sector”, and it includes essentially the projects of the aforementioned three institutions (and it includes also the 17 specific stimulus projects). Second, we separately look at these 17 additional projects and call them “stimulus projects”. Relation between engineer estimate and bids ranked first and second: Table 5 shows the absolute values of the engineer estimate and the winning bid in columns (1) and (2), tabulated by the number of bidders. As the number of bidders increases, the relative difference between winning bid and engineer estimate falls (column (3)). Column (4) reports the number of auctions for each number of bids submitted in the estimation sample. In the lower part of Table 5, “money left on the table”, the relative difference of second-ranked to winning bid, which is a measure of informational asymmetries, is tabulated against the number of bidders. This relative difference is on average 7.7% and it narrows as the number of bidders increases. These statistics are consistent with the findings of Jofre-Bonet and Pesendorfer (2003), Balat (2012) and Krasnokutskaya (2011). Columns (6)–(8) divide the auctions into projects not in the stimulus sector, in the stimulus sector and the 17 specific stimulus projects we could identify. “Money left on the table” appears to be larger for stimulus related projects.","In this section, we describe the estimation results. Based on these results, we calculate bidders׳ markups. We exploit the recent economic crisis as a quasi-experiment and test whether firms׳ pricing power is affected by the recent economic crisis and the stimulus. We then check the results for robustness. Finally, we run a counterfactual analysis to validate the results from the quasi-experiment and test our empirical hypotheses. Estimation results ~~~~~~~~~~~~~~~~~~ Columns (1) and (2) of Table 6 show the estimates for the scale parameter λ, columns (3) and (4) for the shape parameter ρ. Backlog (i.e. own backlog) and backlog sum (i.e. the backlog of the other actual bidders) measures capacity constraints. As expected, both backlog measures display positive coefficients. Firms bid higher if their own backlog is higher since they have higher opportunity costs; and firms strategically bid higher when they know that their competitors have higher backlogs, and therefore higher costs. The effect of the log of the number of bids is negative — consistent with the notion that more bidders increase competitive intensity. Note that the log is used to account for the diminishing influence of an additional bid as the number of bidders grows. The engineer estimate is the most important determinant of the bid. The log of the number of employees, our firm specific size measure, has a positive influence on the bid level. KM measures the driving distance between firm and project site in kilometers. While insignificant, two other variables derived from the distance data are significant: KM sum, defined as the sum of competitors׳ distances, and KM average, the average distance of the bidders in the auction. We interpret these as strategic variables in a limit pricing context: the further competitors are away from the project site, the higher their costs. When competitors are further away, a firm can submit higher bids. Besides that, some control variables are included: same postal code of firm and project site, same district and same state, open format auctions, heavy construction auctions, and bidder is a general contractor. Fig. 4 compares the kernel densities of actual to predicted bids. The prediction matches the actual density closely. To assess our two main effects more closely, we show the effect of the backlog and the number of firms on bids in Fig. 5. For the illustration we chose the median project. The ranges of backlog and the number of bidders are chosen such that 99% of all bids are covered. Both variables have the expected sign in the estimation of the scale parameter. In general, the shape parameter also influences the expected value and the total effect of a variable in the Weibull distribution is not obvious from the estimates. For our results, the partial effect of backlog or number of bidders turns out as hypothesized, too: more bidders in the auction decrease the bid, a higher backlog increases the bid. Quasi-experiment ~~~~~~~~~~~~~~~~ From our structural estimates we proceed with the calculation of the markup for each bid.31 With the calculated markups, we test the first hypothesis on the effects of the economic crisis on markups, i.e., that markups go down in an economic crisis. To obtain our results, we are interested in the mean difference of markups between before the crisis started and afterwards. As has been shown in Fig. 2, new orders and stock of contracts start downward developments in April 2008 and October 2008. If we split the sample in April 2008, the difference between before April 2008 and after April 2008 in mean markups is −0.6 percentage points; if we split the sample in October 2008, the difference reaches significant −1.5 percentage points. The shaded area in Fig. 6 displays these differences on the black line at the left border of the shaded area (April 2008) and right border of the shaded area (October 2008). We further show those values when we split the sample at any arbitrary point in time between April 2006 and October 2009. Note that the sample size is constant at n=14,845, but the subsample of the first part — the pre-crisis period — shrinks when we move the crisis start date to earlier months and the subsample of the second part — the crisis period — grows accordingly. The black line gives the difference of mean markups before and after a specific month. The dotted lines above and below the black line show the upper and lower limit of the 95% confidence interval and indicate whether the difference is significant. Reassuringly, the graph of markup differences displays a downward trend indicating that the crisis truly reduced markups. The actual prices paid are determined by the winning bids. Fig. 7 shows the mean markup differences when only winning bids remain in the calculations. The point estimate of the mean difference between before April 2008 and after April 2008 is −0.9 percentage points and therefore larger than for all bids at the same date, but not significant. The difference becomes −3.3 percentage points if October 2008 is chosen as the crisis starting date, and this number is both economically large and statistically significant. For the (eventually relevant) winning bids, therefore, the reduction of markups due to the crisis is more pronounced. Stimulus sector and stimulus projects: In Table 7, we distinguish between all projects, projects in the stimulus sector and stimulus projects. Distinguishing the markups for these subgroups is of descriptive interest. We observe that markups in the stimulus sector are higher than average markups. This is true before and in the crisis. We also observe that the drop in markups is more pronounced in the stimulus sector than on average. Markups of stimulus projects are larger than markups in the stimulus sector, when we consider all bids. Summarizing, from our structural estimates we conclude that markups decrease in the economic crisis and confirm Hypothesis 1. Threats to identification ~~~~~~~~~~~~~~~~~~~~~~~~~ The identification of our results strongly depends on the assumption that the crisis is a quasi-experiment that exogenously decreased the demand for construction projects. We argue that there was a sudden downturn in the construction sector caused by the crisis and essentially the same number of firms is competing for fewer projects. As a consequence we observe a higher number of bidders per project, more competition in the procurement auctions and, therefore, lower equilibrium prices and lower markups. The advantage of our empirical setup is the clear time-line and its simplicity. The crisis is an exogenous demand shock, and we restrict our attention to a simple static framework. There are, though, threats to our identification. One could think that not only demand was affected, but also bidders׳ cost. In the crisis, credit access and interest rates as well as input prices may have changed. One other source of concern might be the effect of the engineer estimate. Since one firm provided us with the data, it could be biased towards any idiosyncrasy of that firm and yield biased estimates of bidders׳ costs. It is also conceivable that the largest firms behave differently than the fringe firms. In addition, we do not account for unobserved auction specific heterogeneity and endogenous entry. Finally, as we do not consider the dynamic strategic effect of backlogs, we may misinterpret our results due to an omitted variable bias. In order to confront these problems we provide additional pieces of evidence that should lend credibility to the interpretation of our results as market responses to the exogenous demand shock. First, we investigate whether bidders׳ cost distributions have changed in the crisis. We split the sample into the period before and during the crisis and estimate bidders׳ cost distributions separately for these two periods. Second, we address the engineer estimate and look at several modifications of our sample. Third, we include fixed effects for the seven largest firms.32 Fourth, we allow for unobserved heterogeneity, endogenous entry or dynamic considerations. When we model bidders׳ costs, we assume that the set of auction characteristics is known to the econometrician and to the bidders. To relax this assumption, we follow Krasnokutskaya (2011) and allow for unobserved heterogeneity. Neglecting entry cost or the dynamic strategic effect may bias our results. Thus, we estimate a model with endogenous entry similarly to Athey et al. (2011) and a dynamic model following Jofre-Bonet and Pesendorfer (2003).33 Our first robustness check allows the cost distribution to differ across periods. Panel A of Table 8 presents the estimated coefficients of the cost distribution before the crisis and Panel B during the crisis. We observe that, for example, the estimated coefficients for the number of bidders and for the backlog differ. The effect of the number of bidders becomes stronger in the crisis, whereas the effect of the backlog becomes weaker. This reinforces our main argument how the crisis affected this market: competition increases and capacity constraints relax. We then calculate the markups based on these cost distributions and present them in Table 9. Column 1 contains the results from the base model for comparison. The second column, R1, shows the estimated markups when we allow the cost distribution to differ over the two periods. Our main result of lower markups in the crisis holds. Interestingly, the difference is slightly more pronounced on average. This is true for all bids and for the winning bids. The next three robustness checks in Table 9 address the engineer estimate in one way or the other.34 For column R2, the bids of the firm supplying the engineer estimate are omitted from the estimation of the determinants of bids. Results are robust to this modification. In R3, the engineer estimate variable is dropped altogether, nearly doubling the sample size. Although the crisis drop in markups is still present, the magnitude of markups is much larger than in the base estimate. Clearly, one has to control for auction size. For the next check, R4, also auctions with very poor engineer estimates, in terms of large deviations from the bids that were submitted eventually, were included. Consequently, the sample size is a bit larger than that of the main results in Table 6. While average markups increase, the pre-crisis–crisis pattern remains. R5 shows the results when fixed effects for the largest seven firms in terms of total number of bids submitted in the sample are included. The fixed effects permit a somewhat different behavior by the largest firms, similar to the split into main bidders and fringe firms found in Jofre-Bonet and Pesendorfer (2003) or Athey et al. (2011). Results are robust to the inclusion of firm fixed effects.35 In the next column, R6, the model switches to a Weibull-Gamma-mixture distribution to permit unobserved auction heterogeneity similar to Krasnokutskaya and Seim (2011) and also Athey et al. (2011). Average markups go down, but again our main result of decreased markups in the crisis holds up. In column R7, we account for entry. The results are very close to the markups calculated in the base model. Winning bids׳ markups are slightly larger after accounting for entry. Looking at the markup difference pre-crisis versus crisis, both sets of simulated markups yield results close to the base model. In column R8, we depict the results from the dynamic model. Again, our main result prevails, although — as expected — estimated markups are higher when we account for the dynamic strategic effect. Overall, our results are robust to a number of modifications of our main specification. Counterfactual analysis ~~~~~~~~~~~~~~~~~~~~~~~ Three counterfactual experiments help to validate the results: two for the quasi- experiment and one to give a quantitative measure for our second hypothesis, i.e., the markup would have decreased by more had the state not intervened in the construction sector. To model the crisis, we manipulate the number of bidders and the backlog (own backlog, other bidders׳ backlog and new contracts). We conduct the analysis for the base model and the dynamic model. The split date for the pre-crisis vs. crisis period is the start of October 2008. For each experiment, we recalculate the bids and the backed out cost valuations. Bids are predicted using the Weibull model (5) from the base specification. To obtain predictions for bidders׳ cost, we estimate a Weibull model using as covariates variables that influence a bidder׳s cost such as own backlog, new contracts, distance, number of employees and the engineer estimate.36 For the counterfactual analysis, we manipulate own backlog, backlog of others, new contracts and the number of bidders in the bid distribution and own backlog and new contracts in the cost distribution. In the first experiment, we address the following question: what are the markups if the competitive situation did not change in the crisis? We apply the pre-crisis competitive environment in terms of number of bidders and bidders׳ backlogs to projects in the crisis. We define this environment as one bidder less and a backlog that is increased by one standard deviation. We also increase the value of new contracts by 5% to account for bidders׳ outside option in the private sector. According to the model, fewer bidders and higher backlogs will then lead to higher markups. Table 10 shows the implications of the base model and the dynamic model for all bids and for the winning bids. For both, all bids and winning bids, we first show the mean markups for the base model and then for the dynamic model. The results from our first experiment show that we can replicate our quasi- experimental results very well on average. Depicted in column (1), we predict average markups of 12.53% for projects procured in the crisis but evaluated in a pre-crisis environment (+1bl., −1n.). This value is close to the actual markup of 12.04%. Qualitatively we do not observe differences between the static and the dynamic model. Markups estimated from the dynamic model are on average by about two percentage points higher than markups estimated from the static model. Again, the experiment shows that we can replicate our quasi-experimental results very well on average. Evaluating projects procured in the crisis in a pre-crisis environment (+1bl., −1n.), we predict average markups of 14.00% close to the actual markup of 13.96% depicted in column (3). Projects before and during the crisis are not identical. The second experiment uses the pre-crisis projects and fixes the auction characteristics to answer the question: what happens to markups if the increased competition and lower backlog levels in the crisis would have been present before the crisis? We herewith control for changes in the composition of projects. We turn around the counterfactual from the previous experiment and apply it to pre-crisis projects (“−1bl., +1n.”). Given one additional bidder and lower capacity utilization before the crisis on the very same projects before the crisis leads to counterfactual effects that support the evidence from the first counterfactual. Thus, our results are robust to a change in the type of projects in the crisis. To model the government stimulus, we reduce the backlog of those bidders that won projects in the stimulus sector in the crisis. The question here is as follows: what is the quantitative effect of the government stimulus? We take the 17 stimulus projects that we were able to identify through parliament reports. We then randomly draw further projects in 2009 to reduce the backlog by the amount of the government stimulus directed to the construction sector, totalling 363.3 million Euro. The number of bidders increases by another one-third of a bidder relative to the crisis environment and the value of new contracts decreases by 2.5%. In this way, we aim to quantify the effects on bidding were there no government intervention. Lower backlogs and more bidders will lead to lower markups. To account for the fact that the government stimulus was implemented in the crisis, we apply the experiment to the crisis observations. The results of the counterfactual analysis support our second hypothesis. If the government had not intervened, markups would have dropped by more than we observe. Our third experiment provides the evidence. Using the estimates from the static model, we observe a counterfactual markup of 10.02% for all projects, on average. Compared to the observed markup of 10.52%, this is an additional drop of half a percentage point. This drop is slightly larger when we consider the estimates from the dynamic model. This difference may illustrate the lost option of winning tomorrow. Markups in columns (2) and (4) of Table 10 are for winning bids. While the markups are much higher for the winning bids, the general pattern, i.e. drop of the markup in the crisis, partial reversion of the drop in the markup due to the stimulus, is similar. The counterfactual analysis confirms the results we obtained for all bids, although the predictions for winning bids are less precise. Our results are robust to the specification of the model and we can also confirm our second hypothesis. As in the pre–post quasi-experimental analysis, we also observe in our counterfactual analysis that markups are pro-cyclical. The additional demand due to the stimulus package increased markups. Moreover, possible crisis-related changes in the type of projects do not drive our results.","Comparisons of pre-crisis and crisis markups consistently confirm Hypothesis 1: in the crisis, markups drop. Pre-crisis projects in the stimulus sector in general and the 17 stimulus projects in particular exhibit higher markups, but fall almost to the lower crisis markup level of projects not in the stimulus sector. It seems that the change in the competitive environment levels markups over different types of projects. We also provide evidence for Hypothesis 2, which says that the fiscal stimulus weakened the crisis drop in markups. Absent the stimulus, our counterfactual analysis shows that backlogs would have dropped further and (even) lower demand would have increased the number of bidders by more in the remaining auctions. Markups would have decreased by more, given the influence of backlog and number of bidders in the model. Thus, the two pieces of evidence consistently point to pro-cyclical markups. While developing a new method to analyse a structural dynamic auction with unobserved heterogeneity and endogenous entry, Balat (2012) looks at a specialized construction subsector, namely road construction. These specialized construction firms do not face demand reduction from the private sector, because demand comes from the state. As the state increases spending on road construction as part of the stimulus, these firms actually experience increased demand. From the perspective of the economy-wide construction sector this is a special situation, because overall demand is reduced in the crisis, as we have shown. Balat (2012) finds that markups increase in the situation of demand increase, analogously to the markups decrease we find when demand decreases for the whole construction sector. Our results show that the pro- cyclicality of the markup is witnessed in both situations — increases and decreases of demand. Our results also show that the increased prices of government projects in Balat׳s (2012) study carry over to the other, more general construction sectors in our environment. The typical motive for government spending in a recession is stimulation of the economy. We find that the government also faces lower markups and therefore lower prices for its projects. Domowitz et al. (1986) give empirical evidence on business cycles and their influence on the relationship between concentration and (average) price-cost- margins. They observe pro-cyclical margins in concentrated industries. As explanations, models of collusion and models of cost differences are suggested.37 With their industry- level data they do not give evidence allowing to discriminate between the explanations. Other recent studies by Nekarda and Ramey (2010, 2011) are also interested in testing the effects of government stimulus on markups — in particular in the context of different predictions from neo-classical and neo-Keynesian business cycle models. From neo-Keynesian models they expect countercyclical markups, given that prices are sticky while wages would increase due to demand increases, and therefore marginal cost would increase. The authors find pro-cyclical markups as we do, however, neither with economy-level nor with industry- level data do they find an effect of government stimulus on markups.","We estimate the cost distributions of construction firms based on detailed auction specific and firm level data. Based on the results, we measure the effects of competition on the bidding behavior and compare bidding behavior before and during the economic crisis. Our model permits quasi-experimental evidence on the effect of the crisis as well as the government׳s stimulus in the construction sector. The pro-competitive effect from free capacity and increased competition on the markup is estimated to be −1.5 percentage points overall and −3.3 percentage points for winning bids. Our evidence suggests that the fiscal stimulus, by injecting additional demand, counteracts the decline in markups. Thus, from two pieces of evidence we consistently find procyclical markups. As our results are based on careful structural modeling and detailed data, both auction specific and firm specific, our results should not only contribute to research in industrial organization, but could also be a useful input for macroeconomics to model the effects of activist fiscal policy."],["Lower-income countries spend vast sums on subsidies. Beneficiaries are typically selected via either a proxy-means test (PMT) or through a decentralized identification process led by local leaders. A decentralized allocation may offer informational advantages, but may be prone to elite capture. We study this trade-off in the context of two large-scale subsidy programs in Malawi (for agricultural inputs and food) decentralized to traditional leaders (“chiefs”) who are asked to target the needy. Using household panel data, we find that nepotism exists but has only limited mistargeting consequences. Importantly, we find that chiefs target households with higher returns to farm inputs, generating an allocation that is more productively efficient than what could be achieved through strict poverty-targeting. This could be welfare improving, since within-village redistribution is common. Productive efficiency targeting is concentrated in villages with above-median levels of redistribution. --------------------------------------------------------------------------------","Targeting programs such as subsidies to needy households is an important part of what governments do. To do this effectively, governments must first identify who is truly needy, which is difficult in developing countries where government infrastructure and information technology are limited (particularly in rural areas). Governments typically have the choice to administer such selection of eligibles centrally, or to decentralize authority to local communities (usually these programs are officially administered by local leaders).1 Decentralization has two main benefits: (1) local leaders are likely more informed about the relative neediness of people in their village (especially in a context in which most people do not file a tax return); and (2) local leaders will be more accountable to villagers, particularly if leaders face village electoral pressure or are motivated by reputation concerns. On the downside, decentralization may open the door for corruption or nepotism. This paper uses rich panel data collected from a sample of 1559 households over four survey rounds in 2011–2013 to explore this fundamental trade-off in the context of two subsidy programs in Malawi — the well-known farming input subsidy program (FISP) which provides subsidies for fertilizer and hybrid seeds once a year, and a one-time food aid relief program put in place after a financial crisis and drought in 2012. These programs were conceived as anti-poverty programs and the selection of beneficiaries was decentralized to local traditional leaders, called chiefs. How well do chiefs target the programs? This is a setting in which the trade-off between nepotism and information could be severe. On the one hand, nepotism is possible since chiefs cannot be held accountable via electoral pressure — in contrast to the contexts studied in Bardhan and Mookherjee (2000, 2005) or Bardhan (2002), the position of chief in Malawi, as in many other countries in the region, is hereditary and chiefs face fairly weak oversight. There is also no strict eligibility rule provided by the government (only general guidance on who should be “considered” for the subsidy) and no government back-checking of allocations.2 But on the other hand, local information is critical, along two main dimensions: (1) shocks occur frequently and chiefs likely have good information on recent household-specific economic conditions; and (2) the return to inputs will likely be heterogeneous across households within a village and related to factors such as household demographics (especially in regards to availability of family labor), soil type, and access to credit. Targeting inputs to those with the highest returns will increase total village output by the most, and if ex post inter-household transfers can be used to redistribute these gains, then targeting based on productive efficiency rather than neediness may be Pareto-optimal. The paper answers two sets of questions. First, how common are errors of exclusion (truly needy households not getting the subsidy) under the status quo? Do chiefs use local information to target households which have suffered recent negative shocks? Do they favor relatives? Second, do chiefs take into consideration productive efficiency when allocating the input subsidies? Specifically, do they target the agricultural subsidies to households with higher returns to fertilizer? To answer the first set of questions, we use observed food expenditures in the immediate pre-subsidy period as our measure of neediness, and benchmark the targeting effectiveness of the chiefs against that of a counterfactual proxy-means test (PMT). We find evidence that both the chiefs and the counterfactual PMT miss a substantial fraction of poor people, but that the chiefs miss significantly more: chiefs make more and bigger errors. Specifically, mean-squared error is 2–3.5 times higher for the observed allocation as for a counterfactual PMT. We also find evidence of nepotism: chiefs are more likely to target food subsidies to relatives. However, this nepotism appears to have minimal aggregate welfare consequences, since chiefs' relatives are similarly poor as other villagers. We also find that chiefs use their informational advantage to the benefit of households hit with negative shocks: people who have experienced droughts, floods, cattle death, or crop disease are significantly more likely to receive subsidies under the chief than under a PMT-based allocation. The second part of the paper tests whether chiefs target input subsidies to people with higher returns to agricultural inputs. The test is derived from a model of subsidy allocation in which chiefs have preferences over households, but also have information about household-specific returns to agricultural inputs. We assume that there is little heterogeneity in productive returns to food, in which case the allocation of the food subsidy is reflective of the welfare weights. To back out the relative importance of productivity considerations in the chief's objective function, we exploit the wedge between the allocations of the food and input subsidies. Taking this to the data, we find that chiefs indeed allocate relatively more inputs to households with higher gains from fertilizer use, while the PMT would not, suggesting productive efficiency gains from a decentralized system.3 As predicted by the model, targeting based on gains to fertilizer use is observed primarily in villages that exhibit above-median levels of income-pooling: it is only if the extra production can be shared ex post through inter- household transfers that targeting agricultural subsidies based on efficiency rather than poverty considerations can be Pareto-optimal. Our paper paints a nuanced view of the targeting of chiefs. On the one hand, we find evidence that decentralization has the benefit of improved information on recipients.4 On the other hand, we do find evidence of nepotism. As in Alatas et al. (2013), we find that the ultimate welfare consequences of nepotism are likely small, however, since a PMT would not perform much better.5 The main reason for this is that assets like land are noisy predictors of consumption in rural Africa — the R-squared for our PMT regression is only 0.32, and we document similar figures for datasets from Kenya and Uganda. This may be one reason why earlier work — including several previous studies in Malawi (Dorward et al., 2008, 2013; Kilic et al., 2013) had found higher levels of mistargeting and elite capture than we do: they used assets as a proxy for need instead of consumption. Our paper makes several contributions to the literature. Our core contribution is to bring attention to the difference between poverty-targeting and poverty reduction (the ultimate goal of subsidy programs). In communities with informal income-pooling, productive efficiency targeting may be the more effective (albeit indirect) way of reducing poverty. For this reason, looking only at who gets input subsidies rather than how the produced output is allocated is not sufficient to gauge impacts on poverty alleviation. More broadly, we contribute to the literature on the role of traditional authorities in African development. While survey evidence from the Afrobarometer suggests that traditional leaders are perceived to regulate important aspects of the local economy in numerous African countries (Logan, 2011; Michalopoulos and Papaioannou, 2013), the question of whether their existence further undermines weak governance, or instead palliates it, is still unsettled. Acemoglu et al. (2014) find that areas of Sierra Leone where competition among potential chieftaincy heirs was low during and after British colonial rule have significantly worse development outcomes today, but higher levels of respect for traditional authorities. They hypothesize that this reflects the ability of uncontested traditional ruling families to simultaneously capture resources and civil society organizations. Our evidence from Malawi mitigates this view: in our context, traditional leaders are uncontested and popular, as in Acemoglu et al. (2014), but effective at targeting input subsidies to productive farmers, possibly putting their village on a higher growth path.6 The layout of the paper is as follows. Section 2 presents some background on the Malawian local governance structure and decentralized subsidy programs. Section 3 discusses the sample and data. Section 4 presents evidence on poverty-based (mis)targeting. Section 5 tests for productive efficiency targeting. Section 6 concludes. Local governance in Malawi and the role of chiefs ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In Malawi, the democratically elected local government coexists with a traditional chieftaincy hierarchy. There are four ranks within this hierarchy: Paramount Chief, Traditional Authority (TA), Group Village Headman (GVH), and Village Headman (also known as village chief). In our data, TAs have authority over areas smaller than a district. They oversee from 10 to 45 GVHs, and GVHs oversee between 2 and 10 villages.7 Chiefs in Malawi hold little formal power. They do not have direct control over any public funds and are not allowed to raise local taxes. However, chiefs hold other customary responsibilities. The 1998 Decentralization Policy and Local Government Act Malawi Government (1998) recognized the rights of chiefs to allocate communal land and adjudicate matters related to customary law (in particular customary land). Chiefs also play an advisory and coordination role regarding local development projects. Finally — and this is the focus of our paper — chiefs are typically relied upon to identify beneficiaries for targeted government programs. Traditional leadership positions are hereditary, and chiefs who pass away are replaced from within the chieftaincy clan. Chiefs are paid a salary by the government that is known as mswahala, but it is fairly small.8 Chiefs do occasionally charge fees to villagers (in our sample, 44% of villagers report having ever made a payment to the village chief). Interestingly, chiefs are favorably viewed by the majority of the Malawian population. In 2008–2009, 74% of Afrobarometer respondents in Malawi perceived traditional leaders as having “some” or “a great deal” of influence, and 71% thought the amount of influence traditional leaders have in governing the local community should increase — for comparison, the average across 19 African countries for these two questions were both 50% (Logan, 2011). Possibly as a result of this high popularity, chiefs appear able to influence local villagers on whom to support in general elections and local government elections (Patel et al., 2007), an influence that may limit their accountability to elected representatives. Fertilizer subsidy program Malawi's Farming and Agricultural Input Subsidy Program (FISP) is one of the largest fertilizer and seed subsidy programs in the world.9 Though the program has existed since 1998, it greatly expanded after a drought in 2004 and steadily increased in size for a number of years after, until contracting more recently. In 2012–2013, the program reached 4.4 million recipients and took up 10–15% of the government's budget (Dorward et al., 2013; Baltzer and Hansen, 2011). In our data, the percentage of people benefiting from subsidies has increased steadily over time, from 63% in 2008 to 82% in 2012. The subsidy program covers several inputs and comes in the form of vouchers, which are redeemable at local agricultural shops. The four items covered by the voucher subsidy during our study period were planting fertilizer (a 50-kilogram bag of NPK, worth about $40 at market prices in 2013), top-dressing fertilizer (a 50-kilogram bag of Urea, comparable in price to NPK), hybrid maize seeds (a 5-kilogram bag, worth about $7), and hybrid groundnut seeds (a 2-kilogram bag, worth $2.60). The price of the voucher is only $1.7, so the subsidy was worth about 98% of the value of the input during this time period. As a result, take-up of the vouchers in our study sample is universal.10 There is no strictly defined, official eligibility criteria for the subsidy, but the intention is to target the poor and vulnerable. The official FISP guidelines reads that beneficiaries “will be full time resource poor smallholders Malawian farmers” but no threshold is provided for what defines “resource poor.” The program guidelines does hint at particular groups however: “...the following vulnerable groups should also be considered: elderly, HIV positive, female headed households, child headed households, orphan headed households, physically challenged headed households and heads looking after the elderly and physically challenged” (MoAFS, 2009). Many of these targeted groups may have lower returns to inputs than the average poor household, for example because they are unable to farm intensively. The identification of beneficiaries has three main stages (Chirwa et al., 2010). First, the government conducts a national farmer registration census. Second, the central government allocates vouchers to districts as a function of the area's farming population and the acreage under cultivation.11 Finally, within each village, once the number of subsidies available to the village is known, a list of eligible villagers is made. Formally, the selection of beneficiaries at this stage is supposed to be done by a Village Development Committee through open community meetings, and audited by the DADO. However, as we will show below, most authority appears to be de facto delegated to chiefs.12 Once the list of beneficiaries have been received by the DADO, it establishes a date and venue for the distribution of the vouchers themselves. The distribution is done by a staff member from the DADO. Listed beneficiaries have to show their voter registration card in order to receive the vouchers and also to redeem the vouchers at the retail stores (MoAFS, 2009). The identification of beneficiaries and distribution of vouchers is timed to precede the main rainy season (which runs from planting in November/December until harvest in April–August). During our study period, subsidy vouchers were distributed in September/October, in advance of planting. Food subsidy program Malawi devalued its currency in 2012, causing prices to rise 20–30% in 2012–2013 (World Bank, 2015), which made food imports prohibitively costly. There was also a poor harvest in 2012, caused by a drought. In response, a food subsidy program was implemented in late 2012, lasting from November 2012 to January 2013. In our area of study, the subsidies were distributed in kind. As with the input subsidy, the program was targeted at the “poor” but without a precise threshold or formula for what constitutes poverty. Of those receiving the subsidy in our data, the average amount received was 103 kg of maize, 14 kg of soy blend, 18 kg of pigeon peas, 10 kg of beans, and 3 L of oil. We estimate that this package was worth about $72 in 2013 USD. As with the farming input subsidy program, chiefs were given primary responsibility for identifying which households would receive the food aid. Sample ~~~~~~ The data we use for this paper was collected as part of a separate randomized controlled trial to estimate the impact of providing savings accounts to unbanked households (Dupas et al., 2018, henceforth DKRU). The project took place around the catchment areas of NBS bank branches in two districts of Southern Malawi — Machinga and Balaka. The sampling frame for DKRU relied on a census of market businesses and a census of households conducted at the end of 2010 — we use only the household sample for this analysis. The household census listed 9297 households from 68 villages in three Traditional Authorities (TA) areas: Kalembo, Sitola, and Nsamala. Of these, 78.8% met the eligibility criteria set by DKRU: they did not have a bank account and had a female head of household. DKRU randomly selected a subset of this sample for project inclusion, and completed baseline surveys with 2107 households. This set of households is uses for the analysis in this paper, though we must drop some households because their data is incomplete.13 We are ultimately left with 1559 households in 61 villages for our analysis. Given this sampling frame, our data departs from the universe of villagers in two ways. First, we systematically excluded villagers who had bank accounts at baseline (which was about 15% of the sample). These individuals are certainly richer than the average villager, and for this reason our analysis may underestimate targeting errors (if any of the people with bank accounts ended up receiving subsidies).14 Second, even among unbanked households, our dataset includes only a subset of people in each village (roughly 10% on average). However, since these villagers are randomly selected, our results are still internally valid and of interest — our goal is to understand how chiefs allocated subsidies within this sample, and our basic thought experiment is to ask what the gains would be from re- allocating subsidies within this sample. Household panel We have four waves of survey data for each household: (1) a baseline conducted from February to March 2011; (2) a first follow-up survey conducted from February to March 2012; (3) a second follow-up survey conducted from September to December 2012; and (4) an endline survey conducted from February to May 2013. The baseline survey includes a standard set of demographic variables, including a module on asset ownership which can be used to construct the allocation that would have obtained under a counterfactual allocation based on a proxy-means test from baseline assets. Each of these survey rounds included detailed expenditure modules. The follow-up and endline surveys include a module on the farming subsidy. This is used to construct a time series of subsidies received from 2008 to 2013, for each household. The module includes information on which input subsidy was received, whether the household received the voucher itself or shared another household's voucher, and what the household actually did with the subsidized products (used them, sold them, shared them, etc.). The endline survey also asked these questions for the food subsidy, which was introduced in 2012. Finally, the endline included a separate module with questions on how the input and food subsidies were allocated. These include questions on how (in the respondent's opinion) the vouchers were allocated, whether a public meeting was held, whether the respondent participated in the meeting, etc. In addition, between August and October 2014 we collected a fifth wave of data for a random subset of 563 households in the initial sample. This survey asked additional questions on the process through which subsidies were allocated and on respondents' attitudes towards the allocation process as well as their perception of their chief's role, beliefs and objectives in this allocation. Importantly, we also elicited households' beliefs on the returns to farming inputs on their own land. Chiefs survey Between August and October 2014 we collected surveys with all of the 105 traditional leaders in our study area of 61 villages, including 76 village headmen (chiefs) and 29 group village headmen (GVH).15 The survey included questions on their tenure and responsibilities, and included questions about the details of how the FISP and food subsidy programs were allocated. We also measured chiefs' self- reported awareness of whether some farmers had higher returns to inputs than others, and their knowledge of shocks encountered by villagers. Characteristics of households, chiefs and villages ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 1 presents basic summary statistics on the households in our sample. Panel A includes time-invariant characteristics collected at baseline. The first variable shown is the household's self-reported relationship to the chief. We asked the following question to each respondent: “Are you related to the chief?,” to which 27% reported yes. In a follow-up question, we asked: “How are you related?” The modal answer was the chief is an uncle (20% of the related cases), followed by brother (13%), brother-in-law (12%) and grandfather (12%). In what follows, we refer to those who reported as being related to the chief as “kin”.16 Households in the sample are very poor: 90% have mud floors or worse quality, 77% have thatch roofs, and less than 1% have electricity. Only 59% are literate, and average years of education for the household head is just below 5.17 Twenty-eight percent of households have no male head (most of these households are likely widows), and 97% own land. Panel B shows time varying expenditures, shocks and transfers. Across rounds, households report spending only $9.66 per month per capita in total, and the majority of this is on food ($6.80). These figures place these households well below the global extreme poverty threshold of $1.25 per day. Shocks are also quite common: 26% of respondents lost at least 1 day of work in the past month due to illness, 69% of respondents experienced the sickness of another household member in the past month, 28% experienced a drought or flood in the past 3 months, and 20% experienced crop loss or livestock death in the past 3 months. Across survey rounds, 72% of households report being worried about having enough food to eat in the past 3 months. Transfers across households within the village are very common, with 58% of households reporting being recipients of transfers in the last 90 days, and 25% reporting having made transfers. Columns 3 and 4 of Table 1 show, for each variable, the gap between kin and non-kin and its standard error. This reveals that if anything, kin are poorer than non-kin — they are significantly less educated (Panel A), and have slightly lower consumption (Panel B). Lastly, Column 5 shows the correlation between survey rounds for the variables in Panel B. This shows quite a bit of variability over time — the inter-round correlation in food expenditures is only 0.35–0.43, suggesting that neediness varies over time. Table A1 presents summary statistics on villages and village chiefs. The average village in our sample has 309 households and over 7000 acres of customary land. The average village chief is 53 years old and has about 5 years of education. Eighty-two percent of chiefs are male. The average chief has been in power for about 13 years, and 90% inherited the position (most of the remainder were appointed). The vast majority faced no competition from within the family blood line for the position. In principle, traditional leaders can be removed from office or reprimanded, but our data suggests this almost never happens: only one chief reported every being suspended. When chiefs were asked about their main responsibilities, the five most common responses were resolving conflicts among villagers (90%), reporting issues to higher level chiefs (61%), monitoring village projects (56%), disseminating information to villagers (33%), and overseeing subsidy programs (20%).18 Summary statistics on the allocation of subsidies in our sample ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table A2 presents summary statistics on the process through which input and food subsidies were allocated. Panels A and C rely on the latest round of survey data (2014) and presents evidence on how both chiefs (Panel A) and villagers (Panel C) experience and perceive the subsidy allocation mechanisms. Panel B presents data from the earlier household survey waves. The data confirms that chiefs are the primary decision-makers in allocating subsidies. Turning first to Panel A, the majority of village chiefs report that they have control over the subsidy allocation: 62% declare that they decide by themselves, while an additional 3% report that they decide in collaboration with others. Of the remainder, 13% report that the village development committee (of which the chief is a member) decides the allocation, and 13% report that subsidies are allocated in a village meeting (which the chief typically runs). When asked about selection criteria, chiefs report need as the primary criterion. Chiefs also put significant weight on female-headed households, households which recently received a shock households taking care of orphans, and households that the chief believes are hard-working. Panel B shows that community meetings regarding selection happen quite regularly: 95% of villagers report that a meeting was held, and 82% report attending this meeting for FISP (65% attended in regards to the food subsidy). Consistent with chief responses in Panel A, households responses in Panel C confirm that the chief is mostly responsible for allocating the subsidies — 72% report that the chief decides alone (49%) or with others (23%) on the input subsidies, and 73% report that the chief alone decides on the food subsidies. Households report similar inclusion criteria as do chiefs (needy households, as well as elderly and female-headed households). While official FISP guidelines do not endorse sharing of subsidy packages, we find strong evidence that sharing is in practice very common (Web Appendix Table W1). Seventy-seven percent (0.46/0.60) of households who received an input subsidy voucher report sharing it. Moreover, we find that sharing is often at the direction of the chief: of those who shared, 83% say they received instructions from the chief on whether to share it, and 79% received specific instructions from the chief on whom to share with. Food subsidies are similarly shared.19 In what follows, we perform all analyses considering both allocations: the allocation of the vouchers themselves, and the allocation observed after sharing (we call this the “realized allocation”).20 Measuring neediness ~~~~~~~~~~~~~~~~~~~ To measure neediness, we use food expenditures, which we consider a proxy for consumption. Food expenditures have been shown to be better predictors of neediness than other measures such as income (Deaton, 1997; Meyer and Sullivan, 2012). While we measured expenditures on 12 broad food categories (covering all food types), in the main analysis we focus on the 10 categories that are typically purchased rather than self-produced.21 These 10 categories are vegetables, fruits, meat, dairy/eggs, salt, sugar, other cooking items (oil, margarine), coffee and tea, snacks, and juice/sodas. We compute the sum of expenditures on these 10 food categories over the 30 days preceding the survey and then divide the sum by the number of household members to construct “per capita non-staple food expenditure” or PCF, our measure of need going forward (we report this figure in USD).22 The distribution of log PCF in our data is plotted separately for the two main years of analysis, 2011 and 2012, in the top panel of Web Appendix Fig. W. Timing The food expenditure we would ideally use to determine “true need” (PCF eligibility) would be measured at the time that subsidy beneficiaries are identified (which is around August for the input subsidy and November for the food subsidy). The timing of our surveys does not precisely correspond to these periods. Our food expenditure module covered the last 7–30 days (depending on the question) before the survey date. Thus, given the dates of the surveys mentioned in Section 3.2, we have consumption data for the following periods: January 2011 to February 2011; January 2012 to February 2012; August 2012 to November 2012; and January 2013 to April 2013. To study the targeting of the 2011 input subsidy, we thus have to rely on the January 2011 to February 2011 expenditure data, which is substantially before the period of interest. In particular, it is before the March 2011 harvest, which is likely an important determinant of actual neediness as of August–November 2011.23 Fortunately, the data used for the 2012 subsidies is for the correct time period (August to November). For this reason, our 2012 results are our preferred estimates. PMT score rank To construct the counterfactual in which subsidies were allocated via PMT, we repeat this procedure but this time we rank households (within each village) by a “PMT score.” We compute the PMT score as follows: we regress log PCF on household characteristics, including demographic characteristics, dwelling characteristics, assets and occupation, and use the estimated coefficients to predict a score for each household. As in Alatas et al. (2012), we do this in two steps: we first run kitchen sink regressions with all available characteristics and then, using a backward step-wise procedure, keep only those characteristics which are statistically significant at the 10% level. PMT regressions are shown in Table A3. We show the results for both per capita and per adult equivalent food expenditure, and find slightly higher predictive power for per capita values.24 From Column 1, we obtain a R-squared of 0.32, which is somewhat lower than the 0.40 obtained by Alatas et al. (2012) in Indonesia (when pooling districts together).25 For comparison, we also construct a PMT score using data from the 2010–2011 wave of the Integrated Household Survey (IHS3), a representative household survey collected by Malawi's National Statistics Office. We restrict that dataset to the two districts in our sample, and estimate PMT regressions using the same backward step-wise method to identify covariates. Results are shown in Web Appendix Table W2. In the table, we run regressions separately where we restrict to only those variables which were also collected in our surveys (which we call “BDR variables”), which are shown in Column 1, and for all potential covariates available in the IHS3 (Column 2). We find R-squared statistics in both regressions of approximately 0.4. We conjecture that the somewhat lower R-squared we observe in our own survey data is because our sample is somewhat poorer than a representative sample, and their consumption may be more volatile due to lower access to insurance. To shed some light on this, we run similar regressions in samples of unbanked households we have collected in other work in Kenya (Dupas et al., 2019) and Uganda (Dupas et al., 2018). We find an R-squared of 0.31 in Kenya and 0.28 in Uganda. PMT vs. chief allocation Our first set of results is shown in Fig. 1, which plots the probability of receiving the subsidies by quintile of the PMT score distribution (top panel) and quintile of the PCF distribution (bottom panel). These quintiles are across the entire sample, and so include across-village variation. We show the realized allocation (i.e. the allocation after vouchers were shared) as well as two counterfactual allocations: the PMT allocation, our “benchmark” for what could be done under centralization; and the PCF-based allocation, the “optimal” allocation. We pool across villages, which vary in their underlying distributions as well as in the number of subsidies available, which explains why neither of the two counterfactual allocations are perfect step functions of their respective distributions. It also explains why even the PCF-based allocation in Fig. 1 does not reach perfect targeting: there is mistargeting of the number of subsidies across villages, which means that even a perfect allocation within village would yield evidence of mistargeting. The gradient in the PCF-based allocation in Fig. 1 should therefore be considered as the “best possible targeting” given the across- village allocation in our data.26 From the top panel of Fig. 1, it is clear that chiefs target different people than the PMT would: while the PMT, by definition, would allocate subsidies to 100% of people at the bottom of the distribution, the chiefs' allocation has a much flatter gradient with respect to the PMT score. In isolation, this result looks similar to Dorward et al. (2008, 2013) and Kilic et al. (2013), who look at how well chiefs target based on assets and conclude that there is widespread mistargeting. The bottom panel of Fig. 1, which show targeting based on PCF, also show that the PMT does better than chiefs — but the gap is much smaller than in the top panel. In the allocation decision of 2011 (which was contemporaneous to the survey from which the PMT was calculated), the gradient for the PMT allocation is quite a bit steeper than that of the chiefs, but by 2012 the slopes are more similar. This could be because characteristics measured in 2011 become less and less predictive as time goes on, and might suggest that the advantage of a PMT may be short-lived. Fig. 1 also shows that the PMT makes a substantial number of errors. This is true even if we use the PMT formula from the IHS3 rather than the one derived in our dataset. The relatively poor targeting performance of the PMT seems due to the fact that assets (the most important factor in the PMT) are a relatively poor proxy for need in our study context, because PCF eligibility is not time-invariant (the correlation between food expenditures across rounds is only 0.35 as previously discussed and shown in Table 1) and because there are important unobservables in the determinants of PCF. In Fig. 2 we show the allocation against the PCF quantile, both before and after sharing. Before sharing, just over 50% of households received the input voucher and 34% received the food voucher; after sharing, these percentages increase to about 78% and 59%. However, poverty-targeting efficiency does not improve from sharing: Fig. 2 shows that the slope of the realized allocation is identical to the slope of the initial voucher allocation, suggesting that the sharing happens primarily within quantile of the PCF rather than across. Error rates Table 2 shows the average village error rate (averaging first over individuals within villages, and then across villages) under the two allocation schemes (chiefs and PMT). For these calculations, we include only those villages in which the probability of getting a subsidy is between 0 and 100% (so that targeting errors are possible).27 The poverty-targeting error rate is the probability that a household is (1) eligible based on its position in the PCF distribution within the village; but (2) does not make it onto the actual beneficiary list (chief error) or on the counterfactual PMT beneficiary list (PMT error). Note that since the number of beneficiaries within the village is kept fixed in this exercise, this error rate also provides the probability that a household is (1) categorized as ineligible based on its position in the PCF distribution and (2) gets the subsidy. In other words, mechanically there are as many people who don’t get the subsidy when they should (exclusion errors) as there are people who get the subsidy when they should not (inclusion errors). We also show what the expected error rate would be if subsidies were allocated randomly. These are calculated from a permutation test with 1000 draws. Finally, we also compute the squared error for each allocation. We can see that both allocations make a significant number of errors compared to the PCF-based allocation, but that the PMT always has a lower error rate. The error rates for chiefs is 15.8% and 14.4% for the 2011 and 2012 input subsidies, while the PMT's error rate is only 10.3% and 10.9%. For the food subsidy, the chief's error rate is 15.1%, compared to 13.7% for the PMT. Since not all errors are equally important (i.e. denying a subsidy to somebody just barely under the threshold is not nearly as costly as denying a very poor person), a more informative measure of errors may might be the mean-squared error (shown in the bottom of the table). Here too we see consistent evidence that the PMT has a lower MSE than do chiefs, across all subsidies types. Finally, we see that the PMT based on our data does consistently better than that based on the IHS3; however, both outperform the chiefs' allocation. While chiefs do worse than the PMT, they do better than random (see Table 4). For the input subsidy, the simple error rate for chiefs is not statistically distinguishable from random, but the mean squared error is much lower, suggesting that chiefs trade PCF eligible for ineligible only around the PCF cutoff. Chiefs also do better than random on the food subsidy, by both metrics (Table 2). An interesting pattern in these results is that, compared to the PMT, chiefs look worse at targeting the truly needy for the input subsidy than for the food subsidy. A central hypothesis of this paper is that this may be due to productivity targeting of the input, which we will argue is less relevant for food. We dive into this issue in detail in Section 5. Who is favored and who is left out by chiefs? Table 3A shows the results of a multivariate regression of the realized allocation (i.e. receiving a voucher or a share of a voucher) on background characteristics and village fixed effects. Columns 1–4 show regressions for the real-life allocation (decided by chiefs) while Columns 5–8 show a counterfactual allocation if the subsidies were allocated by the PMT formula. Table 3B performs the same analysis, but for receiving the voucher itself (i.e. not including people who received the subsidy via sharing). Comparing the coefficient estimates across the three sets of analyses tells us who is favored and who is left out under each scheme. We consider both the extensive margin (receiving any subsidy) and the intensive margin (the value of the subsidy received, since this varies across households due to sharing).28 The first row of Table 3A confirms the poverty- targeting results discussed above: the gradient in PCF is more negative under the PMT than under the chiefs, and the gap in the gradient is more pronounced for the input subsidy than for the food subsidy. We find evidence of nepotism: conditional on covariates, chief's kin are 11 percentage points more likely to receive the food subsidy under the chief, whereas they would not be favored under the PMT. For the input subsidy, nepotism appears much less pronounced: while chief's kin receive a greater input subsidy package (an extra 3.30 kg compared to a mean of 50.5 kg, significant at 10%), the PMT would also award kin higher subsidy packages (+2.3 kg, also significant at the 10% level). This is due to the fact that chiefs' kin are marginally asset poorer than non-relatives. Turning to other covariates, we find that chiefs target older households, as per the official FISP recommendation. Chiefs also target households that received negative shocks: households who experienced a drought or flood are 4 percentage points more likely to receive subsidized food, while households who experienced crop loss or cattle death are 8 points more likely to get it. By contract, the PMT is not designed to respond to shocks, and indeed we find no correlation between shocks and subsidy receipt in the PMT (such a correlation might exist if shocks are strongly correlated with asset poverty). Discussion of poverty-targeting results ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The results in Tables 3A and 3B epitomize the trade-off between local information and capture: we find that chiefs are able to use local knowledge to benefit households hit by recent negative shocks, while the PMT misses them; but they also favor their kin. These results raise several questions. First, is the fact that kin are more likely to get subsidies evidence of nepotism? An alternative hypotheses is that chiefs have better information on relatives, and therefore are more likely to target kin because they can be certain that they are truly poor. If this is the case, we would expect that subsidies to kin would be more responsive to consumption than to non-relatives. We investigate this in Table A4, in which we include an interaction between log food and kinship. We find no evidence in favor of the information hypothesis: targeting actually appears somewhat worse for relatives for the input subsidy, though there is no effect for the food subsidy. We also do not find that targeting based on shocks is better among relatives (Table A4), consistent with the fact that kin are favored irrespective of whether they faced a shock. While we lack data to definitively rule out an information story, our evidence appears more consistent with nepotism.","In this section, we investigate whether some of the apparent mistargeting of input subsidies by chiefs is due to targeting on farming productivity: if returns to input subsidies are heterogeneous and chiefs have information on this, then they might allocate subsidies in a way that takes both poverty-targeting and productive efficiency into account. We use a simple model that allows for heterogeneity in returns as well as heterogeneity in the welfare weights that chiefs assign to households, to derive a test of whether the mistargeting we observe for input subsidies is in part driven by productive efficiency considerations. Model and prediction ~~~~~~~~~~~~~~~~~~~~ We consider the problem of allocating subsidies across households within a village. The intra-village allocation is done by the village chief. This leads us to the prediction we can test in the data: Allowing chiefs to orchestrate transfers Productive efficiency considerations when allocating farming subsidies increase as the level of ex post income-pooling in the village increases. Prediction 2 brings attention to the fact that the two subsidies we study could be complementary: the input subsidy as a growth instrument and the food subsidy as a redistribution instrument. This logic could also imply that the food subsidy and input subsidy allocations could be related — farmers who received the input subsidy should have larger harvests and be less in need of the food subsidy. Note that this does not invalidate our test: the food subsidy should be allocated based on Pareto weight and current consumption, irrespective of whether the current consumption level was secured through enhanced yields in the previous period thanks to inputs subsidies or not. Relative Pareto weights can still be backed out from jointly observing the food allocation and current consumption, as we do.31 Below we show that our predictions hold under a number of extensions to the basic model. Results ~~~~~~~ To test predictions, we need a measure of A, the household-specific (farm-specific) productivity of fertilizer. In this or any context, estimating the productivity of an input is very difficult, since input choices are endogenous and farmers with higher returns are presumably more likely to use fertilizer in a given season. Returns are also volatile across years and even within farms, so estimating this well would typically require a long panel. Instead of estimating productivity, we therefore opted to simply ask farmers for their expectations of yields with and without fertilizer use. We collected this data in the fifth survey round conducted in the summer 2014. There are several important caveats. First, due to budget constraints the survey could only be done with a random subset of households in each village. The sample includes only about one third of the sample. Second, the questions are about total output with and without fertilizer, rather than marginal returns.32 We show the means of the reported expected yield in Panel C of Table 1, and we plot the distribution of the reported gain in total output in Panel B of Fig. W in the Web Appendix. There is substantial heterogeneity in these reported gains from input use. What drives it? Table A5 examines correlates of self-reported gains. We regress the reported log yield increase on log acres and other observables. We find that reported gains are correlated with many variables, including household demographics (gains are increasing in the age of the head of household), education, log assets, and household size (though this is not statistically significant). We expect that these are the types of proxies that the chief may use to target subsidies, in addition to other characteristics that are unobservable to us, such as land quality. Also of note is that the correlation between estimated production gains from using fertilizer and our measure of neediness, PCF, is fairly weak (Panel C of Fig. W). We also find no systematic differences by kinship status (Table 1 Panel C, column 3, and Table A5). These results are in sharp contrast with those for the counterfactual PMT distribution, in which the value of the subsidy is actually (insignificantly) declining in the gains to fertilizer (because of a negative correlation between returns and assets). In that case, the PMT undermines the effect of the subsidy on total farm output at the village level. In contrast with the chief's allocation, the gap between fertilizer and food subsidy values does not significantly increase with reported gains from fertilizer under the PMT allocation (Table 4, column 6).33 Overall, the results in Table 4 and Fig. 3 are consistent with chiefs taking productive efficiency into consideration when allocating input subsidies — something that the PMT cannot do since information on who has more to gain from fertilizer use is not something that can be elicited in an incentive-compatible way if people expect their subsidy package to depend on it. The magnitude of the effects is not trivial: a household with an extra log point gain from fertilizer gets about 6.5 more kgs of input subsidies under the chiefs than under the PMT. Supportive evidence ~~~~~~~~~~~~~~~~~~~ Is information on the relative productivity of various potential beneficiaries of the input subsidy embedded in the chief, or does it rest in the people themselves? People who have high value for the input subsidy could wait in line more, lobby more or protest more if they don’t get the subsidy, such that the allocation of the chief ultimately favors them in a way that looks as if the chief himself were aware of the heterogeneity. To provide descriptive evidence on this question, in the 2014 survey, we asked respondents if they had ever lobbied the chief to obtain subsidies. Only 9% of respondents reported lobbying for input subsidies, and 4% reported lobbying for food subsidies (Table W1). The likelihood of having lobbied is not positively correlated with returns to fertilizer for the overall sample (see Table 4, column 7), though it is correlated among the chiefs' kin (see Table A6, column 7). However, we argue that the this lobbying is of modest importance, since kin lobby much less on average, and overall the targeting efficiency is not higher among kin as shown in Table A6 columns 1–3. In the survey of chiefs also conducted in 2014, we asked chiefs a number of questions about what they could observe about households, which we present in Web Appendix Table W7. We find that 86% of chiefs report that they can easily categorize farms in their village in terms of productivity of inputs. Chiefs also report that they know who works harder, who has money for inputs, and whose returns are highest. While descriptive, these responses are consistent with chiefs having significant local knowledge. Threats to validity ~~~~~~~~~~~~~~~~~~~ In this section, we discuss several possible threats to internal validity: (1) the fact that returns to inputs are self-reported rather than observed, and (2) the fact that the two subsidy programs we consider exist alongside other social programs which may be allocated simultaneously. Using self-reported returns A possible concern with our analysis is that returns are self-reported rather than directly observed, and so could potentially be correlated with various omitted variables. We present several pieces of evidence to help address this. First, we use our data to construct an agricultural panel. Specifically, we have complete data for the 2010–2011 and 2011–2012 planting seasons in this paper. From this we have at most 2 observations on households about their fertilizer use and output. We utilize this by running fixed effects regressions of output on input usage — relying on variation in input usage that occurs over years. The key variable in this regression is an interaction between self-reported gains and input usage — if the measure is valid, then this correlation should be positive. We show results in Table A7. We find strong evidence that these (admittedly non-random) returns to fertilizer are higher for those with higher self-reported returns. While we do not want to make too much of this since input usage is potentially endogenous, this fixed effects specification does rule out some time-invariant sources of bias — for example, land size and household demographics are held fixed in this analysis. At least descriptively, these results seem to support our interpretation. Beyond these results, we argue that many (though not all) stories for why households might get more subsidies would apply to both food and fertilizer subsidies. For example, one might argue that people with higher returns are more confident and have higher social status and are therefore more likely to get subsidies. But these sorts of stories would not explain our results, since these households would get more fertilizer and food, whereas our main empirical tests are about the difference in the value of the two types of subsidies. We also find that this relationship is stronger in villages with higher levels of sharing. While this is consistent with the framework we have written down, it does not seem likely that we would observe this particular pattern (which was derived ex ante) if the results were driven purely by omitted variables. Other safety net programs Beegle et al. (2017) document that chiefs are also involved in deciding which households are eligible for Malawi's public work program (PWP) — though the responsibility falls more on the Group Village Headmen and the villagers themselves. They report that Malawi's PWP “has been operational since the mid-1990s and aims to provide short-term labor-intensive activities to poor, able- bodied households for the purpose of enhancing their food security.” While we did not collect data on participation in the PWP directly from respondents in our surveys, a fuzzy name match between the original household sample and administrative data on PWP participants obtained from the two districts in our sample yields 167 matches for the 2012–2013 budget year, out of 2107 households in the DKRU baseline survey, suggesting that the PWP coverage in our study area is about 8%. Verification surveys with a subset of those matched and unmatched conducted in March 2015 suggests that an additional 3% may have been participating in PWP, bringing our estimates to roughly 11%.35 While name matching is always prone to significant error, this ballpark figure is not far from the 15% coverage targeted by the program. While studying how the PWP is targeted and the specific role of chiefs would have been interesting, omitting it due to data limitations should not affect our analysis of the other subsidy programs. Notably, Beegle et al. (2017) find no correlation between receipt of PWP and receipt of other benefits, suggesting no “fairness norm” influencing distribution across programs, in particular, no compensation of non-PWP households with input or food subsidies.","Traditional leaders, often known as “chiefs,” have maintained a significant amount of de facto if not de jure power in sub-Saharan Africa. Possibly owing to the weakness of local governance in most of the continent, chiefs are commonly involved in the decisions of how to allocate government resources. One prominent type of resource is subsidies. Developing country governments allocate an important portion of their national budget to subsidies targeted at the poor, and it is common for chiefs to be asked to identify who should be eligible for such subsidies. Do chiefs identify the right beneficiaries? Previous work on this question in Malawi concluded that there was widespread elite capture (Dorward et al., 2008; Kilic et al., 2013), based on evidence that “connected” households are more likely to receive subsidies, and that household assets measures do not strongly predict subsidy receipt. We show that such evidence may not directly speak to the issue of poverty- targeting in environments where assets are a poor predictor of need, and where the subsidized items are productive inputs. We find evidence that chiefs allocate input subsidies to farmers with larger returns to input use. This result underscores how a naive measure of targeting based solely on the neediness of households (even when neediness is well measured) may understate the poverty-alleviation impacts of the allocation: when ex post redistribution is possible through informal transfers, targeting input subsidies based on productive efficiency (i.e. using input subsidies as a growth instrument) can have a larger impact on aggregate welfare than targeting based on poverty would. This issue has not received much attention in the literature up to this point, even though most of the inputs subsidized by governments are productive (farming inputs, health products) that have heterogeneous returns. Future work should explore whether our results generalize to other contexts and countries."],["We study a dynamic social choice problem in which a sequence of committees must decide how to consume a public asset. A committee convened at time t decides on consumption at t, accounting for the behaviour of future committees. Committee members disagree about the appropriate value of the pure rate of time preference, but must nevertheless reach a decision. If each committee aggregates its members’ preferences in a utilitarian manner, the collective preferences of successive committees will be time inconsistent, and they will implement inefficient consumption plans. If however committees decide on the level of consumption by a majoritarian vote in each period, they may improve on the consumption plans implemented by utilitarian committees. Using a simple model, we show that this occurs in empirically plausible cases. Application to the problem of choosing the social discount rate is discussed. --------------------------------------------------------------------------------","Suppose that a society needs to decide on an intertemporal consumption plan for some public asset. A committee is convened at each point in time, and tasked with determining how much to consume in the current period. The members of each committee have differing opinions about the pure rate of social time preference (PRSTP), or utility discount rate, that should be applied to this problem. Some favour a high discount rate, while others believe that different time periods should be treated more equally, and thus favour a low discount rate. Moreover, the current committee knows that future consumption choices will also be made by committees exhibiting similar disagreements on discount rates. How should such committees proceed, given the heterogeneity in opinions on discount rates? Although it may seem abstract, this question is inspired by an important practical problem in public economics: how should governments discount future utilities when evaluating public policy decisions? The appropriate normative value of the PRSTP has been debated at least since Ramsey’s (1928) seminal work on optimal national savings. Subsequent commentators have argued the merits of a variety of values for the PRSTP without a clear ‘best’ value emerging, and different governments have adopted different values for public decision- making. The social time preferences economists prescribe for public decision-making today are still highly heterogeneous (Arrow et al., 2013). This has been highlighted by the long-standing debate about the appropriate value of the PRSTP for the evaluation of climate change policies (Nordhaus, 2008; Stern, 2007). A recent survey of experts on social discounting (Drupp et al., forthcoming) shows significant variation in their prescriptions for the PRSTP (see Fig. 2). Given the persistent normative disagreements about the PRSTP, it is natural to ask whether methods from social choice theory can be used to obtain a compromise between opposing viewpoints. In this paper, we examine perhaps the most common such methods: utilitarian aggregation and majoritarian voting. Under the utilitarian approach, committees seek to maximize a weighted sum of the time preferences advocated by their members in each period, while under majoritarian voting, committee members vote on the current level of consumption, and a Condorcet winner (if it exists) is implemented. The utilitarian approach is appealing, as Jackson and Yariv (2015) have shown that any social choice rule that is non-dictatorial (i.e. sensitive to the preferences of more than one individual) and respects unanimity (roughly, if everyone prefers consumption stream C to C′ then C is socially preferred to C′) is equivalent to utilitarianism in the setting we study. However, while no-dictatorship and unanimity are compelling properties in isolation, they lead to a time inconsistency problem when combined with another assumption: time invariance (i.e. preferences over future consumption streams are identical in all time periods). Millner and Heal (2018) have argued that while time invariance is an excessively strong assumption in intra-group intertemporal decision problems (e.g. allocation between family members), it is plausible when modeling inter- group choices like those facing the successive committees studied in this paper. Thus, if a utilitarian approach to resolving disagreements is adopted, the collective preferences of successive committees will conflict with one another. Rational utilitarian committees will anticipate the actions of future committees, and react optimally to them, inducing a dynamic game between committees. The equilibrium of this game will be seen as inefficient by every committee. The inefficiency of the consumption path implemented by utilitarian committees means that it is possible that voting could give rise to superior outcomes. If each committee holds a majoritarian vote on the level of current consumption, and members of the current committee rationally anticipate the outcome of future votes, we show that the equilibrium consumption path under voting will correspond to the optimal plan of the median member. Further analysis shows that a majority of committee members will prefer this voting equilibrium to the utilitarian equilibrium, regardless of the choice of aggregation weights in the utilitarian objective function. We extend this result to welfare comparisons, finding conditions on the distribution of PRSTPs under which the voting equilibrium is superior to the utilitarian equilibrium according to utilitarian committees' own objective functions. Using survey data on economists' recommended values for the PRSTP, we show that these conditions are often satisfied in practice. There is thus a sense in which voting may be ‘self-stable’ (Barbera and Jackson, 2004) relative to utilitarianism: a majoritarian vote between voting and utilitarian aggregation of PRSTPs will always lead to voting being adopted as the aggregation method. By contrast, a utilitarian comparison of voting and utilitarian equilibria will often favour voting. The paper is structured as follows. We discuss related literature next, before developing our simple model of dynamic public choice with disagreements about the PRSTP in Section 2. This section contains the bulk of our analysis. We first derive the equilibrium behaviour of utilitarian committees, and show that they choose inefficient consumption paths. Next, we derive the equilibrium behaviour of committees that vote on consumption. Finally, we contrast these two preference aggregation methods, deriving results on committee members' ordinal preferences between the implemented equilibria, and comparing them from the perspective of utilitarian committees' own collective preferences. Section 3 discusses the results, and draws some lessons for the choice of the PRSTP in social discounting formulae. Related literature ~~~~~~~~~~~~~~~~~~ The literature on aggregation of opinions about social discount rates stems from the work of Weitzman (1998, 2001), who focuses on aggregation of expert opinions on real (i.e. consumption) discount rates, rather than pure time preferences. Weitzman takes a sample of opinions as to the appropriate (constant) real discount rate for project evaluation, treats these as uncertain estimates of the ‘true’ underlying rate, and takes expectations of the associated discount factors to derive a declining term structure for the ‘certainty equivalent’ real discount rate. As Freeman and Groom (2015) observe, opinions about real discount rates conflate ethical views about welfare parameters such as the PRSTP with empirical estimates of consumption growth rates — they mix tastes and beliefs (see Dasgupta, 2001 , pp. 187–190 and Gollier, 2016 for further discussions of Weitzman's approach). This suggests that it is important to pursue approaches that treat preference aggregation as a distinct problem. Our work highlights difficulties that may arise in practice when decision-makers with a distribution of ethical views attempt to form consensus social preferences, and contrasts the equilibrium outcomes that arise from standard preference aggregation methods. The possibility that utilitarian preference aggregation could lead to time inconsistency when agents favour different values of the PRSTP has been noted by several authors (Marglin, 1963; Feldstein, 1964; Jackson and Yariv, 2015). Millner and Heal (2018) argue that, while this is not a generic feature of utilitarianism as an normative theory (see also e.g. Hammond, 1996), as a positive matter it is likely to occur when distinct groups of agents are tasked with decision-making in each time period, as occurs in the setting we study here. Our work thus falls somewhere on the boundary between normative and positive analysis: we study positive properties of the equilibrium consumption choices that would be implemented by sequences of committees that seek to aggregate their members' normative views on social time preferences. Alternative approaches to the aggregation of time preferences are pursued by Gollier and Zeckhauser (2005), Jouini et al. (2010), and Millner (2018).","We focus on a sequential social choice problem in which a sequence of committees, each composed of N > 1 members indexed by i = 1…N, must choose how to consume a public asset. For the sake of analytical convenience, we assume that N is odd, and that time is continuous. Each committee exists for a single moment in time, and controls the value of consumption in that moment alone. Committee members are drawn from a stable population at each moment, and their tenure lasts for only that moment. The distribution of members' opinions on the PRSTP is assumed to be independent of time.1 Utilitarian aggregation ~~~~~~~~~~~~~~~~~~~~~~~ To determine the limit equilibrium, we must find the linear MPE of the dynamic game between committees. A Markovian strategy in our context is a function σ(S) such that consumption at time t is given by Ct = σ(St) for all t ≥ τ. A strategy σ(S) is an MPE if, in the limit as ϵ → 0, when committees at times t ∈ [τ + ϵ,∞) in the future use the rule σ(S), the best response of the current committee in t ∈ [τ,τ + ϵ) is also to use σ(S). An MPE is linear if the equilibrium strategy is of the form Ct = σ(St) = ASt for some A > 0. The next proposition characterizes the limit equilibrium of the game between committees: The possible non-existence of a linear MPE (and hence a limit equilibrium) when η < 1 is a well known feature of models like ours (see e.g. Phelps and Pollak, 1968). To avoid existence problems, we assume that η ≥ 1 in the remainder of the paper.6 From the perspective of a committee in any period τ, the equilibrium described by Proposition 1 is inefficient. That is, there exist feasible consumption paths that would increase its welfare measure Wτ. However, owing to the time inconsistency of committees' preferences, these paths are not implementable. Any future committee at time τ′ > τ can increase its welfare measure by deviating from the time τ committee's optimal plan. The time τ committee knows this, anticipates the behaviour of all future committees, and reacts optimally to this knowledge. Since all committees behave this way, the resulting equilibrium is inefficient, but fully rational. Thus, although each committee's welfare measure aggregates its members' preferences efficiently, the interactions between successive committees lead to an inefficient intertemporal equilibrium. Voting ~~~~~~ The fact that utilitarian committees choose inefficient consumption plans in equilibrium suggests that alternative methods for aggregating member's opinions could improve on utilitarian preference aggregation. The most natural alternative to consider is majoritarian voting. Aside from being widely deployed in practice, majority rule has been shown to satisfy desirable properties of preference aggregation over a larger domain of preferences than any other ordinal social choice rule (Dasgupta and Maskin, 2008, see also May, 1952). Yet as Sen (2017, p. xxvii) observes, ‘when it comes to welfare economics, majority decision is not a particularly just, or even plausible, way of judging alternatives'. Sen is referring here to the fact that decisions implemented by majority rule will generally not promote more comprehensive measures of social welfare. Indeed, while majoritarian ballots have been a staple of the positive theory of public choice for decades (e.g. Black, 1948; Downs, 1957; Meltzer and Richard, 1981), they are seldom invoked as a means to pursue normative social objectives. We will show below however that our model provides one instance in which majoritarian voting may be desirable according to such a normative objective. Voting on consumption may lead all committees to achieve higher levels of utilitarian welfare Wτ than attempting to maximize Wτ directly. Suppose that consumption Cτ is to be decided by ballot in each period. In each period τ, each committee member may nominate a single value of Cτ. All members vote over each pair of nominated consumption values, and the value that gets a majority of votes wins each pairwise contest. A Condorcet winner (if it exists) is a value of Cτ that wins every pairwise contest. If there is a Condorcet winner, it is implemented. Since the current choice of Cτ influences the consumption choices that will be made in future ballots, committee members must anticipate the outcomes of those ballots when forming their preferences over current consumption Cτ. We assume that members are rational, and thus anticipate the outcomes of all future ballots when forming their preferences over Cτ.7 The following proposition characterizes the equilibrium consumption path that emerges from this sequence of ballots: Committee members at time τ who anticipate the outcome of votes over public consumption in future periods have single-peaked preferences over current consumption Cτ. Thus, the equilibrium of a majoritarian voting model with ballots in every period is the optimal consumption plan of the median member. This result may seem to be in conflict with the analysis of voting over consumption streams in Jackson and Yariv (2015). They show that voting over consumption streams in unrestricted domains is generically intransitive, and thus voting equilibria cannot be represented by the preferences of a single individual such as the median agent. Their analysis assumes however that votes are once off, whereas in our result ballots are repeated, so that in each period members are only voting over a single, unconstrained, value of consumption. The repeated ballot formulation is compelling, as it does not require the somewhat far-fetched assumption that members believe that a consumption plan that is decided on today will automatically be implemented by all future committees. Utilitarian aggregation vs. voting ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We are now in a position to compare the equilibrium implemented by utilitarian committees to that implemented by committees that vote on consumption. Our first result provides a lower bound on the number of committee members who prefer the voting equilibrium to the equilibrium implemented by utilitarian committees: In every period, a majority of members prefer the voting equilibrium to the equilibrium implemented by any utilitarian committee. This result is independent of any assumptions about utilitarian committees' aggregation weights. Thus, regardless of how such committees aggregate preferences, a majority of members will think that voting on discount rates will lead to superior outcomes than attempting to maximize Wτ directly. Next, we take this observation further by investigating when voting dominates direct maximization of Wτ, according to the welfare measure Wτ itself. In addition, denote the median committee member's discount rate by δm. Assume that η ≥ 1. Then, The first part of the proposition shows that utilitarian committees inherit the single-peakedness of their members' preferences on the space of single-agent optimal paths. Although it is not true in general that a weighted average of single-peaked preferences is single-peaked, this does hold for our model. Since utilitarian committees' preferences over discount rates are single-peaked, it is useful to know something about where their ‘bliss point’ discount rate lies. The second part of the proposition shows that it always lies below the discount rate that replicates the utilitarian equilibrium. This is intuitive, as the commitment problem utilitarian committees face always causes them to be more short-termist than they would like to be. The final part of the proposition combines parts 1 and 2 to provide a simple sufficient (but not necessary) condition for voting to dominate the utilitarian equilibrium. This condition applies regardless of the model's primitives, although checking whether Eq. (14) is satisfied requires us to specify these primitives. Logarithmic utility function (i.e., η = 1) When η=1 the equilibrium condition for the dynamic game between utilitarian committees has a closed form solution (Eq. (10)), and members' opinions on welfare in the voting and utilitarian equilibria can be computed analytically. This allows us to obtain a sharp result on when voting will dominate utilitarian aggregation, according to a utilitarian objective function. Assume η = 1. Then voting gives rise to higher utilitarian welfare than direct attempts to optimize Wτ if and only if Fig. 1 plots the set of three element distributions that satisfy these conditions. The figure shows that Eq. (18a) is almost equivalent to requiring that the distribution of discount rates have positive non-parametric skewness, i.e. ⟨δ⟩ > δm. The vast majority of distributions that satisfy Eq. (18a) have this property. However, there is a small set of distributions that satisfy Eq. (18a), but have ⟨δ⟩ < δm, indicated by region C in the figure. The figure also demonstrates that the condition (17a), satisfied in region B of the figure, is sufficient but by no means necessary for voting to dominate the equilibrium implemented by utilitarian committees. The approximate positive skewness condition needed for voting to dominate the utilitarian equilibrium in this example conforms to intuition. When the distribution of discount rates exhibits positive skewness there is a long tail of large discount rates above the median. These discount rates have a disproportionate influence on the equilibrium implemented by utilitarian committees. If a committee with aggregation weights (Eq. (16)) could control consumption for all time, it would choose a consumption path that satisfies impatient members in the short run, and patient members in the long run. Short-run consumption choices are thus always dominated by the concerns of impatient committee members. However, because the current committee cannot bind the hands of future committees, its consumption choice is in effect always short- termist, and thus impatient members' preferences exert a large influence on it. Since this is true for every committee, the equilibrium consumption path implemented by a sequence of utilitarian committees will be biased towards more impatient members. This is reflected in the fact that the equilibrium implemented in this case is observationally equivalent to the optimal plan of an agent with discount rate ⟨δ⟩. It is well known that the arithmetic mean is sensitive to the large ‘outliers' that exist when the distribution of δ is positively skewed. By contrast, the voting equilibrium is robust to the presence of these large outliers, and is thus not subject to the same distortions. More general iso-elastic utility functions (i.e., η > 1) When η > 1, the equilibrium condition (10) must be solved numerically. In order to do this, we must specify the distribution of opinions on the PRSTP. We will use a distribution of discount rate prescriptions elicited from economists who are experts in public project evaluation (Drupp et al., forthcoming). This distribution is illustrated in Fig. 2. In addition, throughout this subsection we assume that the utilitarian committee aggregates preferences using the welfare weights in Eq. (16), i.e. yi ∝ δi. Given these assumptions, standard numerical methods can be used to solve Eq. (10), and the welfare committees achieve under voting and utilitarian aggregation can be computed as functions of the parameters r,η. The results in Fig. 3(a) and (b) are difficult to explain intuitively, as we need to know how both equilibria, and the way they are evaluated, vary with r and η. Since we do not have an analytic expression for the consumption path implemented by utilitarian committees, this is a difficult task. Ultimately, welfare comparisons for η > 1 depend on the empirical details, including the distribution of discount rates. Nevertheless, Fig. 3(b) shows that in our calibration of the model there is a large region of empirically plausible parameter values where voting yields better utilitarian outcomes than the equilibrium utilitarian committees would choose for themselves.","In this paper, we have contrasted two of the most natural preference aggregation methods that committees of decision-makers might employ when attempting to resolve normative disagreements about social time preferences: utilitarianism, and majoritarian voting. While in normal circumstances voting cannot hope to compete with direct optimization of a utilitarian objective function, the time inconsistency of utilitarian committees' collective preferences leads them to implement inefficient consumption plans. We have shown that this can cause voting to yield better outcomes, according to utilitarian committees' own objectives. Indeed, our simple empirical analysis using an elicited distribution of expert opinions on the pure rate of social time preference suggests that this occurs in empirically plausible cases. An interesting feature of model is that, regardless of how utilitarian committees aggregate preferences, a majority of decision- makers will believe society to be better off if committees vote on consumption than if they attempt to maximize their utilitarian welfare measure directly. There is thus a sense in which majoritarian voting is ‘self-stable’ with respect to utilitarian aggregation. The concept of self-stability was introduced by Barbera and Jackson (2004) in the context of their study of constitutions. A voting rule X for ‘ordinary business' (e.g. deciding on discount rates) is self-stable relative to an alternative voting rule Y if when society votes on whether to change the voting rule from X to Y, and uses the rule X to adjudicate this vote, it chooses to stick to X.9 With a little modification, we can adapt this concept to our analysis. If majority voting is also used to decide whether to use majoritarian voting or utilitarianism to aggregate members' preferences, voting will always be selected. By contrast, if the choice between the two aggregation methods is made based on comparisons of the equilibrium utilitarian welfare they achieve, we have shown that there are plausible circumstances under which voting dominates utilitarianism. In these circumstances, majoritarian voting is self-stable, but maximizing utilitarian objectives is not. It is clear from the expression above that the PRSTP δ is a critical input to rt — small changes in its value can have a very large impact on the evaluation of public projects with long-term consequences (see e.g. Heal and Millner, 2014). In practice, governments revise their choices of the social discount at semi-regular intervals (Gollier and Hammitt, 2014), and debates about the appropriate value of the PRSTP are invariably part of this process. While the processes that are used to resolve ethical disagreements about the PRSTP are currently ad hoc and rather opaque, our work takes a more systematic approach to the aggregation of viewpoints on social impatience. Our conclusion is that voting on the PRSTP is likely to have advantages over utilitarian aggregation in practice. Our simple empirical analysis suggests that a consensus value of δ ≈ 0.5%/yr could emerge from such a vote. This value is considerably smaller than that advocated by e.g. Nordhaus (2008) in his analysis of climate change policy (he favours 1.5%/yr), but larger than the value of zero advocated by e.g. Stern (2007) and Gollier (2012) based on their personal ethical views. The main limitation of our analysis is the assumption that committee members share a common utility function. This assumption is made only for technical reasons. It is possible to extend our results on the voting equilibrium to the case where committee members favour different iso-elastic utility functions.10 However, solving for the equilibrium of the dynamic game between utilitarian committees becomes difficult in this case, as the committee's preferences are no longer iso-elastic, and thus the limit equilibrium is generically non-linear. Similar tractability issues arise for more general (i.e. non-linear) production functions. Nevertheless, the qualitative finding that voting may dominate direct utilitarian aggregation will continue to hold even in these substantially more complex cases."],["The discouragement effect of being the lagging player in multi-stage contests is a well-documented phenomenon. In this study, we utilize data from 2447 Davis Cup matches in team tennis tournaments to test the effect of being behind or ahead on individuals’ performance with and without intermediate prizes. Using several different strategies to disentangle the effect of being ahead in the interim score from the effect of selection, we find the usual discouragement effect. However, the discouragement effect disappears after the introduction of intermediate prizes in the form of ranking points. The lagging favorite had close to a 20-percentage point greater probability of winning compared to matches without such a prize. We show that this result is not driven by the selection of better players into tournaments with intermediate prizes. As predicted by previous theoretical studies, our empirical findings suggest that intermediate prizes may mitigate or even eliminate the ahead–behind effects that arise in multi-stage contests. --------------------------------------------------------------------------------","One of the fundamental relationships in the economic environment in general and in tournament settings in particular is the relationship between incentives and performance. It has been well-documented that higher stakes enhance the performance of higher ability agents (Rosen, 1986; Ehrenberg and Bognanno, 1990; Lazear, 2000; González-Díaz et al., 2012; Jetter and Walker, 2015). Another important feature that is frequently found in multi-stage tournaments is ahead–behind asymmetry, where one contestant has an advantage over the other by having a better previous performance. Such situations may occur in R&D contests (Harris and Vickers, 1987), political campaigns (Klumpp and Polborn, 2006), job promotions (Tsoulouhas et al., 2007), and sports competitions (Malueg and Yates, 2010). This ahead–behind asymmetry creates a discouragement effect, according to which a lagging player has fewer incentives to exert costly efforts and therefore is more likely to lose in the following stages.1 There is also a psychological explanation according to which ahead–behind asymmetry creates additional psychological pressure on the lagging player, which in turn harms his/her performance and reduces his/her probability of winning (Apesteguia and Palacios- Huerta, 2010; Genakos and Pagliero, 2012; Palacios-Huerta, 2014; Genakos et al., 2015; González-Díaz and Palacios-Huerta, 2016). The combination between incentives and ahead–behind asymmetry was studied theoretically by Konrad and Kovenock (2009). They showed that intermediate prizes in multi-stage contests might mitigate the discouragement effect on the lagging player. The intuition behind their result is that a lagging player has more incentive to exert effort in every stage, because the player is competing for an additional prize that can be achieved regardless of the interim gap between the players. In a more recent theoretical study, Fu et al. (2015) investigated multi-stage contests, where individuals from two teams compete in pairwise battles. In their model, a team that wins the majority of battles receives a team prize and, additionally, the winner of each pairwise battle receives an individual prize. The authors established the so-called strategic neutrality, according to which the existence of an individual prize eliminates any ahead–behind effect and the probability of winning in every single battle depends only on the players’ innate ability, not on the outcome of the past battles. In general, studying the performance of individuals within a team framework is an important economic and managerial task because in most professions, teamwork is the rule rather than the exception. For example, a recent report by the European Foundation for the Improvement of Living and Working Conditions (Eurofound, 2014) holds that, in 31 out of 37 sectors, teamwork prevails in over 50% of activities. However, studying the performance of individuals in non-experimental contests between teams is not a trivial task because reality rarely creates situations that allow a clear view of the contribution of individuals to a team's output. Therefore, the empirical literature is scarce and based mostly on laboratory experiments.2 A notable exception is the orange grove field experiment conducted by Erev et al. (1993), where the authors found that inter-group competition produced a significantly higher output than in the case where subjects were paid according to their individual output or when they received an equal share of the group's total output. In this paper, we are motivated by the scant empirical evidence from non-experimental settings on the performance of individuals within a team framework, in general, and on the interactive role of incentives and ahead–behind asymmetry in particular. Therefore, the aim of this paper is to test empirically the effect of ahead–behind asymmetry on individuals’ performance in multi-stage contests between teams with and without intermediate prizes using data from tournaments among highly competitive and extensively trained professionals. To that end, we utilized data from tennis matches in the Davis Cup tournament, which is the premier international team event in men's tennis. Each tie between two nations consists of five separate pairwise matches. A team that wins three matches wins the tie.3 Therefore, by construction, before the second and fourth matches of the tie, one of the teams should have more wins than the other. This structure allows us to study the performance of the lagging/leading players.4 More importantly, a change of tournament rules in 2009 makes it feasible to study the effect of intermediate prizes. According to this change, between 2009 and 2015, a player who won a single match in the World Group received individual ranking points. In other years and groups, there were no individual prizes for winning a single match. Utilizing data from professional sports where contestants have strong incentives to win has several advantages. First, it eliminates any possible skepticism about applying behavioral insights obtained in a laboratory to non-experimental settings (Hart, 2005). Second, sports contests involve high-stake decisions that are familiar to the agents. Third, it provides a unique opportunity to observe and measure performance as a function of variables such as heterogeneity in abilities and prizes. Fourth, at each point in time, the contestants have complete information about the interim score and the status of the tournament. Indeed, as Kahn (2000) argues, sports data are very unique in that they embody a large amount of detailed information that can be used for research purposes.5 Since being ahead or behind in the interim score is not determined randomly (for example, home teams or stronger tennis nations have a greater probability of being ahead in the interim score), we use several different strategies to disentangle the effect of leading/lagging from the effect of selection. First, we estimate the average treatment effect of leading/lagging by using the distance-weighted radius matching approach with bias adjustments suggested by Lechner et al. (2011) that has been shown to have superior finite sample properties relative to a broad range of propensity score-based estimators (Huber et al., 2013). We also use Oster's (2019) recently proposed bias-adjusted estimator. Based on the analysis of 2447 matches from 966 international ties, we find a significant ahead–behind influence on players’ performance, which is mostly pronounced in match 4, which is likely because of the unique schedule of the tie. More specifically, we find that the favorite (higher ranked player) has about a 10-percentage point greater probability of winning a match if his team is leading. However, the main contribution of this paper is that we have a unique opportunity to study the performance of players in tournaments with and without intermediate prizes. As already mentioned, in 2009, the Association of Tennis Professionals (ATP) decided to assign ranking points to the winner of a single match in the Davis Cup. These points are taken into account in determining the World Ranking list. Based on this list, players enter the most prestigious tournaments with the possibility of earning large monetary prizes.6 Investigating matches in the World Group with and without intermediate prizes (ranking points, in our case), we find that before the decision to assign ranking points for winning a single match, a favorite from the leading team was more likely to win in match 4 than the favorite from the lagging team. However, from 2009 to 2015, the gap between the probabilities of the lagging and leading favorites’ winning disappeared. We also show that this result is not driven by the selection of better players into the tournament after the change. Our findings suggest that the introduction of intermediate prizes mitigates and may even eliminates the ahead–behind effects that arise in multi-stage contests. The remainder of the paper is organized as follows. Section 2 describes the Davis Cup setting. The data and descriptive results are detailed in Section 3. Section 4 describes the estimation strategy. In Section 5 we present the evidence about the ahead–behind effect. Section 6 reports the effect of intermediate prize. Finally, in Section 7 we offer concluding remarks.","The Davis Cup is an international men's tennis team competition played annually between teams from participating countries. The tournament is structured into five hierarchical levels: World Group, Group 1, Group 2, Group 3, and Group 4. The World Group is the top competition level, comprised of 16 participating nations. Nations that are not part of the World Group compete in one of the lower four groups. Teams in World Group, Group 1 and Group 2 compete in elimination tournaments according to which the winning team advances to the next round and the losing team is eliminated. Groups 3 and 4 use a round-robin structure according to which teams play against each other in pairwise ties.7 A tie signifies a competition round between two competing countries. In the World Group, for example, the 16 nations play eight pairwise ties in the first round (the round of Last 16). The eight winners of this round compete in four Quarterfinal ties. The four winners play two Semifinal ties. Finally, the two winners play the Final tie. Teams from World Group to Group 2 that lose in the first round face the possibility of being relegated to a lower group for next year's tournament. Promotion or relegation in the World Group is decided in Play-off rounds played between losers of the first round in the World Group and winners of Group 1. To be promoted from Group 2 to Group 1, a team needs to win in three different rounds. A team that loses in three different rounds in Groups 1 and 2 is relegated to a lower group. Each tie in the World Group, Group 1 and Group 2 consists of five rubbers (matches), namely, four singles matches and one doubles match. Each team consists of several players who are seeded according to their individual World Rankings. On the first day of each tie, two matches are played between the first seeded player of one team and the second seeded of another team. The schedule of the first day is determined randomly. The doubles match is always scheduled as match number 3, which takes place on the second day of the tie. On the third day, the two top seeded players from each team always compete against each other in match 4 and two second seeded players from each team always compete in match 5. The first team that wins three rubbers wins the tie and progresses to the next round to play a tie against another team. If the tie has not already been decided in favor of one team (no team won three rubbers), then the remaining rubbers are termed live rubbers, which are played in the form of best-of-five sets. Additionally, all dead rubbers are played in the form of best-of-three sets.8 Finally, between 2009 and 2015 a player who won a single rubber in the World Group received ranking points as long as the rubber was defined as a live rubber. In other years and groups there were no individual prizes for winning a single rubber. Data ~~~~ As already stated, since there is a difference between the round-robin and elimination formats, our dataset consists of Davis Cup matches in the World Group, Group 1 and Group 2 that use the latter format. In addition, we consider only matches between individuals and do not use matches between doubles because, in most cases, players do not specialize in doubles and play these types of matches only occasionally. The data were collected from several websites (see Appendix A for a list of all sources). All Davis Cup matches played between 2003 and 2015 are present in the datasets. For every match, information is available regarding the names of the players, their previous head-to-head victories and losses against the opponent, and each player's 52-week ranking prior to the beginning of each match. The ranking is used as a measure of the players’ abilities and is calculated and updated weekly by taking into account all of the player's results in professional tournaments over the previous 52 weeks. Apart from individual level data, information on the location, group type, year, and tournament round for each tie is also available. In all, the dataset consists of 4206 Davis Cup matches. However, we consider only live rubbers (i.e., matches that are still crucial in deciding which team wins the tie) because dead rubbers are in the form of best-of-three sets and usually substitute players compete in these matches. Therefore, 1198 dead rubber matches were eliminated. In addition, another 561 matches lacked information regarding the current ranking of one of the players, or were not played to completion, and therefore were eliminated as well.9 Dropping all of these matches leaves 966 Davis Cup ties, consisting of 2447 matches. Variables ~~~~~~~~~ For each match, we first define the higher ranked player as the favorite and the lower ranked one as the underdog. Then, we estimate the probability that the favorite will win the match. Accordingly, we assign the dependent variable a value of one if the favorite player won and zero otherwise. It is important to note that a favorite is lagging if the interim score of the tie before the respective match is 0:1 or 1:2 in favor of the opponent's team. A favorite is leading if the interim score of the tie before the respective match is 1:0 or 2:1 in favor of his team. Therefore, to estimate the effect of being ahead/behind in the score, we coded a dummy variable that equals one if the favorite is lagging before the respective match and zero otherwise. Similarly, we coded a dummy variable that equals one if the favorite is leading before the respective match and zero otherwise. We also control for the home advantage, which was found to play a significant role in professional tennis (Koning, 2011). Thus, the variable indicating that the favorite has a home advantage receives the value of one if the favorite competes at home and zero otherwise. In addition, we include dummies for each round and type of group categories. Finally, since starting from 2009 a single win in a live rubber of the World Group guaranteed ranking points, we coded a dummy variable that equals one if the match was in the World Group after 2009 and zero otherwise. The descriptive statistics of our dataset are presented in Table 1. It indicates that, on average, the favorite wins in 68.1% of cases if his team is lagging. It also shows that if the favorite's team is leading, his probability of winning is 80.4%. Using a 95% confidence interval, Fig. 1 shows that the favorite's share of wins when his team is leading (1:0 or 2:1) is significantly higher than when the interim score is a draw (0:0 or 2:2) or when the favorite's team is lagging (0:1 or 1:2). However, Table 1 also indicates that if the favorite is leading, he also has more of a home advantage, better head-to head performance, and a lower ranking index, which is associated with greater relative ability. Thus, in order to obtain the causal effect of being ahead/behind, we will use several estimation strategies that control for selection into treatment (leading/lagging). We discuss these strategies in the following section.","Studying whether being ahead or behind before a Davis Cup match gives an advantage to the favorite is a challenging task. A naïve approach of correlating a dummy variable for leading/lagging with the probability of winning a match will yield biased and inconsistent estimates because the status of being ahead or behind is not determined at random. Rather, as mentioned earlier, being ahead is a function of features specific to tennis such as home advantage, previous head-to-head meetings, and the difference in abilities between the other members of the teams. Furthermore, isolating an exogenous source of being ahead/behind in the score by using an instrumental variable approach seems unfeasible because any factor that might be associated with being ahead/behind is also likely to affect the probability of winning the match. In the absence of a valid instrument, we will use several alternative strategies to control for the endogeneity of leading or lagging in Davis Cup matches. Radius matching estimator ~~~~~~~~~~~~~~~~~~~~~~~~~ Our main analysis is based on the radius-matching-on-the-propensity score estimator with bias adjustment (Lechner et al., 2011). Not only was it found to be very competitive among a range of propensity score related estimators, but also a later paper by Huber et al. (2013) actually demonstrated its superior finite sample and robustness properties in a large-scale empirical Monte Carlo study.10 The main idea of this estimator is to compare treated and non-treated observations within a specific radius. The first step consists of distance-weighted radius matching on the propensity score. In contrast to standard matching algorithms where controls within the radius obtain the same weight independent of their location, in the radius matching approach, controls within the radius are weighted proportionally to the inverse of their distance to the respective treated observations to which they are matched. The second step uses the weights obtained from this matching process in a weighted linear or non-linear regression in order to remove biases due to mismatches. Because this approach uses all comparison observations within a predefined distance around the propensity score, it allows for greater precision than fixed nearest neighbor matching in regions in which many similar comparison observations are available. Radius matching analysis ~~~~~~~~~~~~~~~~~~~~~~~~ First, we conducted the analysis for the full dataset. As already discussed, there is a selection into being ahead/behind. Although the purpose of the propensity score estimation is only a technical one, namely, to allow the easy purging of the results from the effects of selection, it is nevertheless interesting to see which variables drive selection. In Table 2 we report the results for the propensity score estimation. We use two different specifications. In the first, we control for differences in rankings, previous head-to- head results and home advantage. In the second specification, we also control for specific features of the ties, such as the round of the tournament, the group, the year and whether the match is a World Group match before or after 2009. We can see that many variables are associated with being ahead/behind. This finding is not surprising because we would expect home players to be more likely to win and players from stronger countries have, on average, better teammates. In Columns 1 and 2 of Table 3 we present the results for the radius-matching estimator where Panel A and Panel B report the average effects for lagging and leading, respectively. The clustered standard errors at the tie level are presented in parentheses. The results show that the effect of lagging is negative and significant. It reduces the favorite's probability of winning by about 5 percentage points. The effect of being ahead on the favorite's probability of winning is between 3.7 and 5.4 percentage points and also significant. This finding is in line with the ahead–behind asymmetry that has been found in soccer (Apesteguia and Palacios-Huerta, 2010; Palacios-Huerta, 2014) and chess (González-Díaz and Palacios-Huerta, 2016). It is important to note that our results do not contradict those of Berger and Pope (2011) who found that being slightly behind (one point) at half-time has a positive effect on the probability of winning in basketball. However, being far behind is less likely to have a positive effect. Since in the Davis Cup there are only five matches, being one match behind is a much more significant lag than being one point behind in basketball, where teams score about 100 points per match. Therefore, we interpret lagging by one match in the Davis Cup as being further rather than slightly behind. Oster's bias-adjusted treatment effect ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ It is important to note that the radius-matching-on-the-propensity score estimator is very flexible, without strict assumptions about the functional form. Nevertheless, as a robustness check we also use Oster's bias-adjusted estimator, which relies on a strict functional form, making it much less flexible than the semi-parametric matching estimator. In order to conduct the treatment effect of leading/lagging, in Columns 3 and 4 of Table 3, we present the results of the linear probability model (LPM) without and with the full set of controls respectively where standard errors clustered at the tie level appear in parentheses. Not surprisingly, we can see that the size of the coefficients of Favorite is lagging and Favorite is leading are much higher in the uncontrolled specification presented in Column 3 than in the specification with the full set of controls presented in Column 4. In Column 5 we present the bias-adjusted treatment effect of lagging/leading. The standard errors obtained from the bootstrap are presented in parentheses. The results show that the estimated causal effect is closer to zero, but still significant. When a favorite is lagging, he is 3.7 percentage points less likely to win with a significance level of 5.3%. The positive effect of being ahead on the favorite's probability of winning is 3.6 percentage points with a significance level of 3.9%.11 Ahead–behind asymmetry in matches 2 and 4 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In this subsection our aim is to investigate only matches where the score is asymmetric, namely matches 2 and 4, where by construction, one team is leading and the other is lagging before the beginning of the match. Fig. 2(a) and (b) shows that on average there is a much larger gap between the share of wins if a favorite is leading in match 4 compared to match 2. Our empirical analysis presented in Table 4 demonstrates that being ahead has a significant and positive effect on the probability of winning in match 4. We find no significant effect of being ahead in match 2.12 This result is in line with several explanations. First, as already mentioned, matches 1 and 2 are played between the first seeded player of one team and the second seeded of another team, whereas match 4 involves the two top seeded players from each team. Thus, it is intuitive that a lagging favorite competes against a weaker opponent in match 2 compared to match 4. Therefore, the discouragement effect is less likely to appear in match 2, because being down 1:0 in a Davis Cup meeting is almost an expected event. However, it is possible that the favorite player loses in the first match, which can have a different effect on the performance of the players in the second match. Therefore, in Appendix D we present the results of match 2 for cases in which a favorite won and lost in the first match. We find that if a favorite won in the first match, then being ahead in the second match has a positive effect on the probability of winning. However, this effect is not significant at conventional levels. We also observe a negative coefficient if the favorite lost in the first match. Nevertheless, the result is far from being significant. There are some additional explanations for the difference in results between matches 2 and 4. For example, it is possible that a lagging player has much more to lose in terms of a team prize in match 4 compared to match 2 because if a lagging player loses in match 4, his team loses the entire tie. Therefore, such a situation may provoke choking under pressure of the lagging player and, as a result, harm his performance and reduce the likelihood of his winning.13 Finally, it is also possible that the leading player values his win more than the lagging player in match 4 compared to match 2, which may also result in a difference in the probabilities of winning. This difference in valuations between the matches may be driven by simple egocentric motives. For example, the winner of the match that determines the tie gets more glory. Although we cannot observe all of the possible prizes the players receive from winning a single match, in the next sub-section we use a unique opportunity to study the effect of the ahead–behind asymmetry in settings with and without intermediate prizes. Introduction of ranking points in the World Group in 2009 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In this section, we take advantage of the change in the rules introduced by the ATP. Up to 2009, players did not receive any ranking points for a single win. However, between 2009 and 2015, the winner of a live rubber of the World Group received ranking points. These points are taken into account in determining the World Ranking list. This list is very important because it determines the entries to the most important tournaments with the largest monetary rewards. In addition, players with a higher number of points may benefit from a better draw, because in the first rounds they play against weaker players. To put this decision into perspective, a win in a main tournament of the Davis Cup was worth 40–75 ranking points, depending on the round. This means that two wins in Davis Cup matches were worth more than two wins in the first two rounds of Grand Slam tournaments (55 points), which are the most prestigious tennis tournaments.14 Theoretical framework ~~~~~~~~~~~~~~~~~~~~~ As discussed, theoretically, the intermediate prizes (ranking points) play a very important role in multi-stage contests. According to Konrad and Kovenock (2009), the introduction of positive intermediate prizes may increase the lagging player's probability of winning. Moreover, according to Fu et al. (2015), if there is an intermediate prize, which is common to both players, the interim score of a tie has no effect on the players’ probability of winning in a single rubber. This probability depends only on the players’ innate abilities. In Section 2 we noted that the last two matches involve the top seeded players from each team in match 4, and the two second seeded players from each team in match 5. In contrast, the first two matches always involve the first seeded player of one team and the second seeded of another team. Intuitively, match 4 is more symmetric than match 2. Indeed, in our dataset, the mean DiffRank value in match 2 equals −1.79, which is significantly lower than the mean DiffRank value in match 4, which equals −1.59 (two- sample mean comparison p-val=0.009). This result illustrates that match 4 is significantly more symmetric than match 2. Therefore, in our model between two symmetric players, we will present only the case of the last two matches, because they bear a closer resemblance to the empirical settings. Empirical evidence ~~~~~~~~~~~~~~~~~~ Although our empirical settings do not fully resemble our theoretical model or the theoretical settings of Fu et al. (2015) and Konrad and Kovenock (2009), we still wish to test whether starting from 2009, the probability of winning is affected by the state of the contest (whether a player is leading or lagging). Based on the World Group ties, Fig. 4(a) shows that the gap between the probabilities of winning when one is leading in match 4 compared to being behind was, on average, 28 percentage points before the change in the rules in 2009.15 However, as Fig. 4(b) illustrates, this gap declined dramatically to only 8 percentage points after 2009. Selection issue One possible concern, however, is that the introduction of the ranking points may attract better players. Therefore, the greater probability of the lagging favorite's winning might be attributed to selection rather than to intermediate prizes. To obviate this concern and show that the players’ rankings are not differently distributed before and after 2009, we use the following two-step procedure. First, we partitioned the data into two parts, where one set contains the World Groups’ match 4 before 2009 and the other the World Groups’ match 4 after 2009. In Table 5 we report the average value of the log2 of the ranking of the favorite, the underdog and the differences between them on the match level, separately for each period. In parentheses we present their standard deviations. Column 1 refers to the matches before 2009, while Column 2 refers to the matches after 2009. We can see that the log2 of the players’ rankings is even somewhat higher after 2009, implying the rankings of those with less ability. Then, we run a set of univariate regressions of each of the variables presented in Table 5 on a dummy variable indicating whether the specific observation was before 2009. The coefficient of this dummy variable and its standard error are presented in Column 3. The results show that none of these players’ characteristics differ significantly between the two periods. This finding indicates that the players’ log2 of rankings and their differences do not differ before and after 2009. Therefore, we can conclude that selection into the sample is not a concern. Radius matching analysis In Table 6, we present the effects of leading in match 4 on the probability of the favorite player's winning before and after 2009. The radius-matching estimator presented in Column 1 implies that there is a significant and positive effect of being ahead before 2009, which is much smaller and not significant after 2009. Finally, we test whether the intermediate prizes increase the probability of the lagging player's winning, as indicated in Fig. 4(a) and (b). In Table 7, we compare the probabilities of the lagging favorites’ winning in match 4 before and after 2009. In total, we have 45 such cases before 2009 and 43 after. The results of the radius-matching estimator presented in Column 1 imply that the effect of the intermediate prizes on the probability of the lagging favorite's winning is 18.4 percentage points with a significance level of 6%. It is important to note that similar to the results presented in Table 5, in the case with the lagging favorite as well, none of the characteristics significantly differs between the two periods, before and after 2009. In fact, as Appendix F indicates, none of the variables has a p-value lower than 0.31. This result serves as additional evidence that selection into the sample is not a concern.16 Oster's bias-adjusted treatment effect As previously, we use Oster's bias-adjusted estimator as a robustness check. In Columns 2–3 of Table 6 we present the LPM's coefficients of Favorite is leading in match 4 in the World Group before and after 2009, where robust standard errors appear in parentheses. Given the very small number of observations, the R-squared is very sensitive to the inclusion of any additional variable. Therefore, we follow Oster (2019) who also offers an adjusted procedure for evaluating the bias- adjusted treatment effect when some variables are considered part of the identification strategy and thus appear in both the controlled and uncontrolled regressions. The idea is to assess the amount of selection on the observables conditional on including these variables in the estimation. Because some of the variables are significant in the propensity score estimation presented in Appendix E, it is worthwhile assessing the amount of selection conditional on these variables being included in the estimation as part of our identification strategy. Therefore, in Columns 2 and 3 of Panel A in Table 6, which represents the matches before 2009, the DiffRank and Year2008 are included in both the controlled and uncontrolled regressions. Similarly, since DiffRank is significant in the propensity score estimation (Column 2 in Appendix E), in Panel B of Table 6, we include it in both the controlled and uncontrolled regressions (Columns 2 and 3). We can see that the coefficients are significant before 2009 and even somewhat higher when including the controls (Columns 2 and 3 of Panel A in Table 6). Therefore, by definition, Oster's bias-adjusted coefficient has to be even farther from zero, as we can see in Column 4 of Panel A in Table 6. However, the bootstrapping procedure in the dataset that includes 89 observations only increases the standard errors and the p-val to 0.125. Nevertheless, the most important result in this case is that Oster's bias-adjusted coefficient becomes even larger compared to the LPM. When testing the effect of leading after 2009, we can see that according to all of the estimators (Columns 2–4 of Panel B in Table 6), the effect of leading in match 4 is much closer to zero and highly insignificant. Finally, in order to conduct Oster's bias-adjusted treatment effect of playing after 2009 on the lagging player's probability of winning, we use the results of the LPM with and without the full set of controls as presented in Columns 2 and 3 of Table 7, respectively. We can see that the effect of the post-2009 period is not sensitive to the inclusion of the controls. Its size is about 20 percentage points with a significance level of 3.6% and 4.7% in Columns 2 and 3, respectively. Not surprisingly, Oster's bias-adjusted coefficient, presented in Column 4, is almost the same as in the LPM, with a significance level of 5.6%. Taken together, our results suggest that the introduction of incentives for winning a single match is likely to affect the performance of players. Although our empirical settings do not fully match the theoretical settings of Fu et al. (2015) and Konrad and Kovenock (2009), our finding that the probability of winning is not affected by the state of the contest when an intermediate prize is introduced is in line with these theoretical predictions. In general, our empirical results emphasize the importance of the strategic allocation of efforts in multi-stage contests that is well known in the theoretical literature. Although we cannot rule out the possibility of some other psychological effects, our findings suggest that the introduction of intermediate prizes may mitigate or even eliminate the ahead–behind effects that arise in multi-stage contests.17","In this paper, we used tournaments among highly competitive and extensively trained professionals to test the effect of ahead–behind asymmetry on individuals’ performance in multi-stage contests between teams with and without intermediate prizes. As in previous studies, we find that being ahead provides players with a greater probability of success. However, the main contribution of this paper is that it empirically shows that intermediate prizes eliminate the usual ahead–behind effect that may arise from psychological as well as from strategic considerations. Our results, obtained from contests between high-profile professionals, underscore the role of strategic motives in individual performance. This is especially important on the team level, because teamwork is probably the most prevalent form of economic activity. Our findings suggest that non- monetary incentives alone, such as a team's pride, are probably not enough to maximize an individual's output. This result may be of great importance in situations that involve a choice between individual and social benefits. Nevertheless, it is important to note that we did not investigate the effect of team-based incentives on individual performance. Therefore, it will be interesting to study whether the introduction of team-based incentives instead of individual-based incentives will also lead to improved performance. Furthermore, individual incentives may improve the utility of other teammates of the lagging team because such incentives do not affect the winning probabilities of leading favorites, who are likely to win the decisive match regardless of the incentives. It is rather the lagging favorite who benefits from additional individual rewards. In addition, his teammates benefit from the greater probability of winning the entire contest. This may explain why companies in difficulty are ready to pay extra salaries to high-profile workers, who are able to stabilize the firm's cash flows or profits. However, other workers who do not receive an additional individual reward may also benefit from the increased stability of their workplace. Finally, it is important to note that despite the fact that our findings are in line with the common ahead–behind effects and with previous theoretical studies on the effect of intermediate prizes, the results of this paper were obtained from the sport of tennis, which is mostly an individual sport. Playing in teams in the Davis Cup is not the usual competitive format for most players. Therefore, it is possible that our results would be different in other settings, where individuals are used to performing in teams. It is also possible that those who are not used to large monetary rewards would also behave differently. Therefore, we call for additional empirical research to test the interactive effect of intermediate prizes and ahead–behind asymmetry in various other environments."],["Economists are increasingly using experiments to study and measure discrimination between groups. In a meta-analysis containing 441 results from 77 studies, we find groups significantly discriminate against each other in roughly a third of cases. Discrimination varies depending upon the type of group identity being studied: it is stronger when identity is artificially induced in the laboratory than when the subject pool is divided by ethnicity or nationality, and higher still when participants are split into socially or geographically distinct groups. In gender discrimination experiments, there is significant favouritism towards the opposite gender. There is evidence for both taste-based and statistical discrimination; tastes drive the general pattern of discrimination against out-groups, but statistical beliefs are found to affect discrimination in specific instances. Relative to all other decision-making contexts, discrimination is much stronger when participants are asked to allocate payoffs between passive in-group and out-group members. Students and non-students appear to discriminate equally. We discuss possible interpretations and implications of our findings. --------------------------------------------------------------------------------","Meta-analysis – a commonplace technique in medical science, psychology and, to a growing extent, economics – holds advantages over literature review in terms of objectivity and analytical rigour. In recent years, the experimental economics literature appears to have reached a critical mass at which researchers are finding meta-analyses useful.1 The benefit of these works is that, by aggregating data across a large number of experiments and exploiting natural between-study design variation, they pinpoint behavioural regularities and the variables that modify them more precisely than could be done through qualitative review. We run a meta-analysis on the body of studies investigating discrimination in lab and lab-in-the-field experiments, a sub-literature which has certainly reached the necessary critical mass for such a venture. Economists׳ interest in discrimination has been strong ever since Becker (2010), and with the growth of experimental economics in the last two decades, experiments have emerged as a popular complement to survey-based econometric studies. These experiments create a controlled environment and therefore allow much cleaner measurements of discrimination than the analysis of naturally-occurring data, avoiding such problems as omitted variable bias and reverse causality. Furthermore, by testing for a very fundamental and general form of discrimination – simply, whether subjects treat others differently depending on which group those others belong to – experimental economists can produce findings of interest not only to their own discipline but also across the social sciences. Also, through the use of incentives, experiments hold a key advantage over questionnaire-based measures of discrimination, in that they elicit revealed rather than reported discrimination. Psychologists had already been studying discrimination in the lab for decades, and experimental economists have drawn on their knowledge, particularly regarding the minimal group paradigm. This technique was first introduced by Tajfel et al. (1971) and has spawned a huge body of experiments wherein group identity is artificially induced in the laboratory. This is often done by, in a preliminary phase of an experiment, asking subjects to state their preference for one artist over another, or to randomly draw a colour. The experimenter then splits the subject pool into groups according to their art preference, or the colour they have drawn, and makes it known to participants that the division is based on these differences. Subsequent stages of such experiments involve interaction tasks between the groups and find discrimination surprisingly (at least to the early researchers) often. To study discrimination, experimental economists set up games such as the dictator game, the trust game or the prisoner’s dilemma, and invite a subject pool segregated along the lines of a particular identity-based characteristic (or else generate this segregation with artificial groups). They make subjects aware of the group affiliation of those they interact with, and then measure how their behaviour varies according to whether individuals they are interacting with share their identity (are in- group) or do not (are out-group). The number of economics experiments of this type has grown rapidly since the turn of the century and now encompasses substantial diversity across several dimensions. Even after omitting many papers which investigate discrimination but do not meet our inclusion criteria devised to ensure a consistent approach (see Section 2), we are left with a dataset consisting of 441 experimental results (significant and null) from 77 studies – more data than most of the other experimental economics meta-analyses have had. In order to aid the progression of this literature, it is worth taking stock of what has been found to date, particularly as casual inspection reveals non-uniformity in the results; the strength of discrimination found against out-groups varies considerably, and some experiments even find discrimination in the opposite direction, i.e. against the in-group. The aim of this meta- analysis is both to yield broad insights on discrimination and to inform the designers of future experiments testing for it. We first investigate the average strength of discrimination across the literature. We then inquire how it tends to vary according to specific experimental characteristics. In particular, we are interested in whether the strength of discrimination depends on the type of identity being investigated. Comparing the level of discrimination between artificial (i.e. minimal) groups and various types of natural groups (such as those based on ethnicity, nationality, religion, gender and social/geographical affiliation) is particularly interesting. One might expect ‘minimal’ groups to yield minimal levels of discrimination. However, it is also conceivable that artificial identity inducement confers an experimenter demand effect in favour of discrimination, or that the experimental priming of sensitive natural identities reduces subjects’ desire to discriminate owing to a preference not to engage in socially unacceptable behaviour. Evidence for these possibilities, in the form of relatively strong discrimination in artificial group experiments, could have implications for the external validity of certain experiments. A further interesting question is whether the strength of discrimination varies according to the type of decision subjects are asked to make. This has implications in terms of the real-world circumstances in which discrimination can be most expected to appear and for the generalisability of findings. We further ask whether experiments with students reveal greater or lesser discrimination than those with non- students. This is also important for the external validity of findings, and is a question worth pursuing as some studies (e.g. Bellemare and Kröger, 2007; Anderson et al., 2013) have found students are not entirely representative of wider populations in economics experiments. This meta-analysis also aims to shed light on the motivations behind discrimination. Some experiments have been designed specifically to distinguish between taste-based discrimination and statistical discrimination – the two models that continue to dominate the theoretical literature in economics. The taste-based model, proposed by Becker (2010), entails individuals gaining direct utility from the act of discriminating against out-groups. Meanwhile, according to theories of statistical discrimination – beginning with Arrow (1972) – individuals aim to maximise their own payoffs given their beliefs and expectations about others׳ characteristics and behaviour, and discrimination occurs when those beliefs and expectations vary depending on the group to which the others belong. Understanding the relative importance of these two motivations will improve the focus of future research and the design of policies aimed at combating discrimination. Finally, we include a subsection on experiments investigating gender discrimination. Gender is unique amongst the identity types in having the same two groups in each experiment. It is therefore simple to make a clean comparison between male-to-female discrimination and female-to-male discrimination. In summary, the meta-analysis presented below aims to address the following questions: (1) What is the general pattern of discrimination across the literature? (2) How does the level of discrimination vary according to the type of identity groups are based upon? (3) How does the level of discrimination depend upon the decision-making context? (4) Do students discriminate any more or less than non-students? (5) Does the experimental literature provide more support for taste-based or statistical theories of discrimination? (6) In gender experiments, how does male-to-female discrimination compare with female-to-male discrimination? Our main results, presented in Section 3, are as follows. (1) We find a moderate tendency towards discrimination against the out-group, with a majority of null results across the literature. (2) The strength of discrimination against the out-group does vary according to the type of group identity subjects are divided by. It is greater when identity is artificially instilled in a subject pool than when it is divided by nationality or ethnicity – minimal groups, it seems, are not so minimal after all. Discrimination is even stronger, though, when participants are divided into socially or geographically distinct groups. (3) The extent of discrimination against the out-group also depends on the role participants are given in an experiment: when subjects are asked to allocate payoffs between inactive players belonging to the in-group and out-group, it is stronger than in any other decision-making context. (4) Students do not appear to be differently inclined towards discrimination than non-students. (5) We find evidence in support of both taste- based and statistical discrimination. Tastes appear to drive the general tendency for discrimination against the out-group, but individual studies have found beliefs to affect discrimination. (6) In gender discrimination experiments the tendency for discrimination against the out-group is reversed, as subjects demonstrate slight but significant favouritism towards the opposite gender. Discriminatory behaviour in these experiments does not differ significantly between males and females. We discuss possible interpretations of these results in depth in Section 4. We are aware of only one other meta-study attempting to analyse the experimental discrimination literature – Balliet et al. (2014),2 who take 214 estimates of discrimination from 78 studies. There is little overlap between our samples; Balliet et al. take studies from across the social sciences but their search and inclusion criteria result in most of the experimental economics literature on discrimination not being included (26 of our studies – around a third – feature in Balliet et al.׳s sample). They exclude decision-making contexts which we consider, such as being the second mover in a sequential game or a third-party allocator. They also exclude interactions between gender groups. The present study and that of Balliet et al. can be viewed as complements. Through focusing only on economic experiments, we enhance comparability and eliminate some studies using methodological elements that may not be acceptable to some social scientists. Our focus on the economic theories of taste-based and statistical discrimination differentiates our study from Balliet et al., who investigate psychological theories of discrimination. Throughout our analysis we compare our results to theirs. Their paper finds a similar overall tendency for discrimination to what we do. They find the extent of discrimination not to differ significantly between settings of natural and artificial identity, but do not split natural identity into subcategories as we do. The clearest difference in results between the two studies is that Balliet et al. find discrimination is stronger by decision-makers who move simultaneously than by first movers in sequential exchanges, while we do not find it significantly differs between these settings.","We chose to restrict our study to the experimental economics literature. Almost all of the economics experiments have been conducted in the last 15 years and can reasonably be expected to have followed comparable procedures, which is important in a meta-analysis. We define an economics paper as follows: it must either have been published in an economics journal or have as at least one of its authors a person trained in economics or a business-related discipline, or who has at least once held a position in an economics or business-related department. Furthermore, we exclude economics papers which, it is clear to the reader, exhibit a breach of standard experimental economics practice – most notably, deception. For inclusion, an experiment must involve interaction between individuals whose decisions determine real material payoffs for participating players. In other words, it must be incentivised. A serious pitfall meta-analyses can face is publication bias, also named the ‘file drawer problem’. Because null results are less likely to be published than significant ones, a meta-analysis risks including a disproportionately low number of studies finding small or no effects (Rosenthal, 1979; Rothstein et al. 2006). This can lead to an overestimation of average effect sizes. It can also, if null results are particularly unlikely to be published when combined with certain other features of a study, result in the meta-analysis overestimating the relationship between strong effects and these features; in our case, for instance, if null results in trust games were never published but null results in other games sometimes were, we would be in danger of estimating a spuriously strong relationship between trust games and significant results. To minimise such bias, a good meta-analysis should conduct the most thorough literature search possible in order to find all applicable studies, whether published or not. Our approach was threefold. In late 2013, we conducted RePEc searches for the keywords, ‘Discrimination experiment’, ‘Identity experiment’, ‘Ingroup experiment’ and ‘Outgroup experiment’, and carefully sifted through the output for candidate studies. We then followed the references and citations of all papers identified as relevant. Finally, we checked our list of included studies against that of Balliet et al. (2014); this step added one study (Spiegelman, 2012).3 One feature of the literature we meta- analyse is that studies tend to include various different treatments, and therefore report multiple results. This may act as a further curb on publication bias – insignificant findings make their way into papers alongside more interesting significant results (indeed, it turns out the majority of results in our dataset are null).4 Previous meta- analyses in experimental economics such as Engel (2011) and Johnson and Mislin (2011), which focus on a single game type, are able to use the average behaviour of subjects (amount sent in the dictator or trust game) as a continuous dependent variable, with one observation and an associated standard error for each treatment. In our case, we are pooling across different game types and therefore need a way of transforming the data to make meaningful comparisons between these settings. Our variable of interest is the difference between decision-makers׳ behaviour towards their in-group and their out-group, whilst all other aspects of the experimental design are held constant – in essence, the discrimination effect size. There is typically one observation per every two treatments (one in-group and one out-group treatment) for each type of player active in the given game. The exception is when a decision-maker interacts with both the in-group and the out- group in the same treatment (either by making one decision which simultaneously affects both, or by playing in the same role twice), in which case a within-treatment measure of discrimination is available.5 The ideal approach would be to record an effect size for each comparison, and we attempt to do this. Consistent with Balliet et al. (2014), the measure we use is Hedges׳ unbiased d: the mean difference in behaviour towards the in- group and the out-group, divided by the pooled standard deviation, with a minor correction for sample size (Hedges and Olkin, 1985). However, a substantial number of studies do not report sufficient data for us to calculate effect sizes. This is particularly the case with null results, as when a difference is not significant authors are less likely to report the test statistic from which an effect size could be derived. We sent data requests to the authors of all papers for which we could not construct the measure using information provided in the paper. After receiving data from 22 of the 36 sets of contacted authors, we ended up with effect sizes on 364 of our 441 data-points. We therefore also employ a binary dependent variable, recording simply whether, for each comparison, behaviour significantly favours the in-group over the out-group at the 5% level.6 The effect size is the inferior dependent variable in that it restricts the sample and may lead to greater under-representation of null results; but the superior one in terms of information content. For simplicity, we define ‘discrimination’ as discrimination against the out-group, and ‘out-group favouritism’ as discrimination against the in-group, and will use these terms hereafter. Unlike some, we make no distinction between nepotism and discrimination; any result of favouritism towards one group relative to a second can equivalently be interpreted as discrimination against the second group. We therefore conceptualise ‘discrimination’ (against the out-group) as something which can be measured on a continuum with positive and negative values. When discussing average effect sizes, we will describe a relatively low value as indicating ‘lower’ or ‘weaker’ discrimination, even if it is driven by highly negative effect sizes (i.e. even if it is driven by instances of strong discrimination against the in-group). For an observation to meet our inclusion criteria, there must be an in-group and out-group, clearly defined on the basis of categorisation by a discrete identity-relevant variable, such as ethnicity, gender or – as with artificial groups – the preference for a particular artist or the colour randomly drawn. There must be controlled interaction within and between the groups, and decision- makers must be aware that they are interacting with individuals belonging to their in- group or out-group. We only consider an in-group to be appropriately defined as such if every one of its members shares the same categorisation as the decision-maker on the basis of the relevant variable. For an out-group to be appropriately so-defined, every member must take a different categorisation from the decision-maker. It is not required that all members of an out-group take the same categorisation as each other. For instance, Guillen and Ji (2011) use as their two groups Australian and non-Australian. In this case, for an Australian decision-maker the Australians are the in-group and the non-Australians the out-group, but for a non-Australian the other non-Australians should not count as their in-group. We then only record the observed behaviour of the appropriately defined group, the Australians in this example. Occasionally, we are forced to make a subjective decision on what can reasonably be considered a group. For example, from Chen et al. (2011), which splits its US-based sample into white and Asian students, we record the behaviour of the white ‘group’ but not that of the Asians, as we believe that in American society white people can appropriately be defined as comprising a shared ethnicity, whilst those of Asian descent comprise a mixture of ethnicities.7 Papers such as Falk and Zehnder (2007) which do not have clear groups but measure each subject’s position on a scale of social distance, based on a continuous variable, are not included. If an experimental design splits the sample up into more than two separate groups, on the basis of a single identity-relevant variable, we record separately how each group treats each other group relative to its own. If such a paper reports that Group A does not significantly discriminate against Group B or Group C but does significantly discriminate against Groups B and C combined, we record two results of no discrimination rather than one result of discrimination; and in the main text of this paper we report our results using this approach. We do this because, although Groups B and C combined could represent a single out-group as defined above, the experiment was set up to treat them as separate out- groups. Similarly, we do not include the reported results of statistical tests run on data pooling two or more treatment pairs. These are grey areas but we have re-run our main regression results for the binary dependent variables in the case of treating every result reported in our sample as an observation: this adds 16 extra data-points and does not qualitatively change our findings. Sufficient data must be reported for it to be clear whether there is significant discrimination in each pair of treatments (or, when applicable, single treatment); if we cannot work out whether there is discrimination in one or more treatment pair, the whole paper is omitted from the study. This is because papers are less likely to report the results of statistical tests finding no discrimination, and if we failed to include a given study’s non-results our analysis would overestimate the likelihood of this particular design finding discrimination. For similar reasons, if an experiment employs a cross-cutting design, dividing its subject pool by multiple identity types, it must report whether there is discrimination on the basis of each category. For example, an experiment which segregates the subjects by both gender and ethnicity must report, for each applicable treatment pair, whether each ethnic group discriminates against each other ethnic group or not, and also whether each gender discriminates against the other or not. Otherwise, we omit the study. Experimenters using artificial groups generally conduct tests on pooled data; rather than reporting whether Group A discriminates against Group B and vice versa, they report whether individuals across the sample pool discriminate against out-group members. This makes sense because there is no obvious reason to doubt the relationship between two artificial groups is completely symmetrical. As such, we use pooled discrimination observations for artificial group experiments. Using similar reasoning, we also admit pooled discrimination observations for experiments dividing subjects by their real-world social groups. The pooling of certain types of data might lead to an increased chance of finding discrimination in certain experiments, which is one reason why we use the size of the sample from which the result is derived as a control variable in our regression analysis. We limit our analysis to lab and lab-in-the-field experiments; we do not include pure field experiments, in which subjects do not know they are participants in a study. We therefore do not include the large body of field experiments in which applications are sent to employers, landlords or others to test for discrimination in markets (correspondence studies). Analytical methods ~~~~~~~~~~~~~~~~~~ Listed in the next subsection are descriptions of the independent variables we include in our regressions. Our basic model contains role and identity type dummies, and some controls. Because our samples are not large and most variables are dummies, we regard linear probability models (LPMs) with errors corrected for heteroskedasticity as the best specifications when employing the binary dependent variables. However, we also run as robustness checks logit models, which we report in Appendix C, Table C.2. In some cases the logits drop observations, which is a major disadvantage. Their results, however, are qualitatively similar to the LPMs. When using binary dependent variables, we treat each study within the meta-analysis as providing a cluster of observations. When analysing the continuous dependent variable, we first use standard random effects meta-analysis procedures to determine average effect sizes for our full sample and for the subsamples based on identity type. These are simply aggregate estimates for the level of discrimination in the relevant subsample; they do not control for independent variables. The procedure takes into account that each observation has an associated standard error. It weights each observation by the inverse of this standard error, thus attaching more importance to results from larger samples and with smaller standard deviations. It then follows an unweighting process, the extent of which depends upon the heterogeneity in effect sizes. The more heterogeneity there is across effect sizes, the more equal will be the weights attached to observations with small or large standard errors (Harbord and Higgins., 2008).8 We then apply random effects meta-regressions, which allow the inclusion of independent variables in the analysis. These models follow the same processes of weighting and unweighting observations as described in the previous paragraph, but are otherwise standard linear regressions. Whereas with the binary dependent variable we must approach discrimination and out-group favouritism separately, the meta-regression analyses both simultaneously, since the effect sizes can be positive or negative. This can be one reason why the results of the meta-regressions may differ from those of the linear probability regressions. Another can be the reduction in sample – therefore, when the results of the meta-regressions do not match those of the LPM regressions on discrimination, we present the LPMs re-run on the reduced effect-size sample, in order to determine whether the disparity is due to the change in sample or the change in analytical approach. Role type dummies We include role type dummy variables to pursue the question of how different decision-making contexts affect the extent of discrimination. The games used in this literature feature either multilateral or unilateral decision-making. When decision-making is multilateral, the outcome of the game is determined by more than one player׳s actions. From such situations, we identify three different role types: First Mover (140 observations), where one׳s move does not finish the game; Second Mover (119 observations), where one determines the final payoffs in response to the co-player(s)’ actions; and Simultaneous Mover (66 observations), where one makes the last move of the game at the same time as one׳s co-player(s). When decision-making is unilateral, the final outcome of the game is determined by one player. From these situations, we identify a further three role types. First, there is the Dictator (67 observations), who allocates payoffs between another player and his- or herself. Next we have third-party allocators (Allocator, 30). These are players who must divide a pie between two or more passive players (who, in these experiments, are members of different groups), but whose own payoff does not depend on this decision. Finally, there are players tasked with selecting a partner (from a choice of in-group and out-group participants) with whom to play a subsequent game. We label this role Partner Chooser (19 observations).9 Identity type dummies A second set of dummy variables records which type of group identity a given experimental sample has been divided according to. We consider identity to have been artificially induced if researchers split subjects into groups that, prior to the experiment, did not exist – in the sense of group members sharing characteristics that are not also shared by members of other groups in the study – and the subjects are aware they have been split into these groups.10 49 studies in the meta-analysis investigate natural identity, 32 artificially generate it, while the remaining four contain both natural and artificial treatments. We have 272 observations for natural identity types and 169 for artificial. We subdivide the natural observations into six specific categories of natural identity. First, we have 82 observations from 13 studies in which subjects are divided by Nationality. Next, nine studies investigate Ethnicity-based identity, adding 63 observations. A further seven studies generate 32 observations on Gender identity. 21 more observations are provided by five studies in which the subjects are split by Religion. 13 studies use a rather different approach, dividing the subject pool into groups based on real-world social and/or geographical identity. This is done in a variety of ways: for instance, using villages (Dugar and Shahriar, 2010), colleges within universities (Banuri et al., 2012) or friendship groups (Brandts and Sola, 2010). However, all such designs share the common feature that each decision-maker has a clearly distinct social and/or geographical in-group – group identity here is defined with reference to the relative frequency with which one interacts with in- and out-group members in ordinary life. The 57 observations generated by these experiments are coded under the variable Soc/Geo Groupings. The remaining 17 results, from four papers, deal with other types of natural identity, which cannot appropriately be fitted into the above categories. These observations relate to political identity (Abbink and Harris, 2012), disability (Gneezy et al., 2012), caste (Hoff et al., 2011) and whether farmers are private or members of cooperatives (Hopfensitz and Miquel-Florensa, 2013). We pool them under the composite variable Natural Other. 11,12 Other variables In our regressions we include as a dummy variable (Students) whether each observation derives from a sample consisting predominantly of students or non- students. Even if not explicitly stated, we assume experiments run at universities have at most a very small number of non-student participants. Likewise, while we accept experiments in the field may include a few student subjects, their proportion is likely to be low (unless otherwise stated). As another control, we include the size of the active decision-making sample from which a given result is derived (Sample Size). What is the general pattern of discrimination across the literature? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In total, as shown in Fig. 1, there are 144 results indicating significant discrimination (32.65%), 28 indicating significant out-group favouritism (6.35%), and 269 indicating no significant discrimination or out-group favouritism (61.00%). 57 of our 77 studies record at least one result of discrimination, while only 15 record any results of out-group favouritism. 10 studies separately record results of discrimination and out-group favouritism. The general tendency, then, leans towards insignificant results, although only 15 studies consist entirely of nulls. For the sub-sample where we are able to generate effect sizes (364 of 441 observations), the random effects meta-analysis finds an overall effect size of 0.256 (95% confidence range: 0.209–0.304). This can be interpreted as, on average, subjects׳ discriminating against out-groups by about a quarter of a standard deviation. This is not significantly different from Balliet et al. (2014), who find an overall effect size of 0.32 (95% confidence range: 0.27–0.38). Fig. 1 also displays point estimates for aggregate effect sizes, conditional on the type of result found for each observation. Observations finding significant discrimination have an average effect size of 0.67, those yielding null results have an average effect size of 0.11, and those finding significant out-group favouritism have an average effect size of −0.51; this confirms that the strength of the effect size tends to be closely related to the type of result found for a given observation. : In general, there is limited discrimination against the out-group. How does the level of discrimination vary according to the type of identity groups are based upon? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 1 displays a breakdown of our sample׳s observations by identity category, and the results of random effects meta-analyses run on these sub-samples. For most categories the tendency is towards null results. Only for Soc/Geo Groupings – which yields no results of out-group favouritism – are observations of discrimination more likely than insignificant results, and this is also the identity type with the highest average effect size. The category for which there is least discrimination and most out-group favouritism is gender; the average effect size for this sub-sample is negative. Table 2a extends the analysis of Table 1 through the use of regressions. LPMa is a linear probability model with the dependent variable discrimination against the out-group (equal to 1 if discrimination is found, 0 otherwise). Metareg is a meta-regression with the dependent variable the discrimination effect size. In both models artificial identity and the dictator game are the benchmark categories.13 In these regressions we test whether our identity-type variables still yield significantly different levels of discrimination after controlling for other factors. Table 2b presents the results of linear restriction tests run on the sets of dummy variables featuring in the models. In both the linear probability models and meta-regression, the identity category linked to the strongest discrimination is social and geographical groupings. In Metareg it yields significantly higher discrimination, at the 1% level, than any of the other identity categories. In LPMa it does the same, except that the differences with Artifical and Natural Other are only significant at the 5% and 10% level respectively. The identity category linked to the weakest discrimination is gender. Both the linear probability model and the meta-regression indicate weaker discrimination between genders than between artificial groups, significant at the 1% level. The meta-regression also finds gender discrimination to be weaker than ethnic and national discrimination (at the 1% level), and religious discrimination (at the 10% level). However, LPMa does not find these differences to be significant.14 In LPMa the coefficients on the ethnic and national identity types are significantly negative at the 1% level, strongly indicating that discrimination is less likely to be observed when subject pools are split along these lines than on the basis of artificial identities. According to Metareg, however, ethnic and national identity experiments are only linked to significantly lower discrimination (i.e. less positive effect sizes) than artificial group experiments at the 10% level. Given the inconsistency, Table 2a also reports LPMb, a linear probability model run on the reduced sample for which effect size calculation is possible. This helps to distinguish whether the losses of significance when moving from LPMa to Metareg are due to the reduction in sample or the change in measurement technique. For the comparison of national and artificial identity, the loss of significance appears to be due to the change in sample, as in LPMb the coefficient is also insignificant. The same cannot be said for Ethnicity, however, as the linear probability model on the reduced sample continues to report significantly less discrimination between ethnic than artificial groups at the 1% level. Doubt, therefore, is cast over the robustness of our finding on ethnicity – although the coefficient’s sign is at least weakly significant.15,16 The strength of discrimination depends upon the type of group identity under investigation. It is stronger when identity is artificially induced in the laboratory than when the subject pool is divided by ethnicity or nationality, and higher still when participants are split into socially or geographically distinct groups. How does the level of discrimination depend upon the decision-making context? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Inspection of the coefficients on role type dummies in LPMa and Metareg (Table 2a) reveals discrimination is significantly stronger when the decision-maker is a third-party allocator than when he or she is a dictator (the omitted category). Linear restriction tests (Table 2b) also show the third-party allocator role is more likely to be associated with discrimination than all the other role types, with the difference always significant at the 1% level under both models. The size of the Allocator coefficients in the meta- regression (1.077) is worth noting – it indicates that discrimination in games of this type tends to be very large indeed, with on average more than one standard deviation between subjects׳ treatment of in- and out-groups. The other role types do not carry significantly different effects from one another. This is at odds with Kiyonari and Yamagishi (2004) and Balliet et al. (2014), who find discrimination to be stronger by simultaneous movers than first movers (Balliet et al. do not investigate second movers). In an attempt to discern why our result differs from that of Balliet et al., we re-ran our analysis keeping only the observations included in their study. We found there was still no significant difference between First Mover and Simultaneous Mover (the remaining sample on which to run this regression was small; however, we also compared the aggregate effect sizes for each category and found they are very similar). This suggests the significance of the finding in Balliet et al. is driven by studies outside our dataset, i.e. outside the economics literature.17 Third-party allocators discriminate more than decision-makers in all other roles. Do students discriminate any more or less than non-students? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Most decision-makers in our analysis were students. Only 101 observations, from 22 studies, are produced by in-groups not comprised (at least in their near-entirely) of university students. 31.8% of the observations for students return discrimination, while 6.8% find out-group favouritism and 61.5% are null; for non-students 35.6% find discrimination, 5.0% yield out-group favouritism and 59.4% are null. The coefficient on Students is not significant in any of our regressions. That experiments with students do not generate significantly different levels of discrimination than those with non-students is an interesting non-result which suggests that, in this literature, working with student samples will not generate a biased perception of the extent and magnitude of discrimination by the wider population.18 : Discrimination does not significantly differ between students and non-students. Does the experimental literature provide more support for taste-based or statistical theories of discrimination? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ For 262 (59.4%) of our observations, as a result of the experimental design any discrimination must be taste-based, as it cannot be statistical. Statistical discrimination cannot occur when a player is making the only or last move in a game, unless the game is to be repeated, or possibly if the move is made simultaneously with others. Discrimination by trust game returners, for example, can only be taste-based, because opponents then have no control over the final outcome and beliefs about their type are therefore irrelevant.19 All observations under the Dictator and Allocator categories preclude the possibility of statistical discrimination, as do all except seven (due to the game being repeated) in the Second Mover category. All observations under the First Mover category permit the possibility of statistical discrimination, as do most in the Partner Chooser category and around a third in the Simultaneous Mover category.20 In Table 3, we run a linear probability regression on discrimination and a meta-regression on the discrimination effect size, with role types re-coded into two types: one, Taste+Statistical, where there is scope for both taste-based and statistical discrimination, the other (the omitted category) where there is scope only for taste-based discrimination. Note that in this literature any game-role contains scope for taste-based discrimination. The coefficient on Taste+Statistical is positive but insignificant in both the linear probability regression (p=0.87) and the meta-regression (p=0.18). This indicates there is no significant difference in the likelihood of observing discrimination, or in the predicted effect size, when scope for statistical discrimination is added.21 This would suggest taste-based discrimination is an important driver of behaviour in these experiments and statistical discrimination is not, but we probe further by analysing the results of individual experiments. Where there is scope for statistically-motivated discrimination, by design for 66.5% of these observations it is not possible to disentangle its effects from taste-based motivations. To be able to do so, an experiment must either use belief elicitation or include a control game in which behaviour can only be taste-based – the most common case of this is adding a dictator game to extricate taste-based from statistical discrimination by trust game senders.22 In the 60 cases that it is possible to distinguish between discriminatory motives, the authors find significant statistical discrimination to occur in 13 cases (10 times against the out-group and three times in favour of it). Within the same sample, for given beliefs or behaviour in a game with a belief-based component, they find significant taste-based deviations from own-payoff-maximisation in 26 cases (16 times against the out-group and 10 times in favour of it). In 26 cases neither statistical nor taste-based discrimination is found at the 5% level. We list all significant findings of taste-based and statistical discrimination from experiments designed to distinguish between the two in Appendix A, Table A.3. Although the sample is small, tastes are found to affect behaviour more often than statistical beliefs. It seems, however, that beliefs do play some role in determining discriminatory behaviour in economics experiments. We conjecture that the insignificant regression results in Table 3 may be due to the fact that beliefs can either increase or reduce discrimination. This would be because individuals have favourable beliefs about the cooperativeness of out-groups, or because unfavourable beliefs about the out-group׳s cooperativeness can in some cases actually lead to statistical out-group favouritism. That is, depending on the game setting, self-serving optimal behaviour can either become more or less generous in response to the perception that one׳s partner is relatively uncooperative. In ultimatum games, for instance, if proposers expect out-group responders to treat them less favourably than in-group responders do, the self-serving optimum is to send them relatively kind offers. This is in contrast to how first mover behaviour would work in trust games, say, where a self-serving sender will send relatively low investments to an out-group responder if it expects to be treated unfavourably by them.24,25 : There is evidence for both taste-based and statistical discrimination. Tastes appear to drive the general tendency for discrimination against the out-group, but individual studies have found beliefs to affect discrimination. In gender experiments, how does male-to-female discrimination compare with female-to-male discrimination? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ An immediately obvious finding is that gender acts very differently from other identity types. It is the only identity category which is more likely to be associated with a bias against the in-group than against the out-group, with eight results of the former and three of the latter out of a total 32 observations. On the reduced sample, the random effects meta-analysis finds an overall discrimination effect size of −0.177 (95% confidence range: −0.301 to −0.053) for gender experiments, representing significant out- group favouritism. There is obvious intuition why gender is different from the other identity categories: it is the only case in which the effects of sexual attraction – towards the out-group more than the in-group, for most subjects – and ׳chivalry’ (Eckel and Grossman, 2001) can be expected. Every experiment on gender in the meta-analysis has a symmetrical male-female design, meaning that for every estimate of discrimination by men against women there is an identical treatment measuring discrimination by women against men. This allows a very clean comparison of these two behaviours across the sample. The only three significant results in our dataset of one gender discriminating against the other are female decision-makers discriminating against males, while six of the eight significant results of one gender favouring the other are male decision-makers favouring females. However, the calculated overall effect size for female decision-makers is actually slightly more negative than for males: −0.181 (95% confidence range: −0.35 to −0.013) for females and −0.173 (95% confidence range: −0.369 to 0.024) for males, although the difference is far from significant. Note that while the effect size indicates females significantly favour males at the 5% level, the equivalent effect for male decision-makers is only significant at the 10% level. : There is significant out-group favouritism in gender experiments. Females significantly favour males; males favour females but the effect is only weakly significant.","A leading result of this paper is that discrimination in economics experiments varies by the type of identity groups are based upon. It is very strong when groups are socially or geographically distinct, and is relatively weak when they are based on ethnicity or nationality. Notably, it tends to be relatively strong in experiments using artificially- induced group identities – so it can confidently be stated that minimal groups do not produce the minimal level of discrimination. At first glance, this seems surprising. It might be that artificial group manipulations are stronger priming instruments than natural identity experiments tend to use – after all, these dedicate an entire preliminary phase of the experiment to inducing feelings of identity, which will remain at the front of subjects׳ minds when they are then offered the chance to discriminate. This explanation is arguably supported by the evidence of Robbins and Krueger (2005), whose meta-analysis of psychology experiments shows subjects exhibit stronger in-group projection – that is, they perceive in-group members to be particularly similar to them, relative to out-group members – when identities are artificial than when they are natural. On the other hand, we do not find that team-building exercises, which are designed specifically to strengthen artificially-induced identity and would seem to amplify priming, have a significant effect on the level of discrimination (this is consistent with the findings of Chen and Li (2009)). Conversely, it could be argued that, for the populations studied in the literature, membership of particular ethnic and national groups does not actually instil strong identity, so that even such trivial identities as can be artificially induced have a greater effect. There is evidence that the process of globalisation has weakened national and ethnic parochialism (Buchan et al., 2009), and in recent decades youth identity in the West and increasingly elsewhere has come to define itself to a large extent upon individuals’ belonging to subcultures based on fashion and music tastes – preferences drawn from choice sets which are not, indeed, so different from the apparently arbitrary minimal group painting dichotomy. However, it would seem highly complacent to draw the conclusion from our results that racism and xenophobia are not big problems in many societies. Another explanation may be that subjects in ethnic and national identity experiments are shying away from displaying ׳politically incorrect’26 behaviour, given that racism and xenophobia are taboo in most societies today. While the link between social acceptability and discrimination has not been well explored, the prejudice literature has yielded relevant findings: that expressions of prejudice correlate with perceptions towards the social acceptability of such prejudice (e.g. Crandall et al., 2002), and furthermore that this correlation is at least partly the result of norm- compliance (e.g. Blanchard et al., 1994). It seems unlikely that discriminating on the basis of a stated preference for Klee׳s paintings over Kandinsky׳s carries any taboo similar to ethnic or national discrimination. Indeed, some subjects may regard an artificial group situation as a game in which they belong to one of the teams, wherein the social norm actively encourages favouritism of one’s own group – the sheer strangeness of the setting may even lead subjects to perceive a demand for discrimination on the part of the experimenter (see e.g. Zizzo (2010)). Concerns about social acceptability could explain also why the Soc/Geo Groupings category produces significantly higher discrimination than other types of natural identity. Of course, it would not be surprising if relational and geographic proximity led to a stronger sense of belonging than shared ethnicity, religion or nationality, but bear in mind too that there is arguably no taboo against favouring friends over strangers.27 If it were shown that discrimination in economics experiments is indeed limited by concerns about social acceptability, it might cast doubt over the external applicability of such studies׳ findings. It is possible that if participants guess an experiment is about a type of discrimination which is taboo, it will systematically generate a lower effect than if the subjects were unaware of its purpose. On the other hand, the very same concerns about social acceptability might also limit certain types of discrimination outside the lab. It is noteworthy that gender is the identity category producing the weakest discrimination: in fact, here the meta-analysis finds a significant amount of out-group favouritism. However, gender discrimination clearly persists in the outside world. It may be that economics experiments do not find it because they poorly reflect the conditions under which it survives beyond the lab – in particular, in the job market. It would be interesting to see more experiments designed to directly compare the effects of different types of group identity. This meta-analysis includes just four. Dugar and Shahriar (2009), Li et al. (2011) and Goette et al. (2012a) all compare discrimination between social/geographical groups and artificial groups, while Abbink and Harris (2012) use artificial groups and political groups (which fall under the Natural Other category). The results of all four studies are consistent with ours – discrimination is always lower with artificial identity. However, direct comparisons between artificial group and ethnic or national discrimination are lacking, and it would be very illuminating to see whether such studies support – and if so, whether they can explain – the findings of this meta-analysis. What implications does our research have for future experiments on discrimination? First, using artificially induced identities as a control against which to pit the results of natural identity treatments may not be recommendable, as the artificial group manipulation appears not so much to capture the minimal level of discrimination that must result from priming any type of identity in a laboratory as to in fact often go beyond that. Regarding role type, we find discrimination by third-party allocators is much stronger than by participants in any other game setting. If social acceptability does indeed limit discrimination, this is a counterintuitive result, as the allocator role essentially invites subjects to overtly and consciously favour one group over another and therefore seems to be the one that most obviously telegraphs the purpose of this type of experiment. One possibility is that the role carries an experimenter demand effect – whereby subjects feel they are encouraged to discriminate – or even an action bias effect, if the equal split feels like a default non- move. Another relevant factor may be that the third-party allocator is unique amongst our role types in the decision-maker׳s payoff being entirely disconnected from the extent to which they discriminate. In any case, experimenters should bear in mind that because they are more likely to identify significant discrimination when they employ the allocator role, they should be less confident that the same groups will discriminate against each other in different contexts. We find the strength of discrimination does not significantly differ between student and non-student subject pools. This suggests – unlike in the context of social preferences (e.g. Bellemare and Kroger, 2007; Anderson et al., 2013) – student subjects are not a generally unrepresentative sample for questions relating to discrimination. However, we do not exclude the possibility that they are unrepresentative in specific instances, or within particular societies. There is scope for more experimental research investigating taste-based and statistical discrimination. We show both are relevant, and the two types manifest themselves to different extents in different contexts. However, relatively few experiments have been designed to distinguish between taste-based and statistical discrimination, and more could be known about the mechanisms underlying them. As a final observation, there is a great deal of variation in the findings of the experimental economics discrimination literature. Our analysis can explain some of it, but our LPM regressions typically have R2 statistics below 0.2, and the meta- regressions׳ Adjusted R2s are rarely above 0.35. As might be expected, discrimination does seem to vary idiosyncratically and is not easy to predict. The results of natural identity experiments do not seem very generalizable – they probably reflect more the characteristics of the specific groups under investigation, and the relationships between them, than aspects of the experimental design. Whilst a drawback for some research questions, this also means there is a great deal of scope for future experimental studies aimed at measuring the levels of discrimination within subject pools of specific interest. Furthermore, given the potential concerns we raise about experimenter demand effects and the external validity of lab experiments on discrimination, the important role of field experiments should be emphasised. Subjects in such studies are unaware they are being observed by experimenters and their behaviour can therefore not be influenced by the fact. Field experiments can test the generalisability of lab findings on discrimination."],["This paper examines the impact of a government programme which facilitated the entry of for-profit surgical centres to compete against incumbent National Health Service hospitals in England. We examine the impact of competition from these surgical centres on the efficiency – measured by pre-surgery length of stay for hip and knee replacement patients – and case mix of incumbent public hospitals. We exploit the fact that the government chose the broad locations where these surgical centres (Independent Sector Treatment Centres or ISTCs) would be built based on local patient waiting times – not length of stay or clinical quality – to construct treatment and control groups that are comparable with respect to key outcome variables of interest. Using a difference-in-difference estimation strategy, we find that the government-facilitated entry of surgical centres led to shorter pre-surgery length of stay at nearby public hospitals. However, these new entrants took on healthier patients and left incumbent hospitals treating patients who were sicker. This paper highlights a potential trade-off that policymakers face when they promote competition from private, for-profit firms in markets for the provision of public services. --------------------------------------------------------------------------------","In the 2000s, there was a widespread push in Europe and the United States to increase the role of user choice and provider competition in public services. In general, these pro- market reforms were designed to increase the quality and efficiency of public services like health care and education, which had previously been run through non-market means like performance management (Gaynor and Town, 2011; Propper et al., 2007, 2010). Often, as part of these market-based reforms, policymakers encouraged the entry of private, for- profit firms to compete against public sector providers. These efforts are exemplified by the growing use of charter schools in the United States and private health care providers in publicly funded health systems in Western Europe (Jost et al., 2006; Fryer Jr, 2012). This paper explores how competition generated by the government-facilitated entry of private, for-profit firms affects the performance of incumbent public providers. In particular, we estimate the impact of the entry of a series of private, for-profit surgical centres in the English National Health Service (NHS). Policymakers steered the entry of these surgical centres to areas with high patient waiting times, with the aims of increasing surgical capacity and stimulating competition. We estimate the impact of this private provider entry on the efficiency of incumbent public hospitals, and examine whether it left incumbents with a riskier and more costly mix of patients. Advocates of diversifying the supply of public services providers argue that private, for-profit entrants will innovate and offer higher quality than incumbents, and that entry of private providers will create competitive pressure on public providers to raise their own performance (Le Grand, 2009; Seddon, 2007). We are particularly focused on testing this latter claim: can the entry of private, for-profit surgical centres improve the performance of incumbent public hospitals? Critics of market-based reforms generally cite the many ways that public services, and health care in particular, differ from highly stylised, perfectly competitive markets, and argue that competition will not improve performance (Jones and Mays, 2009; Fotaki et al., 2008). Moreover, it is sometimes argued that, because new entrants are often much smaller than incumbents (in our case, we analyse surgical centres competing against hospitals), they may not have sufficient scale to affect the behaviour of existing providers (Goddard, 2015). A third criticism is that private, for-profit firms may select customers with desirable characteristics (e.g. better students or less risky patients), leaving public providers treating a riskier or costlier group of users (Los Angeles Times Editorial Board, 2016; Bardsley and Dixon, 2011). More generally, it is not clear that governments are well equipped to determine where to locate entrants in such a way as to engineer effective competition. The English NHS provides a unique environment in which to test the effect of private, for-profit provider entry on public service providers' performance, and in so doing to analyse the extent to which governments can ‘create’ competition. In the 2000s, the British government facilitated the entry of Independent Sector Treatment Centres (ISTCs). ISTCs are private, for-profit surgical centres focused on provision of routine, high volume elective (i.e. medically necessary, non-emergency, scheduled in advance) surgical procedures to public (NHS) patients. This policy was part of a wider policy package designed to tackle waiting times within the English NHS, the centrepiece of which was an ambitious set of targets to reduce waiting times for surgery. ISTCs were established to rapidly expand capacity in regions deemed at risk of not meeting these targets (Naylor and Gregory, 2009). As we demonstrate, while the placement of these specialty surgical centres was correlated with local public hospital waiting times during the pre-policy period, their placement was uncorrelated with measures of the efficiency and clinical quality of these incumbents over the same period. This implies that treatment assignment was unrelated to the pre-policy levels of the outcome variables we study. In addition, we demonstrate that public hospitals close to ISTC entrants had nearly identical pre-entry trends to public hospitals unexposed to ISTC entry across a range of performance measures (other than waiting times). We use this observation to motivate a difference-in-difference (DiD) strategy to estimate the causal effect of ISTC entry on outcomes at nearby public hospitals and highlight that our control group serves a good counterfactual for what would have occurred to the treatment group after 2004/5 in the absence of the entry of ISTCs. Measuring efficiency of health care provision is a long-standing challenge because of the absence or poor standard of data on costs and quality. Faced with these problems, researchers have frequently used patient length of stay (LOS) as a proxy for efficiency (Fenn and Davies, 1990; Martin and Smith, 1996; Gaynor et al., 2013) on the grounds that, provided clinical quality can be maintained, shorter LOS implies lower costs for the same outcomes. However, a key difficulty with using LOS to capture efficiency is that it is heavily influenced by patient characteristics – patients in poorer health before surgery will tend to have longer lengths of stay for reasons unrelated to hospital efficiency. In this study, we use an innovative approach to address the influence of patient characteristics on LOS-based efficiency measures by disaggregating LOS into two components: time from admission until surgery (‘pre-surgery LOS’), and time from surgery until discharge (‘post-surgery LOS’). We show that pre-surgery LOS is less affected by patient characteristics than other components of LOS, and use it – or alternatively, the percentage of patients treated on the day of admission – as a proxy for hospital efficiency. In what follows, we show that the entry of private, for-profit specialty surgical centres led to a 16% reduction in pre- surgery LOS at nearby public hospitals – which translates to a 24 percentage point increase in the proportion of patients treated on the day of admission. However, we also find evidence that these entrants engaged in risk selection, leaving nearby public hospitals with a sicker (and therefore costlier) mix of patients. In particular, public hospitals exposed to the entry of private specialty surgical centres experienced an 11.6% deterioration in average patient health status as captured by the Charlson score (defined in Section 4). This increase in patient severity likely led to an increase in post-surgery LOS at incumbent NHS hospitals. Finally, while ISTC entry may have led to reduced case loads at some public hospitals with which they shared a market, we show that our estimated treatment effects are not driven by changes in volume caused by ISTC entry. This paper adds to several literatures. First, it builds on previous work assessing how the entry of private, for-profit firms impacts the performance of incumbent public service providers (Hoxby, 1994; Barro et al., 2006; Cutler et al., 2010; Sass, 2006). In general, researchers have struggled to assess the causal impact of competition from new market entrants (e.g. surgical centres and charter schools) into markets for public services because the entry location of private firms is usually endogenous. We exploit the fact that siting of surgical centres in England was driven by government policy tied to waiting times, not our efficiency measure, and show that the entry of ISTCs raised incumbent hospitals' productivity. Second, it adds to the broader literature assessing the impact of hospital competition (Kessler and McClellan, 2000; Gaynor et al., 2015; Cooper et al., 2011). We illustrate that, in markets where payments are regulated, competition can raise hospitals' efficiency. Moreover, we find that smaller entrants can affect the behaviour of larger incumbents. Third, it adds to the literature analysing whether private, for-profit surgical centres offering public services risk-select against public incumbents (Barro et al., 2006; Winter, 2003; Cram et al., 2005; Street et al., 2010; Zimmer and Guarino, 2013; Bifulco and Reback, 2014). We find that the entry of ISTCs left public hospitals with a riskier mix of patients. To some extent, this was by design: ISTCs in England were focused on treating uncomplicated cases. While the entry of specialist surgical centres focused on routine procedures could in theory represent efficient patient sorting, such an arrangement is likely to leave existing providers treating a sicker patient mix and worse off financially, unless it is accompanied by a reimbursement system that adequately adjusts payments to reflect patient severity. The consensus is that NHS payments were not adequately risk adjusted during the period we investigate (Mason et al., 2008), meaning that NHS hospitals that had an ISTC enter nearby were likely left worse off as a result of being left with a sicker mix of patients. More generally, this paper highlights the trade- offs that policymakers face when considering policies to encourage the entry of for-profit firms to compete with public service providers. Facilitating entry can lead to competition, which can prompt incumbent providers to raise their performance. However, these for-profit entrants may have very different objectives than incumbent providers, and may have a higher propensity to risk-select in order to draw a more advantageous mix of patients. Our work highlights the need for policy-makers to take risk-adjustment of payments seriously when considering policies to promote competition between firms with different objectives and differing abilities to treat complicated cases. The remainder of this paper is structured as follows. Section 2 presents background information on recent NHS reforms, with particular focus on the ISTC programme. Section 3 explores the potential impact of ISTC entry on incumbents' performance. Section 4 presents the data and empirical strategy. Section 5 reports the results, while Section 6 discusses and concludes.","The English NHS, founded in 1948, is funded through general taxation and, with few exceptions, offers health care that is free at the point of use. Patients must register with a single general practice (GP) clinic for the provision of primary care, and GPs act as ‘gatekeepers’ to the secondary care system. For the most part, secondary care in England is organised around large public hospitals. The NHS long struggled with waiting times for elective surgery, which, in some cases, could exceed a year. In 1997, a new Labour government was elected promising quick action to reduce waiting times. However, 1 year later, waiting times had increased (Klein, 2013, p.200).1 Concerns over waiting times became the catalyst for a series of reforms from 2000 onwards, which included rigorous performance management of public hospitals; introduction of patient choice and hospital competition underpinned by prospective reimbursement; and facilitated entry of specialist private surgical centres to compete with larger public hospital incumbents. The new prospective reimbursement system (known as Payment by Results or PbR) was modelled on the Diagnosis-Related Group (DRG) system used in Medicare in the United States (US). Under PbR, hospital reimbursement is tied to activity rather than to annual budgets or block contracts as was previously the case (DH, 2011). In 2000, The Secretary of State released The NHS Plan (Secretary of State for Health, 2000), in which the government committed to cutting maximum waiting times for elective surgery from 18 months to 6 months by the end of 2005 (later reduced to 18 weeks, by 2008) using a series of targets tied to rewards and punishments. There is substantial evidence that the targets and performance management regime was extremely effective at reducing waiting times (Propper et al., 2008, 2010; Besley et al., 2009). As part of its reform programme, in April 2002 the government announced that it was facilitating the entry of a series of privately run surgical centres (ISTCs) to deliver routine, high-volume diagnostic and elective surgical procedures to English NHS patients.2 Like other NHS services, NHS-funded patients could use ISTCs free of charge. Although the NHS had long made use of private providers in England, ISTCs were distinctive in three ways. First, they were created as a deliberate policy of government, as opposed to being a result of decisions by local commissioners of care. Second, they provided services exclusively to NHS patients, as opposed to earlier arrangements in which NHS patients were treated in settings mainly focused on treatment of private patients (Naylor and Gregory, 2009). Third, whereas NHS physicians are in general permitted to also work in private settings, the first wave of ISTCs (which are the focus of this paper) were not allowed to use NHS doctors. This restriction ensured that ISTCs represented genuine new additions to capacity, rather than drawing away physician labour from nearby public hospitals. More than any other factor, it was local waiting times that influenced where the government sought to locate the new private surgical centres (HCHC, 2006). According to government officials, “In October 2002, the Department [of Health] conducted an extensive forward planning exercise, during which all Strategic Health Authorities were asked to identify, in conjunction with their respective Primary Care Trusts, any anticipated gaps in their capacity needed to meet the 2005 waiting times targets. The results of this exercise led to the identification of capacity gaps across the country, particularly in specialties such as cataract removal and orthopaedic procedures, where additional capacity was needed” (Anderson, 2006). Following this planning exercise, in December 2002 the Department of Health invited expressions of interest to run the first Wave of ISTCs. These invitations indicated the broad geographical regions within which ISTCs were to be placed, but left securing a specific site to bidders. Preferred bidders for these schemes were announced from September 2003. There were 27 Wave 1 ISTCs, all of which operated on a for-profit basis.3 Of these, 23 opened in 2005 or 2006 (see Fig. 1), and most operated from a single site, often in newly built premises that were often co- located with an existing NHS hospital. In March 2005, a second Phase of ISTCs was announced, of which nine were eventually implemented, with most opening between 2007 and 2008. These Phase 2 ISTCs were smaller, provided a wider range of services including diagnosis, and were often on the same site as existing private hospitals. Unlike Wave 1 ISTCs, Phase 2 ISTCs were permitted to recruit NHS staff and employ NHS consultants privately. Phase 2 ISTCs were also required to provide NHS training placements (Naylor and Gregory, 2009). Given these very different characteristics of the Phase 2 programme, in this paper we focus exclusively on analysing the impact of Wave 1 ISTCs by excluding NHS hospitals that were potentially exposed to competition from Wave 2 entrants.4 The ISTC programme had a major impact on the market for some elective surgical procedures (Naylor and Gregory, 2009). From 2006, ISTCs accounted for between 5 and 10% of orthopaedic volume nationally. As the ISTC programme's impact was highly geographically differentiated, the share of patients attending ISTCs was much higher in some areas than in others. In some markets where ISTCs entered, they became the only alternative to large incumbents. As one local NHS official noted when a large ISTC opened next to a dominant NHS hospital, “that's the first time… we've ever had any competition” (McLeod et al., 2014, p.15). The hope amongst the policymakers responsible for the ISTC programme was that these entrants would be less inclined to cooperate with NHS providers, and more inclined to compete (Stevens, 2004). The prohibition on ISTCs employing NHS physicians, and the fact that they were privately owned, may also have contributed to an institutional culture at ISTCs that was more receptive to competition than that of public hospitals (Le Grand, 1999). ISTC contracts specified a range of ‘exclusion criteria’ – acceptable grounds for refusing to treat NHS patients – on the basis that ISTCs did not possess the emergency or intensive care units required to treat sicker and more complex patients. Each ISTC had its own list of exclusion criteria, which typically included demographic factors such as age, social factors such as availability of a carer at discharge, and clinical factors such as health status (Mason et al., 2008). In relation to the latter, a particularly important criterion for rejection was the patient's American Society of Anaesthesiologist's (ASA) score – ISTCs were typically able to refuse to treat patients with a score of 3 or more.5 National Joint Registry data from 2010 indicates that, at NHS hospitals, 20% of hip replacement patients and 19% of knee replacement patients were given ASA scores of 3 or 4. The corresponding figures for ISTCs were only 6 and 8%, respectively (NJR, 2011). Critics of the ISTC programme saw these exclusion criteria as particularly problematic because they allowed the new entrants to dump costlier, more complex patients onto the public hospital system (Wallace, 2006; Kmietowicz, 2006).6","This section examines the likely response of public (NHS) hospitals to the entry of private, for-profit surgical centres (ISTCs). In understanding the impact of the ISTC programme, it is important to note that, although public NHS hospitals are run on a not- for-profit basis, they are financially and managerially independent of central government, and during this period had strong incentives to generate a financial surplus, or at least not to make substantial losses. In the early 2000s, the government introduced a system of ‘star rating’ of NHS hospitals, in which financial performance was a major factor (Bevan and Hood, 2006a, 2006b; DH, 2002). Hospitals given a zero-star rating were ‘named and shamed’, and their chief executives were at risk of losing their jobs. Later, high- performing hospitals (those with Foundation Trust status) were given additional freedoms to retain financial surpluses across financial years. Other hospitals were eventually able to achieve Foundation Trust status in part through good financial performance. These factors meant that, during this period, public hospitals had a strong incentive to generate operating surpluses. It has therefore been argued that it is reasonable to think of public hospitals during this period as maximising profits plus some additional term reflecting altruistic valuation of quality and/or quantity (Gaynor et al., 2013). Ultimately, Wave 1 ISTCs differed from public hospitals in three key dimensions. First, they were explicitly for-profit ventures. Second, they were narrowly focused on offering a small range of elective surgical procedures. Third, given their inability to hire NHS staff, their institutional cultures may have differed sharply from those at NHS hospitals. In what follows, keeping in mind these three differences, we present hypotheses about the response of NHS hospitals to the ISTC programme. Efficiency ~~~~~~~~~~ We expect ISTC entry to lead to efficiency improvements at nearby incumbents. As mentioned in Section 1, we measure hospital efficiency using pre-surgery LOS. Prospective reimbursement systems (like PbR in England) pay hospitals on the basis of outputs rather than inputs. This creates incentives for hospitals to reduce marginal costs by shortening patient LOS (Cutler, 1995). Empirical studies of England (Farrar et al., 2009), the United States (Feder et al., 1987; Guterman and Dobson, 1986; Feinglass and Holloway, 1991; Kahn et al., 1990), Israel (Shmueli et al., 2002) and Italy (Louis et al., 1999) provide evidence that prospective reimbursement leads to shorter LOS. While prospective reimbursement systems provide incentives for all hospitals to reduce patient LOS, these incentives will likely be particularly sharp in more competitive markets. Hospitals located in less competitive markets likely have limited scope to expand their activity because they are constrained by the relative inelasticity of clinical demand within their catchment areas. By contrast, hospitals in more competitive markets have greater opportunity to expand activity by capturing market share from other hospitals. To create capacity for such expansion, in health care systems with prospective reimbursement, hospitals in more competitive markets are likely to take stronger action to reduce patient LOS, so that they can treat additional patients. Consistent with this hypothesis, studies of the 2006 patient choice reforms in the English NHS found that hospitals located in more competitive markets decreased their LOS by larger amounts than hospitals in less competitive markets (Cooper et al., 2012; Gaynor et al., 2013). In light of this theoretical prediction and empirical evidence, we hypothesise that incumbent hospitals exposed to entry by an ISTC will have reduced patient LOS over and above any secular decreases in LOS resulting from the introduction of prospective reimbursement. We therefore identify the effect of ISTC entry on the efficiency of nearby public hospitals using a DiD estimator in which the treatment effect equals the change in efficiency at ISTC-exposed public hospitals minus the change in efficiency at unexposed public hospitals. During the 2000s, the government announced that performing elective surgery on the day of a patient's admission was a key measure of hospital performance, and highlighted that ISTCs would be particularly effective at this. The NHS Institute for Innovation and Improvement (2006, 2008a, 2008b) identified surgery on day of admission as one of the six characteristics of high-performing orthopaedic surgical facilities and argued (2006, p.20) that public hospitals would have to respond to competition from private entrants by streamlining their production: “Same-day admissions [i.e. admission on day of surgery] are seen as imperative by independent [private] providers. Acute [public] trusts will need to reflect this as an integral element of any market strategy when seeking to demonstrate competitive advantage.” This explicit focus on admission on day of surgery means that, in addition to the more general incentives to increase efficiency brought about by surgical centre entry, we expect public hospitals facing increased pressure from private surgical centres to have improved their performance in this dimension in particular. Case mix ~~~~~~~~ The entry of private, for-profit surgical centres could also change the case mix at nearby incumbents due to risk selection by entrants. Whereas in classical private goods markets the profitability of selling to a particular customer is determined solely by their willingness to pay, in health care markets – as in many other markets for the provision of public services, such as social care and education – the profitability of treating a given customer will be influenced by characteristics of the customer that are often imperfectly observed. The influence of patient characteristics on profitability provides all hospitals with an incentive to refuse to treat the sickest patients. However, private, for-profit entrants like ISTCs are likely to be more willing than public hospital incumbents to actively select against costly patients, as for-profit firms are able to redistribute profits to shareholders, whereas public hospitals are, at most, only allowed to reinvest profits into the organisation. The literature on specialty hospitals in the US, for example, has found evidence that these providers select low-risk patients, leaving the sickest patients to nearby general hospitals (Barro et al., 2006; Winter, 2003; Cram et al., 2005). Two further factors add weight to the hypothesis that ISTCs had stronger incentives to risk-select than public hospital incumbents. First, ISTCs could legally decline to treat complicated cases, whereas public hospitals were formally prohibited from doing so. Second, as mentioned previously, ISTCs were prohibited from using NHS doctors, so their workplace culture likely differed sharply from that at incumbents. As Rose- Ackerman (1996) notes, the culture of staff plays a key role in dictating firm behaviour – thus these cultural differences may have led ISTCs to be more willing than NHS providers to engage in profit-driven risk-selection. Prospective reimbursement encourages cream- skimming, since it provides incentives for hospitals to avoid admitting patients whose cost of treatment is likely to exceed the regulated payment (Allen and Gertler, 1991; Ellis and McGuire, 1986; Ellis, 1998; Newhouse, 1989). We use DiD methods to estimate the extent to which ISTCs left incumbent NHS hospitals with a sicker, costlier mix of patients, over and above any secular changes in case mix over this period (either as a result of the introduction of prospective reimbursement, or for other reasons). Previous studies have confirmed that ISTCs treated healthier and less complex patients than nearby public hospitals (Street et al., 2010; Mason et al., 2008, 2010; Browne et al., 2008; Chard et al., 2011; Fagg et al., 2012). However, no one has yet compared the evolution of average patient severity at ISTC-exposed public hospitals with that at public hospitals unaffected by the ISTC programme, and shown that ISTC-exposed public hospitals experienced a larger reduction in average patient health status (measured using a Charlson Index) than public hospitals not exposed to the entry of an ISTC. Providing evidence of such an effect of ISTC entry is important because the case mix differences between ISTCs and nearby public hospitals documented by the existing literature may simply reflect the fact that ISTCs attracted patients who would not otherwise have undergone surgery.7","Our aim is to estimate the causal effect of the entry of private surgical centres on the efficiency, case mix, and case load of nearby incumbent public hospitals. We use difference-in-difference (DiD) regressions in which the impact of ISTC exposure is estimated from the mean change in outcomes for public hospitals in a treatment group (those that had a private surgical centre placed nearby) minus the mean change in outcomes for public hospitals in a control group (those that did not have a private surgical centre placed nearby) before and after entry occurred. This section describes our outcome measures, construction of treatment groups, and identification strategy. Data and outcome variables ~~~~~~~~~~~~~~~~~~~~~~~~~~ Our dataset is derived from the NHS Hospital Episode Statistics (HES) (HSCIC, 2016), which contains the universe of government-funded hospital admissions in England.8 Our data extract consists of all elective hip and knee replacements on patients aged 55–100 performed between financial years 2002/3 and 2008/9 (see Table 1). We focus on hip and knee replacements for two reasons. First, orthopaedic surgery was a major focus of the ISTC programme, as it was recognised in the early 2000s that achieving the government's waiting time targets was going to be more challenging in this surgical specialty than in any other area (Harrison and Appleby, 2005). Second, clinical practice in relation to hip and knee replacements did not change significantly during this period in ways that could affect LOS. As a result, any observed changes in LOS will likely be driven by NHS reforms, not by differential uptake of new medical technologies. We focus on hip and knee replacements performed in NHS hospitals. NHS hospital trusts (firms) often consist of multiple hospitals (individual sites) that can be located up to 100 km away from each other. We therefore analyse the data at site (hospital) level rather than trust (firm) level, and assign hospitals (sites) to treatment and control groups based on the site's proximity to the nearest ISTC. All references to ‘hospitals’ in this paper are therefore to hospital sites, not to trusts (firms). After cleaning and imputing missing values for the site code field, and applying exclusion criteria detailed below, there are 166 public hospitals treating hip and knee replacement patients from 2002/3 to 2008/9. Researchers have generally struggled to quantify hospital efficiency. In the absence of hospital cost data, many studies use proxy measures of efficiency such as LOS (Fenn and Davies, 1990; Martin and Smith, 1996; Gaynor et al., 2013). The logic underlying this measure is that, if a hospital can treat patients more quickly without any deterioration in clinical quality, then it must have become more efficient. However, a critical shortcoming of overall LOS as an efficiency measure is that recovery time after surgery is also heavily dependent on patient characteristics and health status. Moreover, a hospital's average LOS may reflect undesirable hospital behaviour such as cream skimming (prioritising treatment of less costly patients); dumping (avoiding treatment of costlier patients); and quality skimping (discharging patients ‘sicker and quicker’) (Epstein et al., 1990; Martin and Smith, 1996; Sudell et al., 1991). In this study, we use an innovative method to obtain a cleaner proxy for hospital efficiency. We decompose LOS for hip and knee replacements into two parts: the time from admission to surgery (pre-surgery LOS), and the time from surgery until discharge (post-surgery LOS). We hypothesise that, for elective orthopaedic surgery, pre-surgery LOS is not significantly influenced by patient characteristics, as there is rarely a clinical rationale for admitting an elective orthopaedic surgery patient before the scheduled day of their operation. In the early 2000s, fewer than 20% of hip and knee replacement patients had surgery on the day they were admitted to the hospital. Patients were often kept overnight before elective surgery not for clinical reasons, but because operating rooms were not available on their scheduled surgery date.9 The extent to which hospitals are able to schedule patient admissions to ensure that they line up with the availability of surgeons, support staff, and operating theatres will therefore be a direct function of the efficiency with which the hospital is run. By contrast, we view post- surgery LOS as a joint product of underlying hospital efficiency and patient characteristics. Therefore, when we estimate the effect of ISTC entry on post-surgery LOS, we interpret the results as a combined outcome of (i) competitive pressure brought about by ISTC entry, leading to efficiency improvements by nearby public incumbents, and (ii) ISTC cream skimming, leaving nearby public hospitals with a sicker mix of patients. To test our hypothesis that pre-surgery LOS is less influenced by patient characteristics than post-surgery LOS, we regress pre- and post-surgery LOS for hip and knee replacements against a range of patient characteristics. Patient characteristics included in the regression are: Charlson score; number of diagnoses; Index of Multiple Deprivation (IMD) income deprivation score; IMD health and disability deprivation score; dummy variables indicating self-discharge, urban residence, mixed ethnicity, Asian ethnicity, black ethnicity, other ethnicity, and unknown ethnicity; and a full set of case mix dummies with gender interacted with five-year age bins. The results, reported in Table 2, are consistent with our hypothesis – patient characteristics explain less than half a per cent of the variation in pre-surgery LOS, but 12.8% of the variation in post-surgery LOS. We observe changes in patient health status at NHS hospitals before and after ISTCs enter in order to directly measure whether risk selection occurred. To measure patient health status, we calculate a Charlson score for each patient. The Charlson score predicts a patient's 10-year survival probability based on their health status in relation to 17 conditions likely to lead to death. The score varies from 0 to 6, with 0 denoting the absence of any predictors of mortality (HSCIC, 2013).10 As proxies for health status and clinical risk, we also use the patient's age, as well as the IMD income domain (Noble et al., 2004), which reports the percentage of households in the patient's residential Lower Super Output Area (LSOA, a statistical geographical areas containing around 1500 residents) that are income deprived (in our dataset this variable ranges from 0 to 83). The data includes 478,226 hip and knee replacements performed from 2002/3 to 2008/9 that met the sample restrictions. As Table 1 illustrates, during the analysis period pre- surgery LOS, post-surgery LOS, and total LOS fell considerably.11 Treatment assignment ~~~~~~~~~~~~~~~~~~~~ We assign public hospitals to treatment or control groups based on their geographical proximity to the new market entrants, on the assumption that exposure to competition from these entrants is a product of proximity. In particular, we assign treatments by comparing the distance from an NHS hospital to its nearest ISTC with the percentiles of distance travelled by that hospital's hip and knee replacement patients. We measure the distance travelled by each hip and knee replacement patient to hospital using the centroids of a patient's residential LSOA to define home location. We then calculate, for each NHS hospital, percentiles of patient distance travelled (e.g. the distance that captures 25% of a hospital's hip and knee replacement patients). Percentiles of patient distance travelled can be endogenous to hospital quality – for example, a high-quality hospital may attract patients from further afield. To ameliorate this concern, we use percentiles of patient distance travelled based on patient flows from 2002/3 to 2004/5 (i.e. before implementation of either the ISTC programme or patient choice of hospital for elective surgery). Table 3 presents descriptive statistics for quantiles of patient distance travelled and the exposure of NHS hospitals to Wave 1 ISTCs within each of these quantile bands. Panel A presents the mean kilometre distances (averaging over the NHS hospitals in our estimation sample) corresponding to these patient percentile travel distances. The mean value of the 25th percentile of travel distance for hip and knee replacement patients is 4.25 km, while for the 95th percentile it is 26.34 km. Panel B shows how many NHS hospitals had a Wave 1 ISTC enter within each of these distance bands. To define Treatment and Control groups, we start by assuming that, if there is no ISTC entrant within an NHS hospital's 95th percentile of patient distance travelled – a widely-used definition of market size – then the NHS hospital is not exposed to the ISTC programme. These NHS hospitals are assigned to the Control group. We assign the remaining NHS hospitals to discrete treatment groups based on natural breaks in the distribution of distances from NHS hospital to the nearest ISTC in our dataset, measured in terms of patient travel percentiles (see Fig. 2). This assignment yields two discrete treatment groups – a Low Treatment group (ISTC within 95th percentile of patient distance travelled but not within 25th percentile) and a High Treatment group (ISTC within 25th percentile of patient distance travelled).12 There are 11 NHS hospitals in the High Treatment group, 51 in the Low Treatment group and 104 in the Control group. Fig. 3 maps the ISTCs and public hospitals in our data. In the robustness section, we show the effect of changing the threshold used to define the treatment groups on our estimates. One obvious concern is that treatment assignment might be endogenous to our primary outcome (LOS) because ISTCs may have opened near inefficient NHS providers. However, as noted earlier, ISTC placement decisions were driven by local NHS hospital waiting times, not by local hospital LOS or other performance measures. Waiting times are dependent on a wide range of demand and supply side capacity-related factors beyond simply hospitals' LOS. As a result, there is little reason to expect hospitals with high waiting times to necessarily have high LOS. Indeed, as we observe, hospitals' LOS is uncorrelated with their waiting times.13 Table 4 illustrates this point by comparing waiting times, total LOS, pre-surgery LOS, and post- surgery LOS for hip and knee replacement at High Treatment group, Low Treatment group and Control group hospitals in 2002/3 (the year that ISTC placement decisions were being made). Average waiting times in 2002/3 were around 6% higher at High Treatment group hospitals than at hospitals in our other groups. By contrast, there are no systematic differences between total LOS or post-surgery LOS at High Treatment group hospitals and others. Pre-surgery LOS is slightly lower at High Treatment group hospitals – i.e. ISTCs tended to enter near NHS hospitals that were already slightly more efficient, although there is no evidence to suggest that these small efficiency differences were a factor in ISTC location decisions. As we show in later analysis, these differences in pre-surgery LOS are not associated with any statistically significant differences in trends of pre- surgery between 2002/3 and 2004/5, prior to the entry of ISTCs – which is the key assumption of our DiD identification strategy. The fact that ISTCs entered where nearby public hospitals were already more efficient might be of concern if we were to find that ISTC exposure led nearby public hospitals to become less efficient, as it might suggest that mean reversion is driving our results. However, as we find the opposite (i.e. ISTCs entered near public hospitals that were already more efficient, and that became even more efficient in relative terms as a result of ISTC exposure), we have no reason to be concerned that these small differences in pre-ISTC levels of pre-surgery LOS across treated and control groups will confound our DiD estimates. We allocate NHS hospitals to treatment categories by comparing distance to ISTC with percentiles of patient distance travelled, not with kilometre distances, to control for rural-urban differences – treatment assignment based on fixed distances will over-estimate the size of markets in urban areas relative to rural areas, given the impact of urban congestion on travel speeds. In the robustness tests, we examine whether our results change if we use a treatment assignment strategy based on fixed distances from public hospital to ISTC. Treatment start and end dates ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ There is some ambiguity as to the appropriate way to define the policy-on and policy-off dates for a given public hospital exposed to ISTC entry. As the first ISTCs in our analysis opened in April 2005 (financial year 2005/6), we define 2004/5 as the last pre- treatment (pre-ISTC-programme) financial year.14 However, some ISTCs did not begin operations until 6 months to a year after their contracted start date. Moreover, when the initial ISTC contracts (generally around 4 or 5 years in length) were completed, some managed to secure further contracts, but others were shut down or absorbed into neighbouring NHS trusts. The fate of an ISTC was generally announced in the last year of the contract. Thus, if contract end date were used as treatment end date, estimates of treatment effects could be confounded by changes in behaviour due to anticipated contract completion. In response to these ambiguities, we employ a long differences specification using data from the 2004/5 and 2008/9 financial years. We choose 2004/5 as the pre- treatment period in our main specification because it is the last year before the first contract start date amongst the ISTCs we use for treatment assignment, and thus most likely to capture the effect of ISTC exposure as distinct from the effect of other elements of the government's reform programme. We choose 2008/9 as the post-treatment period to allow time for treatment effects to be realised, while avoiding contamination from responses to announcements concerning extension or non-extension of ISTC contracts. For robustness, we show how treatment effects change annually in the post-reform years from 2005/6 to 2008/9, and report estimates that define a public hospital's treatment start date as the contract start date of its nearest ISTC. Regression specification ~~~~~~~~~~~~~~~~~~~~~~~~ We identify the impact of hospital market entry using a DiD regression framework where dummy variables indicating treatment group membership are interacted with a post-policy dummy, which is switched on for the 2008/9 financial year. Regressions are run at the patient level and any non-binary dependent variables are log-transformed (after adding 1 to any variables that have a minimum value of zero) such that the treatment effects are interpretable as percentage changes. In this specification, t denotes the time period (financial year), postt ∈ {0, 1} denotes whether an observation occurs in the post-reform period, yijt denotes the outcome variable under consideration for patient i attending hospital j at time t, and highj and lowj denote dummies for the High and Low Treatment groups respectively. Treatment effects are given by the coefficients on the interaction terms, β4 and β5. We also include a dummy in our regressions to control for the type of procedure a patient undergoes (hip or knee surgery) but suppress this in the notation for simplicity. Our third specification is identical to (2), but includes an extensive set of controls for patient and hospital characteristics.15 All specifications are estimated using ordinary least squares (OLS), with standard errors clustered at the hospital level to account for correlation in unobservables within hospitals (between patients and over time). There are two core threats to our identification strategy. The first is that trends in outcome variables may not have been parallel between treatment and control groups prior to ISTC entry. The second is that there may have been other time-varying policy changes that also affected outcomes concurrently with the ISTC programme. To address the first possibility, we demonstrate that treated and control hospitals had parallel trends for our key dependent variables before the ISTC programme was launched by showing trends in the data graphically, and formally testing for statistically significant differences in trends. To address the risk that concurrent and correlated policy shocks drive our results, Section 5.4 discusses the two most prominent policy changes that could have affected outcomes contemporaneously with ISTC entry – the introduction of hospital competition via patient choice of hospital for elective surgery, and the enactment of differential health policies by Strategic Health Authorities (SHAs) – and illustrates that controlling for these policy changes does not materially influence our main estimates. Descriptive evidence ~~~~~~~~~~~~~~~~~~~~ Fig. 4 presents the evolution of key outcome variables – pre-surgery LOS, percentage treated on day of admission, post-surgery LOS, total LOS and Charlson score – between 2002/3 and 2008/9, for treatment and control groups. The shaded area represents the range of treatment start dates for the Wave 1 ISTCs. We expect that any treatment effects will arise either within the time period captured by shaded region or, if behavioural responses took place with a lag, some time thereafter. Each data point represents a month, but the plots are smoothed using a moving average of the month and the previous quarter. These graphs allow visual examination of whether treated and control groups followed similar trends prior to the entry of ISTCs. Panel A shows changes in pre-surgery LOS and illustrates that High Treatment group, Low Treatment group, and Control group hospitals follow similar trends before ISTC entry. Over and above a secular downward trend, reflecting general improvements in turnaround time, there is evidence of a treatment effect from ISTC entry. After ISTC entry, trends diverge, and by the end of the treatment period the reduction in pre-surgery LOS is notably larger for the High Treatment group than for the Control group. There also appears to be a smaller effect for the Low Treatment group. Panel B shows trends in the percentage of patients treated on the day of admission. All three groups have similar pre-entry trends, but after ISTC entry the percentage of patients treated on the day of admission increases more quickly for High Treatment group hospitals. Overall, Panels A and B provide visual evidence that the entry of private specialty surgical centres in the English NHS made nearby public hospitals more efficient, by reducing pre-surgery delays. Panel C shows trends in post-surgery LOS, which have a markedly different pattern. The High Treatment group, Low Treatment group, and Control group hospitals have similar pre-entry trends. However, there is a sharp increase in post-surgery LOS for the High Treatment group from the middle of the ISTC entry period onwards. Overall, after ISTC entry post-surgery LOS decreases in the High Treatment group less than in the Control group. Panel D presents trends in total LOS, which follow a similar pattern to post-surgery LOS. As discussed in Section 4.1, post-surgery LOS (and therefore total LOS) will be influenced both by changes in hospital efficiency due to increased competitive pressure from the entry of private surgical centres, and by changes in patient characteristics due to cream skimming by entrants. Panels C and D therefore provide suggestive evidence that the negative impact of ISTC cream skimming on nearby public hospitals' LOS may have outweighed any efficiency improvements with respect to LOS arising from competitive pressure from these new market entrants. Panel E looks more directly at the impacts of ISTC entry on public hospitals' case mix by plotting the evolution of average Charlson scores. The pre-policy levels and trends of the Charlson score are similar across treatment and control groups. However, the High Treatment group starts receiving sicker patients from early in the ISTC entry period. This evidence is consistent with our hypothesis that selection of less risky patients by ISTCs left a residual pool of higher-risk patients to be treated by public hospitals. Graphical evidence for other case mix variables is presented Appendix D. Overall, the similar pre- policy trends in treatment and control groups for all outcome variables provides strong support for our argument that DiD estimates are likely to provide an unbiased estimate of treatment effects from ISTC entry. The fact that pre-policy trends (and in many cases levels) of our outcome variables are similar across treated and untreated groups is consistent with our argument that the principal target of ISTC placement was to reduce waiting times for admission to hospital, not to reduce time spent in hospital or to improve clinical quality.16 Regression-based difference-in-difference estimates ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 5 presents our main difference-in-difference estimates of the effect of ISTC entry on log of pre-surgery LOS, percentage of patients treated on day of admission, and log of post-surgery LOS. The sample includes hip and knee replacements, with a hip replacement dummy included to account for level differences in outcomes between the two procedures. In Columns (1) to (3), the dependent variable is log of pre-surgery LOS. Column (1) presents estimates of Eq. (1), without hospital fixed effects or patient controls. The estimate implies that ISTC entry led to a 14.4% (=100(e−0.156 − 1)) reduction in pre-surgery LOS. In Column (2), we estimate Eq. (2), adding hospital and month × year fixed effects. The results are qualitatively unchanged and imply that ISTC entry led to a 16.1% reduction in pre-surgery LOS. In Column (3), we add patient controls and the results again remain qualitatively unchanged (we find a 16.6% reduction). That controlling for patient characteristics barely shifts the estimated treatment effects for the High Treatment group suggests that there is little selection into treatment on the basis of these observable demographic characteristics. This, in turn, implies that there is likely to be little selection into treatment on the basis of unobservable patient characteristics (Altonji et al., 2005). The impact on the Low Treatment group is of the same sign and around one third the magnitude of the High Treatment group effect, but is imprecisely measured and never significant at conventional levels. The most likely interpretation is that there were moderate impacts of ISTC entry on the Low Treatment group, but that our research design does not have sufficient power to detect them. Columns (4) to (6) examine the effect of ISTC entry on the proportion of patients who were treated on the day of admission. The results are similar with and without hospital fixed effects, and with and without patient controls. The results in Column (5), which estimates Eq. (2), shows that ISTC entry led to a 24.3 percentage point reduction in the proportion of patients treated on the day of admission at High Treatment group hospitals, from a baseline of 22.8% in 2004/5. As with pre-surgery LOS, the impact on the Low Treatment group is same-signed, smaller, and never significant at conventional levels. Columns (7) to (9) present estimates of the effect of ISTC entry on post-surgery LOS. In Column (7), we estimate that the entry of ISTCs led to an 8.47% increase in post-surgery LOS at High Treatment group hospitals. However, this effect is only significant at the 10% level. The precision and magnitude of our estimates is reduced when hospital and month × year fixed effects are included, in Column (8). Adding patient controls further reduces the size of the point estimate and decreases precision, as we would expect if patient characteristics affect post-surgery LOS and it is selection of less riskier patients into ISTCs and out of NHS hospitals which drives the treatment effects on post-surgery LOS. That patient controls reduce the magnitude of the post-surgery LOS estimates, but have little impact on the pre-surgery LOS estimates, provides further evidence that patient characteristics are a major driver of post-surgery LOS, but have little influence on pre-surgery LOS. We interpret changes in post-surgery LOS resulting from ISTC entry as a joint product of (i) changes in the mix of patients being treated by public hospitals, due to cream skimming by neighbouring ISTCs and (ii) behavioural responses by public hospital managers and clinicians to competition from new private entrants. Although only significant at the 10% level, the estimates in Column (7) suggest that the increases in post-surgery LOS generated by cream-skimming were larger than the reduction in LOS generated from any efficiency gains in terms of the total time patients spent in the hospital. An important point to note, in relation to the results presented in Table 5, is that the unreported coefficients on the High Treatment and Low Treatment indicators in Columns (1), (4) and (7) are always near-zero and statistically insignificant. For example, in Column (1) the coefficient on the High Treatment variable is 0.0308 (0.0442). These results imply that the treatment and control groups are balanced in terms of the pre-treatment, 2004/5 levels of the outcomes under investigation, consistent with the graphical evidence in Fig. 4 discussed above.17 Table 6 tests for cream skimming directly by assessing whether the entry of an ISTC left nearby public hospitals with a riskier mix of patients. We estimate Eq. (1) and Eq. (2) with hospital and month × year fixed effects; no specifications include patient controls. Columns (1) to (4) indicate that ISTC entry led to an 11.6% increase in the average Charlson score of patients at High Treatment group hospitals – or a 6.2 percentage point increase in the proportion of patients with a Charlson score of three or more – significant at the 5% level. Column (5) indicates that ISTC entry led to a 5.54% increase in the IMD income deprivation score at High Treatment group hospitals, although this point estimate reduces in magnitude and becomes imprecise when hospital and month × year fixed effects are added. We do not find that ISTC entry led to a precisely estimated increase in patient age at nearby NHS hospitals. Table 7 presents event study estimates of year-by-year effects of ISTC entry on log of pre-surgery LOS, percentage of patients treated on day of admission, and log of Charlson score at exposed NHS hospitals, from 2004/5 to 2008/9. The estimates mirror our graphical evidence in Fig. 4 and show that ISTC entry had a statistically significant effect on these outcomes at exposed NHS hospitals from 2007/8 onwards. Appendix J explores the impact of ISTC entry on clinical quality at NHS hospitals by analysing changes in 30-day in-hospital mortality from acute myocardial infarction (AMI) at nearby public hospitals. We find that, after controlling for patient characteristics, ISTC entry did not have a statistically significant effect on AMI mortality at nearby public hospitals. The results suggest that the efficiency improvements reported in Columns (1) to (6) of Table 5 were achieved without any evidence of concomitant deterioration in clinical quality. Treatment assignment using fixed distances ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 8 presents estimates of the effect of ISTC entry when treatment assignment is based not on patient flows but on fixed kilometre distances between ISTCs and NHS hospitals. Specifically, the High Treatment group comprises NHS hospitals that had an ISTC enter within 5 km, the Low Treatment group comprises NHS hospitals that had an ISTC enter within 30 km (but not within 5 km), and the Control group comprises NHS hospitals that did not have an ISTC enter within 30 km. Using this definition, the High Treatment group contains 14 hospitals, the Low Treatment group 78 hospitals, and the Control group 77 hospitals. Table 8 indicates that ISTC entry within 5 km of an NHS hospital led to a 14.7% reduction in pre-surgery LOS, a 21.9 percentage point increase in the share of patients treated on the day of admission, and an 11.2% increase in the average Charlson score. Controlling for contemporaneous NHS policy changes ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ One concern with our DiD identification strategy is that the resulting estimates may be biased by other policies, implemented concurrently with the ISTC programme, that had a differential effect on our outcome variables across treated and control groups. The most prominent such policy is the 2006 introduction of hospital competition within the NHS via patient choice of hospital for elective surgery. In addition, during this era much NHS policy was dictated by ten regional Strategic Health Authorities (SHAs) – differences in SHA policies implemented during the ISTC period may bias our results, to the extent that ISTC entry was differentiated across regions of England. We investigate these possibilities in Table 9; all reported estimates use as their baseline Eq. (2). Columns (1) to (3) test whether the estimates reported in Tables 5 and 6 are robust to controlling for the 2006 patient choice reforms. We control for overall competition intensity by including a measure of market concentration (a time-invariant negative log of hospital HHI) interacted with a post-ISTC-entry dummy variable.18 If the patient choice reforms were driving the results, inclusion of this interaction term would severely attenuate our estimates. However, including this interaction term does not materially change the estimates – they remain precisely estimated and similar in magnitude. Columns (4) to (6) test for other region-specific policy changes that could be driving our results. To do so, we interact dummies for each of the ten English SHAs with a post-ISTC-entry dummy. These additional controls do not materially impact our results. The results are also robust to controlling more flexibly for differential SHA policies and regional trends via separate SHA × year or SHA × year × month interaction terms (see Appendix G). Overall, the estimates reported in Table 9 provide assurance that our main estimates are not driven by the most worrisome potential sources of bias from contemporaneous policy changes. Altering the threshold used to define treatment exposure ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ So far, we have defined the High Treatment group to include any public hospital that had an ISTC enter within the 25th percentile of patient distance travelled between 2002/3 and 2004/5, or, in the alternative specification reported in Table 8, within 5 km. Figs. 5 and 6 show how the estimates of Eq. (2) change as we vary the treatment group definition thresholds. Fig. 5 (Panel A for log of pre-surgery LOS and Panel B for log of Charlson score) shows how the estimates change when the threshold used to define the treatment group changes from one that captures 15% of a hospital's hip and knee replacement patients, to one that captures 95%. Panels A and B illustrate that the estimated treatment effects decrease as the treated group is defined more widely. Fig. 6 performs the same exercise when defining treatment exposure using fixed distances from NHS hospital to ISTC, as in Table 8. Panel A shows that ISTC entry within 5 km of an NHS hospital leads to a precisely estimated reduction in log of pre-surgery LOS at incumbents. As the threshold used to define the treatment group increases, the estimated treatment effects lose precision and asymptote to zero. Panel B, for log of Charlson score, shows a similar sensitivity to the threshold used to define treatment exposure. ISTC entry within 5 km of an NHS hospital leads to a large and statistically significant increase in the average Charlson score at incumbents, with estimated treatment effects decreasing as the distance to ISTC used to define the treatment group increases. Overall, Figs. 5 and 6 demonstrate that our estimated treatment effects are robust to the exact threshold used to define treatment exposure. Furthermore, as the treatment group is expanded, there is a negative gradient to our estimates, which is broadly supportive of the argument that our estimated treatment effects are driven by ISTC exposure. Additional tests of robustness ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 10 reports additional robustness tests for our main estimates of the impact of ISTC entry on log of pre-surgery LOS, percentage of patients treated on day of admission, and log of Charlson score at High Treatment group hospitals. We estimate Eq. (2) unless otherwise noted. Panel A formally tests for parallel pre-reform trends for the High Treatment group relative to the Control group using a ‘placebo’ DiD regression where 2002/3 is the base year and 2004/5 (the last year before ISTC entry) is the treated year. This regression allows us to explore whether there was a statistically significant difference in the change in key outcomes in our High Treatment group relative to the change in key outcomes for the Control group during our pre-period from 2002/3 to 2004/5. If there were statistically significant differences (i.e. a ‘placebo’ treatment effect), it would illustrate that there was a difference in trends over the pre-reform period that would potentially invalidate our identification strategy. Consistent with the graphical evidence in Fig. 4, none of the 2002/3 to 2004/5 point estimates are statistically significant at conventional levels, confirming that there were no statistically significance differences in trends for key outcomes between our control group and high treatment group from 2002/3 to 2004/5 (our pre-period). Though not reported, we also find no statistically significant treatment effects for the Low Treatment group relative to the Control group over the same period. To address any residual concerns about comparability of treated and untreated hospitals, Panel B reports inverse propensity score weighted estimates where we weight our treatment and control groups by the inverse of the probability that an observation belongs in its group. To do this, we first calculate the probability of a hospital being assigned to the High and Low Treatment groups based on hospital and average patient characteristics.19 We then run separate regressions to estimate treatment effects for the High and Low Treatment groups, using as probability weights the inverse of the probability of assignment to the group that the observation was actually assigned to. While it is impossible to reject a null hypothesis of no significant differences between treated and control groups with respect to the above determinants of treatment assignment even before these weights are applied, weighting by the inverse propensity score increases the similarity of the treated and control groups further still, by giving higher weight to control observations that look more like treated observations, and vice versa. The resulting estimates are similar to the headline findings, but more precise. The corresponding (unreported) estimates for the Low Treatment group are not statistically significant. Our main specification accounts for the fact that different ISTCs commenced operations at different times by using a long differences estimation strategy with 2004/5 as the pre-reform year and 2008/9 as the post-reform year. An alternative approach is to define t = 0 (the treatment start date) for each public hospital as the contract start month (or month of ‘full service commencement’) of the nearest ISTC, and to use pre- and post-reform periods defined relative to t = 0 rather than using calendar time. We do not use this as our main specification because the contract start date is not always an accurate indicator of when an ISTC started treating patients. Nonetheless, Panel C presents estimates from such a specification, with months −12 to −1 before ISTC entry designated as the pre-reform period, and months 24 to 35 after entry designated as the post-reform period. Hospitals in the Control group are allocated a placebo ‘treatment’ start date equal to the contract start date of the nearest ISTC, even though this ISTC lies outside the 95th percentile of patient distance travelled. The resulting estimates are very similar to our main results. Panel D reports estimates using a treatment assignment strategy that centres hospital markets on GP surgeries rather than hospitals. Hospital-centred measures of market size based on percentiles of patient distance travelled are potentially endogenous to hospital performance. While we address this concern by basing treatment assignment on percentiles of patient distance travelled between 2002/3 and 2004/5 – before the introduction of patient choice of hospital or the ISTC programme – concerns may remain. To address these concerns, this check assigns treatments by constructing a list of all the NHS hospitals and ISTCs that fall within each GP surgery's market (95th percentile of distance from GP surgery to NHS hospital for that GP surgery's hip and knee replacement patients). If an ISTC is in 95% of the GP surgery markets that an NHS hospital falls within, that NHS hospital is assigned to the High Treatment group. If an ISTC is in 75% of the GP surgery markets that an NHS hospital falls within, but not 95%, that NHS hospital is assigned to the Low Treatment group. All other NHS hospitals are assigned to the Control group. The estimates reported in Panel D are consistent with our main results, providing assurance that they are not driven by assignment of treatments based on hospital-centred market definitions. Panel E reports estimates when we do not take logs of the outcome variables (pre-surgery LOS and Charlson score). We continue to find that ISTC entry made nearby public hospitals more efficient, but also left them with sicker patients; the estimates are nearly equal to the exponent of our main results. A number of other checks are reported in the online Supplementary Material (see Appendices E through G); they provide further confirmation that our results are robust to a wide range of specifications. Ruling out the possible confounding effect of changes in patient volumes ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Increases in local capacity could potentially affect incumbents' efficiency by reducing congestion and overcrowding, independent of any competitive pressure exerted by entrants. Moreover, public hospitals located near new private entrants may have experienced a reduction in demand. Any resulting reduction in case loads at nearby public hospitals could affect average pre-surgery LOS at these incumbents, given the important influence of volume on efficiency. We therefore investigated the impact of ISTC exposure on case load using similar regressions to those used for our main estimates, but found no statistically significant evidence of case load reductions in our High Treatment group, which had the biggest reductions in pre-surgery LOS (results available in Table H.1 in the online Supplementary Material). This suggests that the increases in local capacity brought about by private surgical centre entry did not lead to any reduction in the volume of patients treated at High Treatment group hospitals, but, instead, served to take people off waiting lists and reduce waiting times. That is, ISTC exposure led to shorter pre-surgery LOS in close-neighbouring hospitals without any concomitant reduction in the volume of patients being treated. We did find, however, significant reductions in the volume of patients being treated at Low Treatment group hospitals. This finding seems to suggest that ISTC entry did not simply add to overall clinical capacity, but, at least to some extent, may have reduced patient volume at public hospitals with which they shared a market – although, crucially, these patients seem not to have been drawn from the closest hospitals (i.e. those in the High Treatment group) in which we observe statistically significant reductions in pre-surgery LOS. These findings are broadly supportive of our conjecture that the reductions in pre-surgery LOS reported in Table 5 arose primarily through competitive incentives.20","This paper examines the effect of a UK government programme designed to increase capacity and competition by facilitating the entry of private, for-profit specialty surgical centres into the English NHS. We test the impact that the entry of these facilities – ISTCs – had on incumbent public hospitals' efficiency, case mix, and case load. We exploit the fact that ISTC location decisions were driven by local waiting times, not by other hospital characteristics such as LOS or clinical performance, to construct treatment and control groups that are comparable with respect to the outcome variables examined. Indeed, we demonstrate that trends of key outcome variables – including pre-surgery LOS, post- surgery LOS, and patient case mix – were the same for public hospitals that had an ISTC enter nearby as for those that did not. We find that public hospitals that had a private, for-profit surgical centre enter close by experienced substantial reductions in pre- surgery length of stay for hip and knee replacement surgery. The addition of an ISTC to a public hospital's immediate neighbourhood led to a decrease in pre-surgery LOS of around 16% – or a 24 percentage point increase in the proportion of patients treated on the day of admission. Given that these faster turnaround times were achieved without additional expenditure (as they occurred in an environment with fixed payments per procedure), they suggest that hospitals exposed to competition from new private entrants became more efficient. As well as investigating possible positive effects of ISTC entry on the efficiency of incumbent public hospitals, we looked for evidence of possible negative effects in the form of worsened case mix. We find that ISTC entry led nearby public hospitals to experience an 11.6% increase in patients' average illness severity as captured by the Charlson score – or a 6.2 percentage point increase in the proportion of patients with a Charlson score of three or more. We also find suggestive evidence that this increase in the sickness of incumbent hospitals' patients led to an increase in post- surgery LOS. While our identification strategy is unable to pinpoint how much of this increase in patient complexity at incumbent NHS hospitals is due to ISTCs actively applying their exclusion criteria – as opposed, for example, to differential choice behaviour by sick and healthy patients – the fact remains that surgical centre entry appears to have led to worsened case mix at nearby incumbents, irrespective of the exact channel by which this effect arose. In principle, this sorting of patients between private surgical centres and public hospitals could represent an efficiency-improving division of responsibility between routine and more complex cases. Indeed, this appears to have been the government's rationale for including wide-ranging exclusion criteria in ISTC contracts, which would allow these private entrants to focus on routine cases. Thus, the risk-selection we document was to some extent an intended policy outcome. However, the fact that policymakers explicitly intended such a division of responsibility between private entrants and public incumbents does not automatically imply that the division was devoid of negative consequences. To ensure that a division of responsibility between complex and straightforward cases does not have a negative effect on the financial position of providers that receive the most complex patients, hospitals that treat sicker patients must be appropriately compensated.21 Unfortunately, NHS reimbursement rates during the period we study did not adequately adjust for patient severity (Mason et al., 2008). This situation not only provided private surgical centre entrants with an added impetus to risk select, but it also meant that nearby incumbent public hospitals were left treating a costlier mix of patients without adequate financial compensation. While the prospective reimbursement regime (Payment by Results) was updated in April 2009 to include a more dramatic adjustment for patient severity, Mason et al. (2008) state that providers were still likely underpaid for treating sicker patients, and note that is unlikely that a prospective reimbursement system can ever be designed to fully compensate hospitals for a more costly case mix. Our work highlights one trade-off that arises from the entry of for- profit surgical centres. We show that entry can stimulate improvements in the performance of incumbents, but that entrants may engage in risk selection at the expense of incumbents. These findings offer insights into many situations involving provision of public services where profitability is influenced by characteristics of the recipient: namely that the case for increased private sector involvement should depend, in part, on comparison of the benefits of increased competitive pressure, with the costs arising from cream skimming by private entrants. Our work is part of the broader task of evaluating the overall social welfare implications of the entry of private firms into the market for delivery of public services. Such an overall evaluation would not only consider the impacts of private provider entry on the performance of public incumbents, but would also compare the performance of private and public providers, as well as taking into account any increases in overall capacity resulting from private surgical centre entry. An interesting thought experiment is to consider whether our selection results were a function of the for-profit status of entrants, or of the fact that entrants were specialist surgical centres who competed against full service hospitals. Ultimately, theory and the broader evidence suggest that both distinctions – competition from for- profit ventures and competition from specialist surgical centres – likely drove our selection results (Barro et al., 2006; Dafny, 2005; Cram et al., 2005; Street et al., 2010). Our findings raise a key question: is it possible to reap the positive effects of increased competition resulting from expanded independent sector provision within the NHS, without the negative effects? It is likely that the answer to this question depends on whether it is possible to risk-adjust payments and outcome measures sufficiently to ensure that independent sector providers have an incentive to make profits by raising efficiency and clinical quality, not by selecting against certain patients. Absent suitably risk- adjusted payments, new entrants may take steps to attract patients that are less costly to treat, leaving incumbents with a riskier mix of patients."],["This paper introduces a framework for studying the optimal dynamic allocation of foreign aid among multiple recipients. We pose the problem as one of weighted global welfare maximization. A donor in the North chooses an optimal path for international transfers, anticipating that consumption and investment decisions will be made by optimizing households in the South, and accounting for limits in the extent to which recipients can effectively absorb aid. We present quantitative results on optimal aid policy by applying our approach to a neoclassical growth model, where the scope for aid-funded growth is determined by the recipients' distance from steady-state. --------------------------------------------------------------------------------","Since the 1969 Pearson Commission, there has been a standard benchmark for the generosity of foreign aid programs: developed countries should donate at least seven-tenths of one per cent of their GDP. In practice, however, relatively few countries have achieved this level of generosity. Moreover, its foundations are easy to question. The target was calculated more than forty years ago using a combination of the Harrod–Domar model and financial programming. This rather mechanical approach would not command much support today. In this paper, we introduce a new framework for studying optimal aid policies. We model foreign aid as a form of global redistribution. A utilitarian, forward-looking social planner seeks to maximize a weighted average of welfare in the global North and the global South. The planner decides on an optimal path of international transfers, anticipating that consumption and investment decisions in the North and South will be made by optimizing households within each region. These households cannot borrow or lend internationally, and the scope for global redistribution is limited by diminishing returns to aid. This framework can be used to study the optimal generosity of aid, and its relationship to absorptive capacity. It can also be used to inform the timing of aid, and its allocation across countries that differ in development levels. We describe circumstances in which donors should seek to increase generosity over time, and examine whether middle-income countries should ever be a priority for aid. Importantly, the framework is tractable and could be extended in many directions; for example, it could be adapted to study the implications of capital mobility for optimal aid policies. The starting point is to model both the global North and South as neoclassical Ramsey economies. These economies differ in their levels of TFP and income, and in their distances from steady-state. Obstfeld (1999) analyzed the effect of exogenous transfers on a Southern economy; in this paper, that problem will be nested within the problem facing a Northern donor, so that the level and timing of transfers will be endogenously determined. At first glance, the Ramsey model may seem too stylized for this purpose, since it neglects political economy forces that will often be central to aid effectiveness. But the Ramsey model casts sharp light on a direct consequence of aid flows, which is to relax intertemporal resource constraints. Studying this in isolation should enhance our understanding of the choices facing donors. A further motivation is the growing interest in cash transfers to households as a form of poverty alleviation. Hanlon et al. (2010) argue that transfers direct to households would be more beneficial than more traditional forms of aid. In 2013, the Indian government launched an ambitious Direct Benefit Transfer scheme, intended eventually to replace multiple welfare programs with cash transfers to households. Using evidence from randomized trials, Gertler et al. (2012) and Haushofer and Shapiro (2013) find that cash transfers to poor households are partly invested. It is therefore interesting to ask: what happens when aid is used to relax household budget constraints, and what are the implications for optimal aid policies? In our framework, utility functions are concave and some global redistribution is optimal, but its extent will be constrained by diminishing returns to aid. Obstfeld (1999) concluded that the welfare benefit of (exogenous) aid was modest even without an aid absorption constraint, but our analysis turns that logic on its head. If the donor is also seen as a Ramsey economy, the opportunity cost of aid for the donor is similarly modest. Hence, we sometimes find that donors should be generous, especially for recipient economies that are close to subsistence. The welfare impact of aid may be substantial, and aid can be justified even when a substantial fraction of it is wasted. To investigate the quantitative implications, we use simulations. Donor and recipient initial conditions are based on data from the Penn World Table, combined with assumptions on structural parameters. We consider isoelastic (CRRA) preferences and Stone–Geary ‘subsistence consumption’ preferences. Under Stone-Geary preferences, the effects of aid on investment and growth are especially strong, and the donor maintains a higher level of aid generosity for a longer period. The limits to recipient absorptive capacity play an important role throughout. When recipients are capable of absorbing relatively high levels of aid, the optimal path of aid is typically front-loaded. But if the South lacks absorptive capacity, the North may want to increase the generosity of its aid over time, relative to Northern GDP. This reflects our assumption that, as the South develops, it can use aid more effectively. We examine what happens if the donor nevertheless maintains aid as a fixed share of its GDP, as in the Pearson Commission benchmark, and compare this to the fully- flexible optimal path. We show that, perhaps surprisingly, the welfare costs of restricting aid to a fixed proportion of donor GDP are often borne in equilibrium by the North, rather than the South. Another issue of recent interest has been whether donors should make transfers to middle-income countries (see, for example, Kanbur and Sumner, 2012). We therefore study what happens in the case of two aid recipients, where one is a low-income country and one a middle-income country with a larger population. Our simulations of optimal policies indicate large changes over time in aid generosity and in the division of aid between recipients. Given diminishing returns to aid intensity, most of the aid may be directed at the middle-income recipient initially, with the allocation later switching towards the smaller, poorer recipient. Throughout the analysis, we assume that the capital account is closed. This assumption is common in the literature, but some commentators argue that poor countries do not need aid when investment can be financed by capital inflows. In practice, capital has tended to flow to middle-income countries rather than the poorest countries. And even when the capital account is open, there is still a role for aid, to finance both consumption and the accumulation of assets owned domestically. It is likely that the welfare gains from aid would be less in this case, since initial consumption is higher; nevertheless, financing higher consumption would remain valuable. The quantitative investigation of this would be an interesting topic for future research.1 Since the model is stylized, it is worth clarifying the intended contribution. The paper provides a new way of framing the decision problems facing aid donors. The model could be extended, and made more realistic, in many directions; it is a first step towards richer quantitative frameworks that could inform future aid policies. The current simulations serve two more limited aims. The first is to investigate some qualitative results about aid policies when absorption constraints matter: these results are not special cases, but can emerge as important under reasonable parameter assumptions. The second aim is to illustrate what might be learnt by future research using more complicated models. For example, the current analysis is too preliminary to suggest a replacement for the Pearson Commission benchmark, but it draws attention to some relevant considerations. A companion paper, Carter (2014), uses related ideas to study aid allocation rules of the type sometimes implemented by donors. The existing literature has generally considered simpler models of North–South interactions, often in models with just one or two periods. The majority of this work is theoretical, and studies the effects of exogenous transfers rather than deriving their optimal time path. Some papers explore the transfer problem, and especially the possibility of transfer paradoxes driven by terms-of- trade effects; Eaton (1989) surveys the older literature. An alternative approach is taken by Chamon and Kremer (2009), who construct a multi-country model in which developing countries gradually integrate with the world economy. They discuss the potential role of aid in accelerating this process, but aid is not included in the version of the model they calibrate. The paper is structured as follows. The next section provides the formal description and analysis of the donor's decision problem. Section 3 describes the assumptions used in our simulations. Section 4 presents initial analysis of the optimal path of aid and Section 5 extends the analysis to subsistence economies. Section 6 considers two recipients which differ in population size and income. Section 7 carries out a sensitivity analysis and discusses potential extensions, before Section 8 concludes.","Our paper deliberately takes a narrow view of the donor's problem, seen exclusively in terms of international resource transfers. We hope to show that even this narrow view could inform the design of aid programs. Since the 1960s, cumulative spending on foreign aid has exceeded three trillion dollars in nominal terms, a figure that would be even higher in today's prices. Yet basic issues remain contested and, in some cases, rarely studied. How generous should aid flows be, and should donors choose aid targets relative to donor resources, or to recipient GDP? To what degree should this generosity be greater, early in the development process? When allocating aid across multiple recipients, how sensitive should aid flows be to recipient income levels? These are all debates that can be informed by the approach that we develop here. We use the framework to revisit some long-standing questions. Some of these questions concern the Pearson Commission benchmark, which appeared in United Nations Resolution number 2626, from October 1970. The 0.70% target was originally justified using a financing gap calculation, based on assessed capital needs for a Harrod–Domar economy (see Clemens and Moss, 2007, for a full history). It seems unsatisfactory to base aid policies on a model and form of analysis that are clearly outdated. Even the basic assumption that aid should be a fixed share of donor GDP might not be supported by a more rigorous approach. We agree with Clemens and Moss (2007) that it seems backwards to determine aid levels based on the size of donors, rather than conditions in recipient economies. The optimal degree of redistribution should be dependent on the extent of between-country differences, the nature of preferences, and the extent of absorption constraints, all of which play a role in our analysis. We formulate the North's decisions in terms of an explicit dynamic optimization problem. The North decides on an optimal path of transfers to the South. As in Obstfeld (1999), aid will benefit Southern economies in two ways. Aid accelerates the rate at which the South converges to its steady-state, and allows the South to sustain a higher level of consumption than otherwise. Stripping the aid problem down to these two roles will clarify their implications for aid policies and allocation decisions. In our simulations, we find that accelerating convergence to steady-state plays only a minor role. This does not mean the aid is wasted, however. In a model of this type, aid is effective to the extent that it raises consumption immediately, later, or both.2 Framing the donor's problem as one of weighted global welfare maximization has several advantages. First, the opportunity costs of aid arise endogenously from the structure of the model. Second, our approach provides a mapping between an explicit weight ω on Southern utility and optimal aid policies.3 We can readily study normative questions: how generous should aid be, and what does the optimal time profile look like? We now set out the decision problems formally. The North and South are both characterized as Ramsey economies. We study the South's problem first, and then the optimal control problem facing the Northern donor, where the optimizing behavior of the South represents a constraint in the North's decision problem. To simplify the presentation, initially we set the rates of technical progress and population growth to zero, but the necessary extensions are straightforward. We assume that aid is ultimately distributed to Southern households, and these are individually too small to internalize the effects of their actions on donor policies. Hence, we can consider the donor's problem without needing to allow strategic interactions between donors and (multiple) recipients. That would require analysis in terms of a dynamic game, and would not be straightforward. For the same reason, we do not model political economy forces explicitly, but these simplifications allow an analysis that is richer in other ways. The decision problem for the South ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Political economy approaches often imply that some aid is wasted or diverted. Rather than develop a specific structural model, our aid impact function translates Northern aid donations into the actual transfers received by Southern households, allowing for wastage or diversion.4 Our formulation nests the case where the marginal benefit of aid is declining in aid intensity, the ratio of aid to recipient GDP. This idea has often been investigated empirically, as in Burnside and Dollar (2000) and Clemens et al. (2012).5 One possible story is that, as aid intensity increases, so does the proportion of aid that is wasted or appropriated by a local elite. High aid intensity may also lead to Dutch Disease effects. It could have adverse effects on a domestic political equilibrium, partly by influencing rents to sovereignty, and perhaps by undermining long-term accountability and state capacity. A proliferation of aid projects and programs could overwhelm the capacity of the recipient government. These mechanisms have been widely discussed (see, for example, Temple (2010)). As a consequence, it seems essential to allow for limits to absorptive capacity. In our framework, these limits are eased by growth: as the GDP of an aid recipient increases, the recipient can use a given amount of aid more effectively. In the first version of the South's problem, the donor's optimal policy is not time- consistent, due to the externality. We follow much of the literature on conditionality, and assume that the donor has access to a commitment technology. This could take the form of explicit commitments to aid made through domestic legislation or international agreements that would be costly to reverse. This simplifies the analysis, and is relatively natural given our interest in normative questions. Further, in the cases considered below, the quantitative effects of the externality will be modest. The decision problem for the North ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Optimality at the initial date further requires zS0 = 0.8 In fact, the Northern planner's problem has a solution with zSt ≡ 0 at all dates. The Southern planner internalizes the capital externality, and the shadow value of Southern capital is the same for both planners. The level of aid is pinned down by the first-order condition (3). The solution is time-consistent, as the South values capital in the same way as the Northern planner. Again optimality requires zS0 = 0, but no solution exists with zSt ≡ 0. The Northern planner's Euler Eq. (8) would then be incompatible with the South's Euler Eq. (1) because of the capital externality. (For example, one can see that zS > 0 in the steady-state.) The externality leads to time inconsistency. The social returns to capital are higher than perceived by Southern households, given the benefits of higher output for aid absorption. To maintain the returns to capital and encourage capital accumulation, the Northern planner is less generous to the decentralized South than it would be to a social planner in the South. But at later dates, in the wake of past investment, the Northern planner would like to transfer additional aid in order to increase Southern consumption. Hence, the North's aid decision is time inconsistent. The model based on decentralized investment decisions is arguably more realistic, and will be the focus of our simulations. When we later examine the contrasting case of a Southern planner, the optimal generosity of aid is similar, but with higher investment in the early stages of the transition. These effects are quantitatively modest, suggesting that time-consistent aid paths would not look greatly different from those under full commitment. As noted earlier, we assume that the donor can commit to an entire path for aid flows, and leave the investigation of time- consistent aid policies to future work, perhaps using the approach of Cohen and Michel (1988). Contributions related to ours include Kopczuk et al. (2005), Weyl (2014) and especially Kemp et al. (1990). The latter paper uses optimal control methods to study the timing and generosity of aid needed to maximize weighted global welfare, just as we do. A crucial difference with our paper is that, although they briefly acknowledge the possibility of absorption constraints, they do not analyze them. Instead, they concentrate on the case where the South's resource constraint is linear in aid. They show that, at time zero, the optimal policy involves a transfer of part of the North's capital stock to the South.9 In our work, we rule out stock transfers, a step which is a natural counterpart to the assumption of absorption constraints, as they note. A further difference is that Kemp et al. (1990) is purely theoretical, whereas we present quantitative results based on simulations. The steady-state ~~~~~~~~~~~~~~~~ The first equation has a natural interpretation: the Northern donor balances the marginal cost of aid – represented by the marginal utility of forgone Northern consumption – against the marginal benefit of aid, namely the marginal utility of Southern consumption weighted by ω and LS, and multiplied by the derivative of the aid impact function with respect to aid.10 Under our assumptions, the equation has some interesting implications when parameters are such that cN∗ > cS∗. First, under absorption constraints, it is not optimal to equalize steady-state marginal utilities even when ω = 1. Second, it can be shown that steady-state aid is increasing in ω, as expected. Third, another simple result applies under isoelastic preferences: steady-state aid is decreasing in the intertemporal elasticity of substitution. All of these results are in line with intuition. In other respects, the externality complicates the comparative statics. We saw earlier that, with isoelastic preferences, increasing the intertemporal elasticity of substitution reduces the extent of global redistribution, as expected. But this is not always true in the version of the model with an externality. In principle, there could be reasonable functional forms and parameter values for which steady-state aid is increasing in the intertemporal elasticity of substitution over some range. This counter-intuitive possibility arises because the externality-related wedge term depends on the elasticity. Since the effects of the externality are modest in our simulations, we do not pursue this question further. The structure of the model does not rule out negative transfers. In fact, for a sufficiently low weight ω on Southern welfare, the North may choose to transfer income from the South to the North (‘negative aid’) at some points in time. This never happens in our simulations for the values of ω that we consider, but would emerge for weights close to zero. A richer analysis, at the expense of greater complexity, would either constrain aid to be non-negative — this is what Kemp et al. (1990) call the ‘non- cooperative’ case — or modify the aid impact function so that negative transfers are costly for the North to implement. Note that, in our simulations, both the South and North will start below their respective balanced growth paths. We could assume that the North is on its balanced growth path and decides to donate a fixed share of its GDP, so that aid to the South grows steadily over time. This would remove the need to model the North explicitly, and Carter (2014) and Carter and Temple (2014) use this simplification. But if we want to study aid generosity, and allow the share of aid in Northern GDP to change over time, then the more general approach of this paper is useful. Aid generosity has implications for Northern consumption, and modeling the North explicitly is a natural way to capture the time-varying opportunity cost of aid.","We now describe the assumptions used in our simulations. As we noted earlier, the simulations that we present are best seen as preliminary. This is partly because the model is stylized, and partly because the appropriate parameter choices are uncertain. The simulations illustrate what might be learnt from richer quantitative analyses in the future. In the past, the broad principles of aid programs, and some of their details, have often drawn on rather simple and unsatisfactory models, as Easterly (1999) emphasizes. The parameter choices for preferences are relatively straightforward. We assume that σ is equal to two, a common choice in the literature.12 We set the discount rate ρ equal to 0.03; this corresponds to the choices of Obstfeld (1999) and Gourinchas and Jeanne (2006). Note that both σ and ρ will influence the optimal generosity of aid. For simplicity, we assume that donors and recipients have access to a Cobb–Douglas technology with a common exponent on capital, but different levels of TFP. We assume the exponent on capital is 0.50. This is higher than most estimates of physical capital's share of income, but some authors argue that a broader notion of capital is needed for neoclassical growth models to be consistent with the data: see Mankiw et al. (1992) and, especially, Barro and Sala-i- Martin (2004). Our choice of 0.50 has been used in related contexts, such as in Kraay and Raddatz (2007). The choice matters because it will influence the speed at which the South converges to its steady-state, and the rate at which the marginal product of capital falls along the transition. To explore the predictions of the model, we also need to make assumptions about the long-run growth rates of GDP per capita and population, the size of the South relative to the North, and their initial capital–output ratios. Data on output, investment and population are taken from the Penn World Table version 6.3. We estimate the capital stock for 110 countries using the perpetual inventory method over the period 1960–2007, following Bernanke and Gürkaynak (2002) and adopting their depreciation rate of 0.06. We then aggregate countries into two units, the global North (the donor) and the global South (the recipient). The 33 countries aggregated into the Northern economy are those with output per capita above 20,000 in 2007 international (PWT) dollars, while the remainder are classed as the South. For our initial investigations, we sometimes exclude China and India, with their large populations.13 This keeps the quantitative analysis comparable with the aid decisions made in practice. As is well known, aid receipts per capita are low for China and India, partly reflecting the ‘small country bias’ in aid allocation; excluding these two countries helps to keep the model close to the data. This leaves us with an aid recipient whose population size is 2.7 times that of the donor, denoted S′ in Table 1. The final two columns of Table 1, S1 and S2, show the South sub- divided into low and middle income recipients, for use in Section 6. The cut-off of 20,000 international dollars roughly corresponds to the upper quartile of the GDP per capita distribution. We then calculate total output, capital stock and population for the two regions, for the most recent year in the PWT 6.3 data (2007), making no distinction between population and labor force. This procedure yields a capital–output ratio for each region. Since we assume Cobb–Douglas production functions for both North and South, the capital–output ratios imply the initial levels of capital per effective worker. We can then infer the relative TFP and GDP per capita of the South. We assume that rates of technical progress and population growth are the same for donor and recipient, helping to ensure a balanced growth path. The first assumption is common in the empirical growth literature. We adopt a rate of technical progress of 2% a year, as in Mankiw et al. (1992). This is also approximately the average growth rate in our Northern group of countries for the most recent decade in the data. The assumption that population grows at the same rate in donor and recipient can be justified as a long-run outcome, given that population growth rates are falling in the developing world. We assume that the long-run population growth rate is 1.5% a year. This is approximately the average rate over the last decade in the Southern group of countries, but somewhat higher than in the North over the same period. Under these assumptions, the North begins the simulation with capital per effective worker about 10% below its steady-state value (see the ‘N’ column of Table 1). A convenient way of interpreting this function is to ask when aid intensity at/YSt is sufficiently high that the marginal benefit of aid is zero for the South. This happens when at/YSt = υ/2. The appropriate severity of this absorption constraint is a matter of debate. In our baseline, we set υ so that an extra dollar of aid has zero impact when the aid/GDP ratio is 25%.14 A final choice relates to ω, the relative weight of Southern utility in the objective function of the North. The use of such a weight seems essential to any normative study of transfers between the North and South. We are not seeking to estimate the weight that Northern citizens currently adopt, nor to defend a particular choice of ω on prior grounds. Instead, our framework provides a mapping between assumptions on ω and optimal aid generosity. A possible parallel would be with Kopczuk et al. (2005), who build on the earlier work of Mirrlees (1971): these papers allow a mapping between alternative degrees of inequality aversion and tax rate schedules. In our simulations, we choose a baseline weight ω = 0.1 to represent a donor that is genuinely altruistic, but imperfectly so. We will investigate later how optimal aid changes when ω takes higher values. As we will see, even the choice of ω = 0.1 — so that Northern citizens have ten times the welfare weight of Southern citizens — can lead to shares of aid in Northern GDP substantially higher than in the data.","In this section, we study the optimal time path for aid using simulations. To do this, we use the relaxation algorithm of Trimborn et al. (2008) to solve the system of differential equations implied by the optimal control solution. In our baseline case, aid generosity should be highest at the beginning, so that the South accelerates towards its steady- state. But this front-loading result is sensitive to absorption constraints. There are scenarios in which the North should increase the generosity of its aid (relative to its GDP) as the South grows, since the South can then absorb a given level of aid more effectively. The first panel of Fig. 1 shows the optimal path of aid in our baseline case. Aid relative to Northern GDP is initially 7.1%, falling to 4% after 17 years and to 2.4% in steady-state. Aid flows on this scale are clearly too high to be realistic, and were not seen even under the Marshall Plan.15 Importantly, however, long-run generosity is highly sensitive to the assumed curvature of the utility function. If we reduce σ by just 10%, to σ = 1.8, the North will continue to be generous early on, but much less so at longer horizons. Hence, diminishing returns to consumption can provide a powerful motivation for international transfers, but at longer horizons, parameter assumptions matter a great deal. Moreover, it is likely that the North places even less weight on the utility of the South than we are currently assuming. There are other possible reasons for the divergence between the model and the flows observed in practice. We do not model the marginal cost of public funds: aid is financed by lump-sum taxation rather than distortionary taxes. Nor do we incorporate any political economy constraints on the donor, such as taxpayer resistance to large international transfers. Hence, natural extensions to our model would yield smaller ratios of aid to Northern GDP for given values of ω and σ. Aid intensity in the South is initially 16.8%, falling to 3.3% in the long run. Fig. 1 also shows the effect of aid on the rate at which Southern output and consumption converge towards Northern levels, and the effect of aid on the Southern saving rate (expressed as the deviation from a zero-aid counterfactual). The front-loading of aid enables the South to raise consumption by 14.4% immediately. The effect of aid on growth is familiar from Obstfeld (1999): there is an initial, but modest, acceleration which is eventually followed by slower growth relative to the autarky counterfactual. Growth is ultimately slower because aid brings capital accumulation forward, so that capital is accumulated quickly early on and more slowly later. Even though aid intensity is high in this baseline calibration, the effect on output growth is small. In annual terms, the initial growth rate is 0.57 percentage points higher, with the difference eliminated after 14 years; roughly 25 years later, the growth rate is 0.08 percentage points lower than it would have been without aid. It may seem surprising that the growth and welfare effects of aid are not larger. It is clear that, with isoelastic utility, exogenous transfers of this magnitude have only limited effects on optimal investment. The explanation is that, when the benefits of investment are high, it will be undertaken even in the absence of transfers, because forgoing consumption is not costly under these preferences. Hence Obstfeld (1999) finds modest growth effects, and he notes (p. 136) that the result is ‘likely to be a robust feature of any plausible model in which aid is funneled through the private sector’.16 More broadly, our results on welfare effects are also consistent with the work of Gourinchas and Jeanne (2006) on capital mobility. They showed that the welfare benefits of capital inflows in a calibrated Ramsey model are unexpectedly modest. This is partly because convergence will be rapid even in the absence of foreign capital, and partly because accumulating capital more rapidly brings forward a reduction in its marginal product. This intuition is helpful in understanding why the effects of aid on productivity are also relatively modest in our setting. But in the case of aid, we can take this logic further: if relaxing the South's resource constraint is not all that valuable to the South, it is likely that tightening the North's resource constraint is not all that costly to the North. This is why we find large transfers to be optimal, even though the productivity benefits of a given transfer are limited. Scenarios in which aid increases over time are shown in Figs. 2 and 3. For some combinations of ω (the weight on the South's utility) and υ (the aid absorption parameter) it may be optimal for the North to back-load aid: as the South develops, it becomes better able to use aid effectively, and the North should increase the generosity of aid relative to its GDP. What happens if the North gives a fixed share of its GDP, as in the Pearson Commission 0.70% benchmark? When the fixed share is chosen optimally, the welfare losses for the South relative to the fully-flexible optimal aid policy are modest. Perhaps counter-intuitively, if the North has to donate a fixed share of GDP, the costs of this restriction are sometimes borne by the North in equilibrium: the North becomes more generous in the long run than it otherwise would, and the South sometimes does better than it otherwise would.17 We also find that, if aid is fixed as a share of Northern GDP, departures from the optimal choice of this share are not particularly costly: the donor's objective function is relatively flat as a function of the aid share. Fig. 4 compares outcomes when the North can vary aid as a share of its GDP, and when the share is fixed. By construction, the donor must always do at least as well when free to choose a flexible path for aid, but the effect of fixing the share on the donor's objective is surprisingly modest. In our baseline case, the associated utility loss is equivalent to an ongoing reduction in consumption of 0.2% for the North and South. Looking at the two economies individually, the restriction to a fixed share leaves Northern households worse off (− 0.8 %) and Southern households better off (+ 0.6 %). But this is not always the case. For some cases with higher ω, the restriction to a fixed share leaves the North better off and the South worse off. The role of ω hints that these results arise from the interplay of the benefits of aid at short horizons, the benefits at long horizons, and whether fixing the aid share is wasteful due to absorption constraints. Finally, we limit aid to the Pearson Commission benchmark, 0.70 % of Northern GDP. Compared to our optimal flexible path, the much lower generosity means that the HEV welfare gain in the South falls from 11% to 2%.","We now consider preferences with a subsistence level of consumption. It is well known that the single-sector Ramsey model with isoelastic preferences makes some unrealistic predictions: fast convergence to the steady-state, and a sharp decline in investment rates and in the marginal product of capital as capital is accumulated. Nor can it readily accommodate periods of international divergence in GDP per capita, or explain why investment rates are sometimes low in poorer countries. Ben-David (1998) and Steger (2009) argue that Stone–Geary (SG) preferences overcome these problems. Low investment can co- exist with high returns to investment, because the opportunity cost of investment is high when households are close to subsistence. We examine the implications for the generosity and timing of aid, including whether donors will want to front-load aid. As expected, under SG preferences, aid generosity should be high early in the development process. We present some results in Fig. 5, which could be compared with the earlier CRRA results in Fig. 1. The optimal level of aid is initially higher under SG preferences, at 9.6% of Northern GDP. It declines to the same asymptotic level of 2.4% (recall that the CRRA and SG growth paths are asymptotically equivalent) but converges to this level more slowly in the SG case. Under SG preferences, Southern output converges more slowly towards the Northern level. This is a natural result: when the Southern economy is close to subsistence, the opportunity cost of investment is high. But the optimal time path for aid helps to close the gap between North and South to a greater extent under SG preferences than under CRRA preferences. This is another natural result, since relaxing the intertemporal resource constraint is more valuable in the SG case. The effect of SG preferences on the generosity of aid means that aid has a greater impact on consumption in the SG case. The peak change in consumption, relative to autarky, is similar (17% under SG, 14% under CRRA) but the duration of the effect on consumption differs greatly. Under SG preferences, Southern consumption is 11% higher than the zero-aid counterfactual 50 years after the commencement of aid. Under CRRA preferences, the equivalent figure is just 6%. This difference reflects the greater generosity of aid in the SG case: aid is 5.5% of Northern GDP at t = 50 under SG preferences, but just 2.5% under CRRA. Although the absolute change in the initial annual growth rate, relative to the zero-aid counterfactual, is similar to the CRRA case at 0.59 percentage points, autarky output growth is slower under SG preferences, at 3.5% per year.18 So the impact of aid on growth with SG preferences is more substantial, multiplying the initial growth rate by 1.17, compared with a multiple of 1.08 under CRRA preferences. To compare the SG and CRRA cases in more detail, Fig. 6 shows outcomes based on the same fixed quantity of aid.19 The higher aid intensity in the SG case reflects the fact that Southern output grows more slowly under these preferences. The effects of aid on net investment, consumption and growth are markedly greater in the SG case. Since our assumptions on preferences vary across the two cases, more direct welfare comparisons are not meaningful.","In this section, we study the donor's problem when there are two aid recipients. Additional recipients can be integrated into the Northern planner's optimal control problem in the obvious way. With multiple recipients, allocation decisions are connected because aid to one country reduces consumption in the North, increasing the opportunity cost of aid to other countries. We construct a low-income recipient, S1, using the countries in the lowest quarter of the GDP per capita distribution. The other recipient, S2, is an aggregate of ‘middle income’ countries, defined as those between the lower and upper quartile of the GDP per capita distribution. Table 1 lists some of the relevant numbers. We now include China and India, unlike in our baseline case. As a result, the middle-income country is ‘large’: the population of S2 is 6.1 times greater than the population of S1. Output per capita in S2 is about 19% that of the North, whereas the poor recipient, S1, has output per capita only 5% that of the North. Hence, in this first experiment, output per capita in S2 is about four times greater than in S1. The long-run ratio of output per capita between the two is smaller, however, at 2.5. This is because the initial capital stock in S1 is further beneath its long-run level than S2. The ratio k0/k∗ is 0.13 for S1 and 0.35 for S2. Hence, without aid, S1 initially grows faster than S2: annualized growth is 13% for t = 0 in S1, 6% in S2. This calibration can address an important policy issue: does low consumption help to motivate aid even for middle-income countries? Fig. 7 illustrates the optimal aid allocation under CRRA preferences. The contrast between the two recipients is striking. A substantial amount of aid is allocated to the middle-income recipient, and its response is familiar from our baseline case. But despite the poor recipient S1 starting further beneath its balanced growth path, and with greater scope for aid-induced growth, the effect on investment (which mirrors the effect on output convergence) is negligible. Instead of a front-loaded aid path that induces accelerated investment, the donor keeps aid intensity high throughout and increases generosity in absolute terms over time. This reflects the aid absorption constraint: the donor cannot be too generous to S1 initially, because the optimal quantity of aid already takes this economy close to the point where the marginal benefit of aid is zero. The recipient S1 uses aid to fund a higher level of consumption, with little change in investment behavior, because this recipient correctly anticipates that aid will increase over time. As for the middle-income recipient S2, where the population is larger and the absorption constraint is less binding, the donor is generous early on but then scales aid back over time as the recipient grows. This is reflected in the shares of the two recipients in total aid: the middle-income recipient receives the bulk of the aid at shorter horizons, and the low-income recipient S1 at longer horizons. To explore further, we change the population of S2 so that it matches that of S1. Fig. 8 shows what happens: there is a sharp reduction in the quantity of aid given to S2, but qualitatively the paths are unchanged. The level of aid given to S1 is virtually unchanged, despite the large reduction in aid given to S2 with an associated reduction in the opportunity cost of aid to the North. This result arises because S1 is already being given as much aid as it can effectively absorb.20","In this section, we carry out some sensitivity analyses, and then discuss possible extensions. First, recall the externality that arises in our setup: capital accumulation by Southern households increases the ability of the South to absorb aid. What happens if we take our baseline calibration and assume Southern decisions are made by a social planner? The optimal policy for the Northern planner now leaves Northern households slightly worse off (equilibrium aid is slightly more generous) and the South slightly better off. The effects on the North are modest and the optimal path for aid barely differs, compared to the decentralized version. The peak difference in aid between the two cases, as a percentage of Northern GDP, is only 0.7 percentage points. But the change in Southern consumption is noticeable: the Southern planner anticipates the favorable effect of investment on future aid absorption, and chooses higher investment rates in the first few years of the transitional dynamics. Overall, however, the quantitative importance of the externality is modest. We discuss these results in more detail in an online appendix. Next, we briefly consider the effects of varying two parameters: σ, the curvature of the utility function, and α, the output–capital elasticity. In modifying these parameters, we still require the Northern and Southern capital–output ratios to match those in the data. Relative to our previous experiments, there is a direct effect of changing each parameter, and an indirect effect which arises from matching the observed capital–output ratios under new parameter assumptions. The alternative route, of only focusing on the direct effect, would lead to an inconsistency between the capital–output ratios in the simulations and those in the data. As noted earlier, the curvature of the utility function has a major effect on optimal generosity, especially at longer horizons. When the marginal utility of consumption diminishes less rapidly with consumption, the motivation for global redistribution is weakened. In our simulations, with slightly less curvature (σ = 1.8 rather than σ = 2.0) aid generosity is reduced and the effect on consumption is smaller. Given technical progress, the level of the balanced growth path is raised by a reduction in σ, and the effect of aid on growth is initially stronger than in our baseline. This effect is modest, however. Increasing the output–capital elasticity has similarly complicated effects. It makes aid more effective, since capital accumulation becomes more important; it also raises the balanced growth path. We do not present the results in detail, but note that as α increases, the effect of aid on investment is lower but the effect on growth larger. Intuitively, because investment now has greater benefits, the South is able to consume more early on. We now briefly discuss some possible extensions. Perhaps the most important modification to the analysis would be to introduce political economy considerations within the South. This would introduce a major complication, namely strategic interactions between the donor and recipients. One possibility would be to allow aid to finance public investment, as in Chatterjee et al. (2003) and other work summarized in Turnovsky (2009); see also Lowe et al. (2013). Expenditure on public investment projects could be wasteful to some degree, with the extent of waste increasing in expenditure as in Berg et al. (2013). Politics in the North could also play a role: in principle, the North's objective function could reflect the North seeking political influence, as in the work of Antràs and Padró i Miquel (2011). Another extension could allow aid recipients to borrow internationally, with a risk premium that is increasing in external indebtedness. van der Ploeg and Venables (2013) examine the optimal response of recipient governments to windfall revenues in such an environment, but treat the time path of revenues as exogenous, which seems better suited for natural resource windfalls than aid flows. Perhaps their decision problem could be integrated within that of an altruistic North, using the nested structure we adopt here. A related extension would model recipients as two-sector economies that produce traded and non-traded goods, in which case absorption constraints could arise endogenously through Dutch Disease effects. Staying closer to our current framework, other possibilities include country-specific population and technical progress dynamics, costs of distortionary taxation in the donor, adjustment costs in the North for aid disbursements, finite-duration aid commitments by the donor, CES production technologies, and the introduction of capital varieties; see Hoxha et al. (2013) on the latter. A more ambitious extension would introduce output volatility in the South, so that aid would have an additional insurance role. This would bring the analysis closer to Arellano et al. (2009), at the expense of greater computational complexity.","This paper has introduced a framework for analyzing optimal aid policies using the neoclassical growth model. In contrast to most previous research, the findings are based on a clearly-defined optimization problem for the Northern donor. It takes into account the opportunity cost of aid, optimizing behavior by Southern households, and constraints on the effective absorption of aid. This framing of the problem is a first step in the direction of a richer quantitative analysis, and the simulations indicate what might be learnt from future exercises of this type. Since any tractable model must be stylized, there will always be a need for subjective judgments on the part of donors; but formal models can make visible some considerations and possibilities that would otherwise be obscure. We find that optimal transfers are influenced by the weight a donor places on recipient welfare, the curvature of the utility function, the recipient's capacity to absorb aid, the relative level of the recipient's balanced growth path, and the recipient's initial distance from that growth path. In our simulations, the scope for aid to raise growth rates plays some part in aid decisions, but the effect on the level of consumption often dominates. The optimal generosity of aid, relative to Northern GDP, should vary over time to reflect conditions in aid recipients. In our baseline case, aid generosity declines over time. But in some scenarios, where donors run up against absorption constraints, there is a time interval over which generosity should be increasing. As for optimal aid intensity — that is, aid relative to Southern GDP — this should generally decline over time. But in cases where the aid absorption constraint is close to binding, the optimal policy may dictate that high aid intensity is maintained for a long time. In the case of multiple recipients, the optimal policy may involve large changes over time in the absolute quantity of aid and the division of aid between recipients. As expected, the case for donor generosity is strengthened when consumption is close to subsistence. The effects of aid on Southern consumption, investment and growth can be dramatic in this case. In general, our numerical results indicate optimal aid levels that exceed those in the data, at least at short horizons. An interesting task for future research is to pin down the features of reality that could explain this disparity. That citizens of donor economies may place little weight upon the welfare of aid recipients is only one possible explanation; others could include the potential roles of political economy forces and corruption in undermining the effectiveness of aid. Given the obvious importance of these considerations, our approach is a partial view. Nevertheless, it provides some insights and results that could help to inform the design of aid policies, and is simple enough to be extended in many directions."],["We study the impact of a supply management mechanism (SMM) similar to the Market Stability Reserve proposed in 2015 which preserve the overall emissions cap and we comment on the recent cap-changing amendments. We provide an analytical description of the conditions under which an SMM alters the emissions abatement paths, affecting the expected length of the banking period and its variability. While abatement strategies of risk neutral firms solely depend on the former, for risk-averse firms changes in the latter would lead to higher risk premia, accelerated depletion of the bank and, consequently, further reduction of abatement and allowance prices. Cancellation of part of the reserve could partially outweigh the effect on risk premia sustaining allowance prices. --------------------------------------------------------------------------------","Despite an emerging use of supply control mechanisms, in most existing cap-and-trade programmes the environmental reduction target (the cap) is fixed and the supply of allowances is inflexible and determined within a rigid allocation programme. In theory, as long as the regulator makes allowances available before they are needed, the programme will deliver a cost-effective solution Hasegawa and Salant (2015). However, observations from recent cap-and-trade schemes – in particular the European Union Emissions Trading System (EU ETS) – have raised concerns over excessive allowance price variability and price collapse. These maladies seem to stem from a problem of ‘over-supply’, wherein unexpectedly low levels of allowance demand have led to the accumulation of a significant surplus of allowances. An article in The Economist (2013) lamented a surplus of allowances equivalent to an average year’s emissions. This surplus is often attributed to two effects. On the one hand, the economic recession and renewables-promoting policies have led to a significant drop in allowance demand; on the other, the system has been unable to respond to changes in economic circumstances and policies, see Grosjean et al. (2014) and Ellerman et al. (2015). The resultant drop in allowance prices has policy makers and other stakeholders concerned that the current imbalance in supply and demand, if left unchecked, could reduce incentives for low-carbon investment and ultimately impair the ability of the EU ETS to meet its targets.1 There are already provisions within a cap-and-trade framework that, in theory, should compensate for unforeseen changes in allowance demand. For example, most ETSs have banking provisions that should provide firms with a tool to respond to demand shocks. Several studies have explored the effect of banking and borrowing provisions as cost ‘smoothing’ mechanisms which decrease allowance price variability; Hasegawa and Salant (2014) provide a comprehensive and critical review of the literature on bankable emissions allowances that has developed over the last two decades. Other studies demonstrate how hybrid systems, combinations of quantity- and price-based instruments, lower expected control costs ultimately mitigating allowance price variability (Fell and Morgenstern, 2010; Grüll and Taschini, 2011; Fell et al., 2012b; 2012a). However, these provisions alone may not be sufficient when the market is faced with severe demand shocks. This leads to the question of how to amend an existing ETS to deal with an unexpected under- or over-supply of allowances. Namely, how should the allowance allocation programme (the supply, which can be controlled by regulators) be changed to better cope with unexpected changes in allowance demand. In the case of the EU ETS, the European Commission (EC) has proposed a structural reform of the ETS, including the implementation of the Market Stability Reserve (MSR) that has started to operate in 2019 (EC, 2014a; 2014b; EP, 2015). The MSR amends the allowance allocation programme. In particular, it adjusts the number of allowances auctioned based on the size of the aggregate bank, i.e. the sum of firms’ individually held banks of allowances. Hereafter, we refer to quantities concerning the entirety of regulated firms as ‘aggregate’ whereas the respective quantities for each firm are referred to as ‘individual’. In a given year, if the aggregate bank of allowances exceeds 833 million, a pre-defined percentage of the size of the aggregate bank will be withheld from auctions and will be placed in a dedicated reserve. There are two intake rates: 24% from 2019 to 2023 and 12% from 2024. These allowances are returned to the market in batches of 100 million as soon as the aggregate bank drops below the threshold of 400 million. In its original 2015 design, the MSR changes the allowance allocation programme but leaves the total number of allocated allowances (the cap) unchanged within the regulatory period. As such, the reserve is temporary in nature and the initially proposed version of the MSR preserves the original cap. The alternatives of temporarily versus permanently placing allowances in the reserve have been heartily debated in the past years. In late 2017, after numerous stakeholder consultations and more than two years of negotiations, the European Commission decided that, starting in 2023 the volume of allowances that can be held in the reserve will be capped at the previous year’s auction volume. The resulting difference in the reserve will be cancelled, providing a mechanism for allowances to be retired and thus reduce the long- run supply of allowances. In an earlier paper (Kollenberg and Taschini, 2016), we examined a similar dynamic allocation programme where the cap could be varied in response to exogenous shocks, as is the case for the 2017 version of the MSR. Accordingly, the applicability of this earlier framework to the original 2015 MSR is limited. With the additional objective to comment on the recent proposed MSR amendments (allowances cancellation), we focus our analysis on a generalised, cap-preserving supply management mechanism (SMM for short) similar to the original 2015 MSR legislation. The proposed SMM allows us to abstract from the operational details of the EC MSR2 and to provide a conceptual framework that enable us to transparently illustrate (1) how firms’ abatement strategies vary in response to changes of the allowance allocation programme and (2) how an SMM affects the risk-premium associated to holding allowances or any equivalent investment in abatement. As such, we draw from and contribute to the literature on inter- temporal permit trading under uncertainty and to the emerging literature on the assessment of mechanisms that vary allowance allocation according to market conditions, such as indexed regulations (among others, Newell and Pizer, 2008; Kollenberg and Taschini, 2016; Lintunen and Kuusela, 2018), output-based allocation (Meunier et al., 2017), and price- based mechanisms (Aldy et al., 2017). In particular, the implications of the changes in risk-premia speak directly to the policy debate playing out among experts on no-cap adjustments versus cap adjustments. Our analysis ultimately suggests that a permanent cancellation of part of the reserve could keep in check the premium that risk-averse firms demand for abatement investments. As a by-product, the relevance of our findings extends to the policy debate in California, South Korea and the member States of the Regional Greenhouse Gas Initiative, where similar supply management mechanisms were adopted. The results of previous theoretical and empirical analyses of intertemporal trading of emission allowances reveal that, under the usual assumption that marginal abatement costs are increasing in emissions reduction, firms start accumulating allowances and then draw them down, see Rubin (1996), Schennach (2000), Ellerman and Montero (2007), and Ellerman et al. (2015). Banking of allowances is thus a manifestation of the inter-temporal trading problem. The rationale for banking is quite intuitive: if tomorrow’s discounted expected cost is higher than today’s cost, it is worth banking allowances, whether obtained by abating more emissions today or by purchase, and either using them to cover some of tomorrow’s emissions or selling them later on. The expected duration of the banking period, i.e. the period of time during which firms prefer to hold allowances, depends on the amount of abatement implied by the cap and, as long as the original abatement path is feasible (see Perino and Willner, 2016), it is independent of the allowance allocation programme. In the analysis that follows, we explore the impact of an SMM on firms’ abatement strategies using a model of the inter-temporal pollution control and allowance trading. We consider the inter-temporal optimisation problem of each entity in a continuum of small regulated firms. At each point in time, each firm has to decide by how much she wants to offset her individual emissions, considering current and future costs of reducing emissions, as well as her existing individual bank of allowances and future allowance demand and allocations. The chief decision state variable is the firm’s expected required individual abatement, the difference between counterfactual emissions (individual cumulative emissions in the absence of emissions restrictions) and the number of allowances individually allocated. Every firm adjusts her abatement and trading strategies at each time t based upon this state variable, taking into account her current bank of allowances and any change in the required abatement. Under uncertainty, changes in firm’s expectation about the required individual abatement affect how much individual abatement and banking will occur in the future – and for how long. We thus frame our analysis of the impact of a cap-preserving mechanism that amends the allowance allocation programme, similar to the 2015 version of the MSR, in terms of two main state variables: firms’ expectations about the required abatement and the length of the banking period. Previous studies have demonstrated that firms’ strategy adjustments and the overall efficacy of the 2015 MSR are highly dependent on the constraints on temporal provisions (i.e. limitations on borrowing) and on the design of the mechanism implemented to adjust the allowance allocation, see Salant (2016), Perino and Willner (2016), Fell (2016). In the absence of borrowing constraints, abatement decisions are independent of the temporal distribution of allowances. If firms can always borrow from future allocations, any change to the allocation programme that maintains the overall emissions cap is irrelevant. Firms will simply borrow the required allowances needed to remain on their original cost-minimised emissions path and the adjustments of the allowance allocation programme will have no influence. Under borrowing constraints, a change in the allocation programme can affect abatement and allowance price paths only when the amount of allowances presently available to firms to cover emissions is insufficient. This is the availability condition in Salant (2016) or the feasibility condition in Perino and Willner (2016). Our analytical results are consistent with these results: an SMM can only change abatement and allowance price paths if and only if the onset of the SMM changes the expected required abatement, i.e. the expected future net demand of allowances. Specifically, abatement strategies are unaltered when neither the expectation about the length of the banking period (τ) nor the post-SMM expected required abatement change. Conversely, when the adjustments in the allowance allocation programme determined by an SMM affect the expected required abatement, the expected length of the banking period τ and its distribution vary.3 When considering the impact of previously unexpected changes (e.g. demand shocks), we note that changes to the timing of allowance allocation can affect the instantaneous likelihood of the event of an instantaneous depletion of the bank. We term this instantaneous breakdown. That is to say, changes to the distribution of τ by an SMM can change the probability that firms are not able to compensate for a demand shock with their current individual bank of allowances. This is related to the discussion of price variability in Perino and Willner’s analysis. They show that the short-term scarcity produced by the (binding) MSR can drive prices up and increase price volatility when allowances are removed from the market. Our findings support the conclusion that a cap-preserving supply control mechanism increases price variability overall. Crucially, changes in price volatility due to an SMM are immaterial for risk neutral firms. Their abatement strategies solely depend on the expected required abatement. However, for risk-averse firms, differences in price volatility matter and should be reflected in the risk premium demanded by those firms for holding allowances or for investing in abatement. Thus, we expand our analysis to risk- averse firms and show that changes in the probabilistic distribution of τ brought on by an SMM that lead to higher price variability (compared to no-SMM) generate higher risk premia. The higher the risk premium, the more quickly firms will deplete their bank, which leads to lower levels of abatement and lower prices. However, abatement and allowance prices are affected to a lower extent during different periods of the bank. Thus, compared to the no-SMM case, the consequence of higher price variability are more compelling when regulated firms are not perfectly risk-neutral. This could have significant implications for the overall impact of an SMM like the MSR proposed in 2015. While one of the goals some stakeholders attributed to the original MSR was to increase prices during periods of over-supply, the building up of the allowance reserve by a cap-preserving mechanism would have the opposite effect. When the behaviour of risk-averse firms is taken into account, the impact of an SMM is more striking: the rise in price volatility would lead to higher risk premia, accelerated depletion of the bank and, consequently, abatement and prices are reduced even further. Cancellation of part of the reserve could partially outweigh the effect on risk premia and sustain allowance prices. The remainder of the paper is organised as follows. In Sections 2.0 and 2.1 we describe the model assumptions and define the key decision making variables for each of the agents on the allowance market. In Section 2.2 we present the market equilibrium in terms of aggregate quantities and provide an analytical description of the conditions under which an SMM alters the emissions abatement paths. In Section 2.3 we relax the assumption of risk neutrality and explore the effect of an SMM on a time-dependent risk premium. Section 3 concludes.","Regulated firms are assumed to be atomistic in a perfectly competitive market for emission allowances. Firms face an inter-temporal optimisation problem where, at each point in time, they have to decide how much they want to offset their emissions (either by abating or by trading allowances), considering the current and future costs of reducing emissions. Each firm accounts for her current individual bank of allowances and the number of allowances she expects to be allotted in the future. In this context, the required abatement, the difference between the cumulated individual amount of emissions without abatement requirements (counterfactual individual emissions) and their future allocation, is the key quantity each firm has to assess at each point in time. Under uncertainty, changes in a firm’s expectation about the required abatement affect how much abatement and banking will occur in the future – and for how long. Crucially, the impact of these changes is relevant only during the banking period. Once the bank is depleted, the inter- temporal problem breaks down: each firm uses every allowance available to cover contemporaneous individual emissions and instantaneously abates her residual individual emissions (Schennach, 2000).4 Thus, we focus our analysis on the banking period [0, τ) and investigate under which conditions an SMM can alter the length of the banking period τ and its probabilistic distribution. Required abatement under uncertainty ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To capture the impact of uncertainty on banking in a cap-and-trade programme under the SMM, we identify two key state variables of the system: the time-t expectations of (i) the instant τ when the aggregate bank is completely depleted and (ii) the corresponding required aggregate abatement, that is counterfactual emissions over [0, τ) minus the total number of allowances allocated in the same period (including the initial aggregate bank of allowances). When new information becomes available, firms update their expectations and adjust their strategies. That is, abatement and trading strategies are adapted at each time t, taking into account the current aggregate bank of allowances and the change in the required aggregate abatement. Equilibrium solution for risk-neutral firms ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The following analysis extends the efforts of these authors by studying how a cap- preserving supply control mechanism affects price volatility and – under risk aversion – the risk premium associated to the instant when firms prefer to deploy their bank. The rationale is the following: under risk aversion the impact of an SMM on price volatility is reflected in the risk premium and, consequently, in firms’ discount rate. The latter signals whether returns from allowance-related investments should promise higher or lower returns with consequent effects on allowance banking. Changes in expectations: a dynamic view and risk-aversion ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ During the building up of the allowance reserve, the change in allowance allocation increases the likelihood of an instantaneous breakdown. This is reflected in an adjustment of the risk-premium qt. More precisely, Eq. (6) reveals that a change in gt generates a change in σt which equally transfers to a change in qt. Consequently, risk-averse firms adjust their abatement behaviour, as quantified in Eq. (5). As previously discussed, the impact of qt on abatement is larger in the short term, when the expected end to the banking period lies in the distant future and the SMM is removing allowances from the market. Thus, in the short run the potential losses associated to abatement investments due to an instantaneous breakdown are high. As time goes by, the expected time τ of complete depletion of the aggregate bank approaches and allowances from the reserve are released. As we can see from Eq. (5), the reduction in qt is determined by the combined effect. Abatement and allowance prices increase, but to a lower extent than both were decreased earlier during the building up of the reserve. In conclusion, we find that under risk aversion, the changes brought on by the SMM to the distribution of the instant when firms prefer to use their entire aggregate bank, lead to higher price variability – compared to the risk-neutral case – and, consequently, higher risk premia. However, we note that the impact of a change in the likelihood of an instantaneous breakdown affects prices and abatement more in the short run than it does in the long run. These findings corroborate concerns about rising price volatility, as raised by Perino and Willner (2016) and Fell (2016).","The supply of allowances in the European Union Emissions Trading System (EU ETS) has been inflexible and determined within a rigid allocation programme. As such, the system lacked provisions to address severe imbalances in demand and supply of allowances resulting from economic shocks. In 2015 the European Commission proposed a structural reform of the EU ETS, including the implementation of a Market Stability Reserve (MSR), operative since 2019. The MSR will adjust the allowance allocation programme based on the aggregate bank of allowances: In times of a large bank, allowances are transferred to a dedicated reserve to be released in times of scarcity. In its original 2015 design, the MSR preserves the total number of allowances issued over the regulatory phase. After two years of negotiations and an extensive impact assessment, the European Commission decided that allowances held in the reserve above the previous year’s auction volume will no longer be valid. The findings of our work support the decision for regular cancellation of excess allowances. We develop a stochastic equilibrium model of inter-temporal trading of emission allowances to investigate under which conditions a supply management mechanism (SMM) similar to that proposed for the EU ETS can alter allowance price and emissions abatement paths. Similar mechanisms were adopted in California and South Korea. We show that the timing of allocation is largely irrelevant as long as changes in expected net demand of allowances are such that the resulting bank remains essentially unaltered. Conversely, when the transitory scarcity brought on by the SMM changes the net allowance demand, the mechanism affects the expected abatement and the price paths. In this context, we consider unexpected changes in firms’ expectations that triggers an instantaneous depletion of the bank of allowances (what we termed unexpected breakdown). Risk neutral firms are indifferent to changes in the variability of this event. However, when firms account for the risk in the change of the variability of their future required abatement – i.e. counterfactual emissions minus total number of allowances allocated over the same period – and equivalently, risk in the variability of the value of their abatement investments, adjustments in the allowance allocation programme matter. We then expand our analysis to study how risk-averse firms’ strategies are affected by an SMM at different points in time of the banking period. We show that changes in the distribution of the time of the unexpected breakdown brought on by the SMM lead to higher price variability and, consequently, higher risk premia. The higher the risk premium associated with holding allowances, the more quickly firms will deplete their bank, which is associated with lower levels of abatement and, importantly, lower allowance prices. This has clear policy implications for the current debate on cap adjustments vs. no-cap adjustments: the influence of a generalized, cap-preserving supply management mechanism like the 2015 version of the MSR could be counter-productive, especially when the behaviour of risk- averse firms is considered. Importantly, while increased price variability in the short run may prevail even under the amended MSR, (the anticipation of) a permanent cancellation of part of the reserve will, at the very least, lead to lower risk of low-carbon investments (such as purchase of allowances) and, accordingly, higher prices in the short run with lower but less risky long-run returns. The late 2018 increase in allowances prices to almost three times its value since the amendment of the MSR may well be attributed to a market perception of such decreased risk."],["We use data from six cohorts of university graduates in Germany to assess the extent of gender gaps in college and labor market performance twelve to eighteen months after graduation. Men and women enter college in roughly equal numbers, but more women than men complete their degrees. Women enter college with slightly better high school grades, but women leave university with slightly lower marks. Immediately following university completion, male and female full-timers work a very similar number of hours per week, but men earn more than women across the pay distribution, with an unadjusted gender gap in full-time monthly earnings of about 20 log points on average. Including a large set of controls reduces the gap to 5–10 log points. The single most important proximate factor that explains the gap is field of study at university. --------------------------------------------------------------------------------","Since the influential survey by Altonji and Blank (1999), the economic literature has offered a variety of new explanations about gender gaps. Bertrand (2010) provides an insightful review of recent contributions, drawing on advances in the psychology and experimental literatures. Her review emphasizes the importance of gender differences in risk preferences, attitudes toward competition and negotiation, and the strength of other- regarding preferences as well as the importance of social norms that may induce differential sorting of men and women across occupations.1 In another recent survey looking at a large sample of high-income countries, Olivetti and Petrongolo (2016) stress the role played by changes in the industry structure with the shift from manufacturing to services, which might have increased female employment and reduced (but not eliminated) the gender wage gap. We make use of a unique data source which allows us to analyze a representative survey of six cohorts of university graduates between 1989 and 2009 from all fields of study and across the range of Higher Education institutions in Germany, soon after they complete their college education. The survey also follows the same individuals five to six years into their careers. In this paper we focus only on the first snapshot after graduation. The reason is that career opportunities within the firm (which may involve firm-specific investment, or ability to negotiate with employers, or internal promotions) as well as other crucial family-related decisions (such as starting a family and committing resource allocations within the household) are likely to be less relevant than later on in life.4 Our results indicate that 12–18 months after graduation, the raw (unadjusted) gender gap in full-time monthly earnings is about 20 log points on average, even though male and female full-timers work relatively similar hours. Including a large set of controls reduces (but does not eliminate) the gap to 5–10 log points, with the lion share in the reduction being accounted for by field of study. In light of this, assessing the extent of gender gaps and deepening our understanding of their nature within a human capital paradigm seem important steps, before looking for alternative explanations related, for example, to psychological cues, attitudes, social norms, and biomarkers. Our analysis provides new evidence for Germany on a variety of facets of gender gaps. We consider two broad sets of outcomes. The first refers to the educational performance of men and women while they are at university and explores gender differentials in enrolment rates, marks at university entry (i.e., secondary school grades), graduation rates and graduation marks. The second set refers to early labor market performance and focuses on gender differences in the probability of having a full-time job and in pay within 12–18 months after graduation. Existing analyses of gender gaps in Germany focus on wage differentials among all workers and not just university graduates, while gender differences in university attainments have not yet been fully explored.5 Using social security data, Fitzenberger and Wunderlich (2002) document large and persistent gaps between 1975 and 1995: they estimate that West German full-time female employees earned about 35% less than their male counterparts at the beginning of that period, and the gap went only slightly down to about 25% at the end. These figures refer to all workers and may not apply to the most educated workforce, where censoring in the administrative data occurs frequently. In a more recent assessment with data from the 1996 Labor Force Survey, Machin and Puhani (2003) find that the (raw) gender wage difference among university graduates in Germany is about 28 log points, and about 40% of the explained gap can be accounted for by field of study. By looking at graduates across all ages, however, these differences will in part reflect choices and constraints that emerge well after graduation. Using data from the European Community Household Panel over the 1995–2001 period, Arulampalam et al. (2007) observe an unconditional average wage gap of 20%. Conditional on covariates, they find an increasing profile of the gap along the wage distribution ranging from 6 to 17% among public sector workers and from 14% to 20% among workers in the private sector. Analyzing administrative data, Huffman et al. (2017) confirm the large gaps found in the earlier studies and show that roughly half of the 29% raw gap in 2008 is accounted for by observables, including age, schooling, establishment size, collective bargaining status, and industry dummies.6 A number of institutional features specific to Germany may potentially not be gender neutral and may thus influence our outcomes of interest.7 One of such features is that Germany tracks student into different schools based on ability as early as the end of elementary school (ages 10–12). There is evidence that early tracking has no lasting effect on later wages, employment, and occupational choice (Dustmann et al., 2017). However tracking may give differential penalties or advantages by gender (e.g., Pekkarinen, 2008 for Finland). For example, if girls mature earlier than boys during secondary school years, they could perform better in school irrespective of the track chosen, and this in turn could affect their relative performance later on, at university and in the labor market. Alternatively, more competitive boys might disproportionately attend the top school tracks, and this might in turn affect performance at university or in better-paid jobs. Our analysis will assess the extent of differential performance in secondary school among university graduates and consider whether this in turn has an impact on outcomes. Another feature is that, since the late 1970s, Germany has also developed (and provided working mothers with) generous maternity leave coverage, whereby mothers are currently eligible for three years of partially paid leave. Analyzing the impact of five major expansions in maternity leave coverage occurred between 1979 and 1993, Schönberg and Ludsteck (2014) find that the impact of such expansions on overall maternal employment, employer continuity, and labor market income 3–6 years after childbirth is small. The small magnitude of the impact could be the result of pre-labor market selections, either into school track or field of study and occupation. In our work, we focus on the role of decisions and human capital accumulation inside the university and education system, such as field of study. Our work is related to a number of existing studies which focus on highly skilled individuals, primarily in the United States. Bertrand et al. (2010) study the careers of men and women who gained a master's degree in business administration (MBA) from a top US business school between 1990 and 2006. They find that, 10–16 years after MBA completion, the male (unadjusted) earnings advantage is nearly 60 log points. Differences in business school courses and grades, differences in career interruptions and differences in weekly hours worked are identified as the three most important explanations for the large and rising gender gap. Interestingly, important differences are observed also at the outset of men and women's careers. One year after MBA completion, men earn approximately 15% more than women. Although the gap goes down to about 6% after controlling for a large set of characteristics (e.g., pre-MBA characteristics, MBA performance, labor market experience, and weekly hours), it remains statistically significant. Looking at pharmacists, Goldin and Katz (2016) find that conditioning on hours of work reduces the gender earnings gap from about 28% to 4%–7%. Among those without children, female pharmacists earn only 1% less than comparable male pharmacists, and this difference is not statistically significant. It is unlikely, however, that this pattern is observed across the range of other professional occupations, and it is part of our objectives to explore this for Germany.8 Two recent studies look at the experience of countries other than the United States. Building on the previous work by Albrecht et al. (2003), the first study is by Albrecht et al. (2017). They analyze Swedish matched employer-employee data on men and women born in the 1960s who completed their university education in business and economics and are followed for 20 years after graduation.9 At the start of their careers, at age of 25–26 years, men and women have virtually the same monthly full-time equivalent wages, but by age 45 there is a sizeable gender wage gap of about 25 log points. Albrecht et al. (2017) argue that, over and above the differences in firm characteristics and firm-to-firm mobility by gender, the main driver of the gap appears to be the greater wage gains experienced by men as opposed to women. The second study by Bütikofer et al. (2017) looks at Norway. Using detailed register data, they focus on the effect of parenthood on the careers of men and women with MBA, law, medical, and STEM degrees.10 They show that, among top earners, the share of women holding a degree in all subjects, except medics, has increased substantially since the mid 1980s. Interestingly, when focusing on men and women at the start of their careers, the average earnings prior to childbirth are identical in levels for men and women. As illustrated by Bertrand et al. (2010) for American MBAs, also in Norway differences in career interruptions that are associated with childbirth lead to substantial pay gaps by gender, which do not decline even ten years after childbirth.11 Comparing to our findings, our results indicate that the differentials faced by young German male and female college graduates are somewhere in between those faced by their American and Scandinavian counterparts. Section 2 describes our main data source. Section 3 presents the main results related to gender gaps at entry into, and exit from, university. Section 4 discusses the estimates found in relation to gender gaps in the labor market. Finally, Section 5 summarizes our main findings, contrasts them with the existing literature, and puts forward a number of areas for future work.","We analyze the trajectories of college graduates using survey data collected by the German Centre for Higher Education Research and Science Studies (DZHW). These data come from nationally representative longitudinal surveys of individuals who complete their university education across all fields of studies and the full range of Higher Education institutions in Germany.12 The DZHW sampled university graduates from the graduation cohorts 1988–89, 1992–93, 1996–97, 2000–01, 2004–05, and 2008–09.13 This gives us a large sample of high-skilled individuals spanning a 20-year period. For each graduation cohort, individuals are first interviewed about 12-18 months after graduation. The same individuals also participate in a second, follow-up survey about five years after graduation (see Appendix Fig. A.1). In our analysis we focus only on the first survey: as we have explained above, this allows us to abstract from crucial decisions other than those related to early careers (e.g., family formation and fertility choices) and from later career opportunities (which involve internal promotions and negotiation with employers). Each survey contains detailed information on graduates’ personal characteristics, family background, study history, and labor market experience. For each graduate, we observe the university attended and the field of study completed.14 We complement this information with administrative data published by the German Federal Statistical Office. Although the DZHW data have been gathered systematically across cohorts, there are some differences in the information available across the cohorts. For instance, information on marks from the exams (known as Abitur) taken at the end of the academically oriented secondary school track (Gymnasium) are not available for the 1989 cohort. We therefore present two sets of estimates, either excluding the 1989 cohort or dropping the secondary school mark variable from the analysis. Table 1 reports the summary statistics of key variables from the DZHW data. High school marks are scored on an approximately continuous scale from 1 to 4, where 1 is the top score and 4 is the lowest passing mark. On average, across all cohorts, female graduates enter university with a slightly better mark, showing a gap of 0.065 points which is statistically significantly different from zero at conventional levels and represents about 10% of a standard deviation. University final score is measured on the same scale, with 1 and 4 meaning top and lowest passing marks, respectively. At the end of their university career, women also graduate with slightly better marks than men across all cohorts (again statistically significant). However, as we document in our subsequent analysis, these aggregate average statistics conceal heterogeneity across cohorts. About 87% of females and 88% of males have ever been in any kind of employment since graduation and up to the survey. Between 12 to 18 months since graduation, about 71% of women are in paid full-time employment (relative to part-time employment), as opposed to 81% of men. The large 10 percentage point difference is statistically significant. Among full-timers, the mean number of hours worked per week is around 40, with no significant difference between men and women. Average full-time monthly log earnings (excluding additional bonuses or variable pay) are 7.587 (corresponding to just below 2000 Euros, measured in 2001 prices) for all workers across all graduation cohorts, with men earning 28.3 log points (or nearly one-third) more than women.15 This difference is statistically significant. In terms of median earnings (not reported in the table), the members of the 2009 DZHW cohort have median monthly earnings of 1747 Euros for women and 2329 Euros for men (including full-time and part-time workers). Among full-time workers, median earnings in the 2009 cohort are 2183 and 2620 Euros for women and men, respectively (in 2001 prices). At the time of interview, men are older, suggesting that women tend to obtain their degrees earlier than men.16 It is worth noticing that male and female graduates are also different along a number of other observable, predetermined characteristics. For instance, almost one-third of men have completed an apprenticeship before starting their degree, as opposed to only one-quarter of women. Moreover, compared to their male counterparts, female graduates have parents (both mothers and fathers) with substantially higher education, greater labor market involvement, and in higher-level occupations. For example, female graduates’ mothers have on average almost 0.9 more years of education, compared to their male counterparts (which partly reflects time trends in share of females and parental education).","To see if there are compositional differences in university education we first analyze enrolment rates by sex. We then look at potential gender differences in “quality” at entry into college, which we proxy with the Abitur grades obtained before starting a university program. Next, we consider the possibility of differential graduation rates by sex, since those who complete their program of study should be more likely to be in graduate employment than those who do not complete. Finally, we look at gender gaps in graduation marks, which are taken as a indicator of quality of the university qualification attained and could be a relevant proxy (for ability or productivity) which is easily observed by employers. Gaps in the quantity at entry: enrolment numbers ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ There might be a “quantity” imbalance at entry, that is, differential female–male enrollment into university courses. To explore this possibility, we investigate enrollment counts by gender, using aggregate administrative data on university enrollment.17 Fig. 1 shows the trends from 1947 to 2015. From the beginning of the period to 1989, the data refer to the former West Germany only, while from 1990 onwards the data has coverage of the whole of Germany, including the former German Democratic Republic. From the post-WWII period to the early 1990s, substantially more men than women entered a college program, with a female-to-male student ratio of 0.4–0.6. By 1995 and up to the most recent data point, however, the numbers have become very similar.18 The share of females enrolled has therefore steadily increased since 1995, from a ratio of 0.91 per male student, and reaching parity in 2015, when we observe a ratio of 1.002. Gaps in the quality at entry: high school grades ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ New male entrants into post-secondary education institutions may have greater high school marks than their female counterparts. We assess this possibility by looking at gender differences in Abitur grades, conditional on having obtained a degree. Grades range between 1 (top) and 4 (lowest passing grade) on an approximately continuous scale. The pooled mean of 2.23 and standard deviation of 0.63 have remained fairly constant across cohorts and the cohort-specific distributions are also relatively similar. Panel A in Table 2 reports the female coefficient from a series of regressions. In the first row, we only control for cohort dummies; while in the second we also include apprenticeship status, and family background variables (all predetermined characteristics at the time of university enrollment) as well as age, and in the third we further control for subject and university fixed effects. Notice that a negative sign means that women have a better mark than men, because values are on a reverse scale. Across all graduation cohorts, women enter university with slightly better high school marks (row (i)). The size of the uncorrected gap is 0.069 (or about 0.11 of a standard deviation in high school score). Correcting for predetermined characteristics reduces the gap to 0.015 (or 0.02 of a standard deviation, row (ii)). In fact, men in the earlier cohorts (1993 and 1997) have better marks, corresponding to a period when also more of them enrolled (see Fig. 1): in such cases, the size of the gender gap is about 0.05 (0.07–0.08 of a standard deviation). But from 2005 onwards, women have substantially higher entry marks, of the order of 0.12–0.13 of a standard deviation. This ‘overtaking’ is a phenomenon also highlighted by Fortin et al. (2015) for the United States, although in our case we consider only individuals who actually enroll at university and complete their university studies, and not all secondary school students. Controlling for field of study and institution fixed effects leads to a gap of about 0.07 of a standard deviation across all subjects (row (iii)). Panel B of Table 2 shows there is heterogeneity across specific degree areas, although the broad pattern emerges within each field. Among those who enroll in medical sciences and economics/business, women tend to have better secondary school marks. This is not the case among those who enroll in STEM (which are typically male dominated) and humanities subjects (which are typically female dominated). In summary, our starting point is to assess whether there are gender differences when men and women start their university careers, over the period that pertains the DZHW data. There is a greater fraction of men enrolling at university up to the mid 1990s, when men also had a small advantage in terms of quality at entry. In the most recent data, however, women and men enter college in equal numbers and female graduates begin their university careers with slightly better high school marks. Gaps in the quantity at exit: graduation numbers ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We now move to the end of the university process and focus on two graduation outcomes. We first consider whether there are differences in the quantity of female graduates as opposed to male graduates. To describe the time trends in the numbers of male and female graduates, we use administrative data on the number of graduates by field of study. Fig. 2 reports the raw numbers, while Fig. 3 shows the corresponding female/male ratios. From the early 1990s to the mid 2010s, the total number of graduates has gone up from 181,000 in 1993–339,000 in 2015 (see panel (a) in Fig. 2), implying an increase of about 87% over the period (an average increase of nearly 4% every year). The share of all graduates represented by women has also increased from about 40% in 1993 to approximately 52% in 2015, or equivalently from a female/male ratio of 0.66 to a ratio of 1.09 (see Fig. 3). Some fields of study, such as fine art, art history, and humanities (including foreign languages), have been traditionally female dominated, with 60–70% of graduates in the mid 1990s being women. Over time, they have become even more female dominated, with 65%–75% of graduates being women in 2015 (see panels (c) and (f) in Fig. 2). Conversely, other fields—such as science, technology, engineering, and mathematics (STEM)—were male dominated in the 1990s and continue to be so in the 2010s, although the share of all graduates represented by women has increased from about 15% to almost 25% in engineering and computing and from around 35% to 40% in all hard sciences and mathematics (see panels (d) and (g) in Fig. 2). Clear increases for women in graduation rates can be observed in health sciences (including medicine and pharmacy), social sciences (including law, business studies, economics and psychology), and agricultural studies (including nutrition science). In the first two areas of study, in particular, the growth in the number of female graduates has been quite remarkable (see panels (e) and (h) respectively in Fig. 2, as well as Fig. 3): up by about 12 points in the case of social sciences (from 44% to 56% of all graduates) and up by nearly one quarter in the case of health sciences (from 45% to almost 70%). Notice that both the time trend over our sample period and the actual fraction of female graduates in all health science programs in Germany are very similar to those shown by Goldin and Katz (2016) for pharmacists in the United States. Gaps in the quality at exit: graduation grades ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A second outcome at exit is degree quality, which is measured by the final graduation mark. To assess whether there are gender gaps in grade performance, we use data from across the six DZHW cohorts and perform three different exercises. Means In the first exercise we estimate five different specifications of a final grade equation, keeping in mind that marks range from 1.0 (best) to 4.0 (worst passing mark), and thus a positive sign indicates lower marks for female graduates. Using data pooled across five cohorts (excluding 1989) and with a regression that controls for survey wave dummies, we find a small (but statistically significant) gender gap at 0.022 points (s.e.=0.006) that favors men (see column (a) of Table 3). This difference corresponds to just 0.032 of the sample grade standard deviation. The addition of demographic and socioeconomic controls to the regression increases the gender gap to 0.04 points (column (b)), suggesting that the lower marks obtained by women are unlikely to reflect differences in socio-demographic and family background variables. Adding high school marks, field of study fixed effects, and university fixed effects (columns (c)–(e)) leads to gradually higher gender gaps reaching 0.063 points (column (e)), which corresponds to 0.09 of a standard deviation. These estimates demonstrate that conditioning on covariates to control for differences in individual characteristics and family background, high school results, subject of study, and university attended does not eliminate the small grade gender gap that favors men over women. In fact, if anything, such differences magnify the grade gender gap, although this remains arguably modest. These patterns are in line with the descriptives presented in Table 1 above. When such estimates are disaggregated by graduation cohort (see Table 4, panel A), we find a strong time gradient, whereby the gap is about twice as large at the beginning of the 1990s as it is in 2009. The estimates in panel B, where we do not control for high school marks (and therefore the 1989 cohort also contributes to the analysis) show an even stronger gradient. Although the gap at the start of the sample period is almost 0.1 points in favor of men, in 2009 this is 70% smaller. Comparing the results in panels A and B suggests that accounting for high school marks accentuates the gender difference in terms of graduation marks from 2001 onwards. This is in line with the significant differences in terms of high school marks seen in Table 2(iii) for these cohorts, and the strong link between high school grades and graduation marks. High marks By looking at linear mean regressions we cannot detect nonlinearities. With the second exercise, therefore, we examine whether there are gender differences at the top of the final grade distribution. This allows us to see if there are different effects of gender among high achievers. Using the same specifications employed for the mean graduation marks regressions before, we estimate linear probability models in which the dependent variable takes value one if an individual obtains a graduation marks of 1.5 or better (i.e., a mark between 1 and 1.5), and zero otherwise. The results from this analysis are reported in panels C and D of Table 4, where, besides the full set of controls, we investigate the role of secondary school marks as additional control. Irrespective of this inclusion, we find that, across all graduation cohorts, women are 3.5–4 percentage points less likely than men to graduate with top marks, a nontrivial effect that corresponds to about 14%–16% of the sample average probability of obtaining a mark at 1.5 or better. An important question, therefore, is to understand what happens during the college career of male and female students, given that women start university with an indicator of higher quality (i.e., better secondary school grades). In part, this might be due to a higher dropout rate among men (since equal numbers of men and women enroll and more women graduate), something however we cannot test directly with our data. Dropout rates are substantial: for students starting 2006/07, Heublein et al. (2012) estimate dropout rates for Bachelor students in Germany at 35% at universities, and 19% at universities of applied sciences. Comparing across genders, males are found to have a 6–10 percentage points higher dropout rate (at universities, 38% for males versus 32% for females; at universities of applied sciences, 23% for males versus 13% for females). If academically weaker men abandon their university studies (and do so more than women), each cohort of graduates could be composed of academically stronger men. This in turn may be reflected in their performance in the labor market, which we will analyze in the next section. Of course, there could be other explanations, for instance, universities might have academic curricula that fit men’s abilities better than women’s, or build men’s skills more efficiently than women’s. These questions are an interesting issue for future work. Heterogeneity by field of study In the last exercise, we look at differences in graduation marks in the four subject categories we already analyzed above (economics/business, STEM, humanities, and medical sciences). These have different gender compositions (e.g., humanities are female dominated and STEM subjects are male dominated) and have gone through major changes in terms of gender representation in the last two decades, as shown in Figs. 2 and 3. Table 5 reports the gender gap estimates by cohort, using the same specification as that shown in panel A of Table 4, i.e., including secondary school marks in all regressions. There is large effect heterogeneity across degree subjects. Male graduates in STEM subjects and humanities have better marks at graduation than women. The difference over all cohorts is 0.13–0.18 of a standard deviation in graduation marks and is statistically significant at conventional levels. In the humanities, the gap has become smaller (and statistically insignificant) for the most recent graduation cohort, while in STEM subjects the gap has remained fairly constant over time. In economics/business, where the fraction of male and female graduates is more similar, the gender differential is substantially smaller, and we cannot detect any significant gap. As documented earlier, medical studies have witnessed a rapid and considerable increase in female graduates. In this field, the gap (favoring men) has grown among graduates in the most recent cohorts, when fewer men obtain medical related degrees. This might indicate an increasing positive selection of men into medical subjects. We close this section with a summary of its main results. We emphasize six findings. First, since the mid 1990s, roughly equal numbers of men and women enroll in Higher Education programs in Germany. Second, conditional on having completed college, women enter university with better secondary school marks. Third, at the end of their university career, more women than men obtain a degree. Fourth, there is some persistent educational specialization by gender, with substantially more men in STEM subjects and more women in arts and humanities, although this segregation has lessened in recent years. Fifth, female graduates do not outperform male graduates in terms of final exit marks. The difference reveals a better performance among male graduates. This is clearer when we look at the top of the final university grade distribution, whereby men are 3.5 percentage point more likely to obtain top grades than women over the whole sample period. Sixth, gender gaps in graduation marks differ by field of study. When we pool graduation cohorts, we detect larger differentials among graduates in the humanities and STEM subjects. The reversal of relative performance at the end as opposed to the start of the university career is interesting and, to our knowledge, new.","To assess the extent of gender differentials in the labor market we begin with an analysis of the earnings gap among full-time workers observed 12–18 months after graduation. We will start by looking at differences at the mean, but then also analyze effects across the earnings distribution. We then consider the gender difference in the probability of being in full-time employment and the number of hours worked. Finally, we perform a decomposition analysis of the earnings gap so that we can better gauge the role played by different sets of predictors, especially those that are not normally observed in nationally representative datasets, such as university attended and fields of study. At the mean Table 6 reports estimates of the gender gap obtained from different sample selections and with or without the inclusion of different sets of correlates. To avoid issues related to choices or constraints which we do not model here, we focus on full-time workers only. The gender gap in full-time earnings over the whole period is 28.7 log points, including only graduation cohort and state dummies (row (i)). This is remarkably close to the figures reported in Machin and Puhani (2003), although those refer to all graduates, and not just recent graduates. The temporal pattern of the gap is only slightly negative, going down from 34 log points in 1989 to only 24 in 2009. This gap may in part reflect a differential propensity by gender to enter into different types of paid professional training (e.g., lawyers, medical doctors, and teachers) or to start a salaried doctoral program of study. Although paid, both training programs and Ph.D. studies are likely to be paid below market wages. Excluding individuals in training or Ph.D. programs leads to the estimates shown in the second row (row (ii)). About 30 per cent of the sample are dropped, and the gender gap becomes one-third smaller, but without a clear time trend. Selecting the sample further (i.e., requiring the availability of secondary school marks data, and thus dropping the 1989 cohort) leads to virtually identical gender gap estimates (row (iii)). The inclusion of a large set of controls (i.e., personal characteristics, family background variables, graduation university score, university fixed effects, and field of study fixed effects) reduces, but does not eliminate, the gender pay gap. When all individuals are included (row (iv)), this goes down to 7.8 log points for all cohorts pooled together, a reduction of more than 70% in comparison to the estimates shown in the first row. We also see a more pronounced reduction over time, with the gap almost halving from 10.1 log points in 1989 to 5.6 log points in 2009. The inclusion of controls in the sample in which men and women in professional training or doctoral programs are excluded (row (v)) leads also to a smaller gender gap than in the case without controls in row (ii). The reduction is about 55%–60%, with a resulting gender gap of 8.4 log points (across all cohorts). Once we account for covariates therefore we obtain similar estimates of the gap irrespective of whether we exclude or include the subpopulation of individuals in professional training or Ph.D. programs. Controlling for secondary school (Abitur) marks does not affect the estimated gap regardless of the selection on training (rows (vi) and (vii)), nor does the inclusion of weekly hours worked (row (viii)), although this last addition can only be assessed for the 2009 cohort (because detailed information on hours of work among full-time workers was collected only for the 2009 cohort). Previous research has emphasized the importance of the relationship between pay and hours, with nonlinearities indicating a higher career cost of family (Goldin and Katz, 2016). If there is linearity, a worker who works, say, 80 hours a week will be paid twice as much as another worker who works 40 hours. Female full-timers may earn less than their male counterparts, but this might be entirely driven by the fewer hours they work; and women might deliberately choose to work fewer hours because they wish to have greater flexibility and balance job and family responsibilities. The results in Table 6 suggest that hours differences do not play a major role in driving gender earnings gaps in our context, since including hours worked in the analysis does not eliminate the pay differential between men and women.19 The estimates in Table 6 are remarkably close to those reported by Bertrand et al. (2010) for the Chicago MBAs at 1–3 years since MBA receipt. There are two striking features in this similarity. The first is that our estimates refer to a different country, with a different school system, and different labor market institutions, employment protection laws, and maternity leave regulations. The second is that they emerge across all university graduates in Germany, not just those in the corporate and financial sectors in which MBA holders are typically employed and where it is known that women are not getting ahead fast enough (e.g., Bertrand and Hallock, 2001; Philippon and Reshef, 2012; Wolfers, 2006). A formal decomposition of the earnings gap using gender differences in means of the explanatory variables and the coefficients from the specification in line (v) of Table 6 shows that the single most important contribution to the mean level of the gender earnings gap across all graduation cohorts is given by field of study (see Table 7). This proximate factor alone can account for 9.4 log points of the explained difference between male and female earnings (or 83% of the 11.3 log point explained average gap). The same occurs when we examine each graduation cohort separately. Similar results also emerge when we analyze the larger sample that does not exclude graduates in professional training and Ph.D. programs (see Appendix Table A.1). Leuze and Strauß find that descriptively those subjects with high share of females tend to be associated with lower labor market returns. These findings are consistent with the evidence put forward by Machin and Puhani (2003) for Germany and the UK, even though their study focuses on all graduates and not just those who are observed one year to 18 months since graduation. They are also consistent with recent evidence that shows that different fields of study have substantially different labor market payoffs, even after accounting for individuals characteristics and awarding institution (e.g., Kirkeboen et al., 2016). At Different Quantiles One well established result for Sweden is that, although male and female wages are close to equal at the bottom of the wage distribution, they become extremely unequal at the top of the distribution (Albrecht et al., 2003), where presumably there is a greater concentration of college graduates. Similar evidence of greater gender differences at the top is also found among recent cohorts of full-time American workers (Blau and Kahn, 2017) and English graduates (see Britton et al., 2016). For Germany, Gallego-Granados and Geyer (2015) find that the gender wage gap decreases for higher quantiles, while Antonczyk et al. (2010) find a U-shaped gender wage gap.20 In what follows, we explore the gender earnings gap at different percentiles of the distribution for university graduates. Fig. 4 shows the results. Regarding the raw gender difference, we observe a larger gender pay gap among lower paid graduates than among their better paid counterparts. To investigate what share of the raw gap can be explained by the covariates, we use the decomposition method described in Chernozhukov et al. (2013). The results are reported in Appendix Table A.2. It is striking that once our rich set of covariates is controlled for, the gender wage gap is very similar across the quantiles, at around 7% (see Table A.2). As a result, the variables we use in our models can explain substantially more at the bottom of the distribution than at the top. We find that our covariates explain 65%–75% of the gap in the bottom three deciles, but only 30%–40% in the top two deciles. Interestingly, further results (not shown for space constraints) show that this pattern is essentially driven by the role of major choice, this is in line with the results in Table 7. Full-time employment and hours ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Previous research has emphasized that a major reason gender differences in earnings are observed is due to differences in job experience and hours, and a large part of such differences comes from job interruptions, with most job interruptions being due to children (Angelov et al., 2016; Bertrand et al., 2010; Bütikofer et al., 2017; Goldin and Katz, 2016; Goldin and Mitchell, 2017; Juhn and McCue, 2017; Sasser, 2005; Wood et al., 1993). By looking at gaps occurring 1–1.5 years out since graduation, we are likely to mitigate most of these channels which possibly emerge only after a longer time span. Here we explore the role played by the differential selection into full-time employment and by differences in weekly hours of work. Table 8 shows the results of this analysis. In the first two columns we report the estimates from linear probability models in which the dependent variable equals 1 if individuals are in full-time employment, and 0 part-time. Even soon after graduation, female graduates show a significantly lower propensity to be employed in a full-time job. When we consider all cohorts together, the (raw) employment gap is 9 percentage points (column (a), where we only control for cohort dummies). This gap has varied between 7 and 13 percentage points between 1989 and 2005, while only among graduates from the 2009 cohort do we observe a substantial reduction in the gap to less than 6 percentage points. When we include the whole set of (individual, family background, pre-university and university) controls, the gender employment differential becomes considerably smaller, less than 2.5 percentage points among graduates across all cohorts, but still statistically significant (column (b)). For the most recent cohort, however, we find a reversal in the gap, with female graduates experiencing a greater likelihood of being employed full time by 1.1 percentage points than their male counterparts, although this gap is not statistically significantly different from zero. Standard decomposition analysis shows that field of study plays again a major role in explaining the gap by gender (not reported for the sake of brevity). In columns (c) and (d), we take advantage of additional detailed information on forms of employment other than full- and part-time jobs, such as employment in temporary jobs and freelance work. This information is only available from the 2001 cohort onwards. The results show that, in the most recent cohort, women are 3–4 percentage points more likely to be ever employed than their male counterparts. The last two columns of Table 8 (columns (e) and (f)) display the estimates on the gender differences in hours worked. These refer to the 2009 cohort only because, as mentioned before, information on hours of work for full-time employees was collected in the 2009 round but not in the earlier waves. Every week women work about 0.2 fewer hours than men. When all controls are included, the gap grows to about 0.39 hours a week. This difference of around 25 min of work per week is statistically significant and accounts for about 12% of a standard deviation in hours worked. The difference cannot account for more than 15% of the 6 log point difference between male and female pay found for the 2009 cohort (see Table 6, row (viii)). Thus, hours of work are important in this context, but are unlikely to be the whole story among full-timers. In sum, 12–18 months out of university, gender gaps on the intensive margin are quantitatively small and appear to play only a minor role in explaining the observed gender pay gap. Gender gaps on the extensive margin instead are more pronounced, and field of study turns out to be a key proximate factor in determining such differentials. To understand more about the role played by degree categories, we next analyze the extent of full-time gender earnings gaps across different fields of study. Heterogeneity by field of study ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We look again at the four different degree subjects we have already analyzed before (Table 9). Panel A shows the estimates obtained on the sample that excludes individuals in training and Ph.D. programs, whereas panel B includes such individuals. In interpreting the estimates one needs to keep in mind that different subject fields have different propensities for further training, as seen from the difference in sample sizes between panels A and B in Table 9. The gap between male and female economics/business graduates is insensitive to sample selection and type of specification. Just controlling for cohort and regional dummies results in a gender gap of about 10 log points (column (a)). After all controls are accounted for (specification (e)), the gap remains essentially unchanged and always highly statistically significant. More variation is instead observed for the other fields. For example, when we only condition on time and regional dummies (specification (a)), the full-time gender earnings gap is 30 log points among STEM graduates if we include individuals involved in training/Ph. D. programs (panel B), but it reduces to 13 log points if we exclude them (panel A). In the case of medical studies, we find a significant pay gap of about 9 log points in panel A. Including individuals in training and internship programs leads a much larger sample (panel B), in which however we cannot detect any significant pay differential by gender. Similar patterns emerge in the case of humanities, albeit with different values of the estimated gender gaps. Overall, in Panel A the point estimates are quite similar across subjects (column (e)), even though the estimate for humanities is not statistically significant. In Panel B, however, there is a striking difference between Economics/Business and STEM on the one hand, with significant gender gaps close to 12 log points, and Humanities and Medical studies on the other hand, with an insignificant gender gap of around 2 log points. Given a large fraction of graduates in the latter fields are involved in some sort of further training, the next step it to see what happens to their labor incomes after its completion. This is another area to be explored by future research. Public versus private sector ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Recent work by Mas and Pallais (2017) and Wiswall and Zafar (2018) suggests that there are significant differences in male versus female valuation of non-wage benefits, and job stability and work flexibility in particular. We investigate this point by comparing the gender-specific probability of entering the public sector. Clearly there are many potential differences between public and private sector jobs. Nonetheless, the additional job security and the flexibility in terms of work-time (including the ability to adjust between full-time and part-time as family circumstances change) is a potentially important aspect in this context. Table A.3 shows the results, where indeed we find a striking difference between gender. Across all subjects, females are 4 percentage points more likely to enter the public sector than males, which is a substantial difference given the baseline rate of about 13%. In particular, when we look at STEM subjects, females are 8 percentage points more likely to work in the public sector, compared to an overall mean of less than 10 per cent. In our data, monthly earnings of public sector workers are significantly lower (controlling for field of study).21 This suggests that these tradeoffs maybe very relevant in the population we are studying. Occupational and sectoral choice are therefore important areas for future research.","We have examined gender differences in university performance and labor market outcomes among German college graduates. A key aspect of our labor market analysis is that we have focused on the short run, looking at gender gaps soon after graduation, thus limiting the role played by marriage, children, career interruptions, negotiations with employers, and promotions which many have found to put greater constraints on women’s careers than men’s. Before 1995, more men entered college than women; but in the following 20 years, the numbers have become roughly equal. More women than men, however, complete their degrees. This is consistent with differential dropout rate by gender which might have consequences on the cohorts of male and female graduates entering the labor market. In terms of quality, female graduates enter college with slightly better high school marks,22 but leave university with slightly lower grade levels. This again might reflect a greater dropout rate among men. But it might also reflect other factors (such as men achieving maturity and catching up with women in terms of academic skills, or universities offering programs that are better suited to men’s than to women’s abilities), which deserve more research in the future. Immediately following university completion, full-time men earn more than full-time women, even though they have very similar weekly hours of work. Between 12 and 18 months after leaving college, the gender gap in full-time monthly earnings is nearly 20 log points across all graduation cohorts, if individuals in professional training or Ph.D. programs, who are likely to be paid below market wages, are excluded. The single most important proximate factor that explains the sizeable, early gender gap in pay is the field of study at university. There is heterogeneity in the magnitude of the gender pay gap by field of study, with the largest differentials emerging among graduates from economics/business and STEM subjects. Once the full set of controls is taken into account, the remaining wage gap is about 8 log points when we estimate across all available cohorts. Several channels may be at work here. One could be related to human capital considerations. The importance of field of study in our results indicates the relevance of pre-market choices. These also interact with subsequent market decisions (such as occupational choice) at the very beginning of professional careers (e.g., Canon and Golan, 2017; Liu, 2016). In turn, such choices could be partly driven by gender differences in preferences (e.g., risk aversion), self-confidence, competitiveness, earnings expectations, and valuation of non-wage benefits (e.g., Buser et al., 2014; Fortin, 2008; Mas and Pallais, 2017; Reuben et al., 2017; Wiswall and Zafar, 2015). Joining some (if not all) of these drivers into one coherent analysis would be another important area for future work. Another possible channel is related to statistical discrimination against women, based on employers’ difficulty in distinguishing more from less career oriented women (e.g., Gayle and Golan, 2012; Reuben et al., 2014; Thomas, 2018). Also this mechanism deserves more attention in future research."],["World economic development is associated with growing food consumption. Agricultural land, however, suffers from over-exploitation and is subject to environmental shocks which are projected to become more severe due to climate change. We present a stochastic model of a dynamic economy where soil is an essential input and natural disasters are sizeable, multiple, and random. Expansion of economic activities raises effective soil units but contributes to an aggregate loss of soil-protective ecosystem services, which exacerbates soil degradation at the time of a shock. We provide closed-form analytical solutions and show that optimal development is characterized by a constant growth rate of stocks and consumption until an environmental shock arrives causing all variables to jump downwards. Optimal soil management consists of spending a constant fraction of output on preservation measures, which is an increasing function of the shocks hazard rate, degradation intensity of agricultural practices, and the damage intensity of environmental impact. We derive the optimal propensity to save and discuss the impact of human pressure and risk exposure on soil and output. We also discuss quantitative impacts of climate change on optimal soil management. --------------------------------------------------------------------------------","Rising food demand and growing risks of large-scale land degradation suggest bringing back to macroeconomics what is by now the “nearly forgotten resource” – soil.1 It is estimated that by 2050 agricultural production will have to rise by 70% to meet the needs of growing world population (FAO, 2009).2 However, already nowadays about one third of all soils are degraded and the global amount of productive land per person could in 2050 be only one- fourth of the level in 1960, if current management practices and policies remain unchanged.3 A moderate interpretation of U.S. President Franklin D. Roosevelt’s quote that “A nation that destroys its soils destroys itself” would read that important aspects of well-being like food security, human health, clean air, and clean water are at risk when quantity and quality of world’s soils are not taken care of Wall and Six (2015). The challenge is amplified by changing weather patterns and climate variability, which force farmers to adopt ecosystem-unfriendly practices. This provides ample motivation to study the consequences of endogenous stochastic land degradation for long-run development.4 Soil cultivation raises availability and productivity of arable land but involves major risks and uncertainties. In fact, risk is an inherent element of all agricultural activities. The development of agriculture itself was a response to the risks of relying on hunting and gathering for food (Hardaker et al., 2004). The impact of risk has not disappeared, of course, but it has changed its character and has shifted to different areas. It is still inevitable, because food markets are closely interlinked with the rest of the economy and, in particular, because ecological systems and weather conditions are increasingly subject to perturbations. Many developing countries experience large variations in rainfall and are subject to frequent extremes of flooding or drought, both of which contribute to soil erosion and land degradation (UNEP, 2015). Drought and erosion are exacerbated by poor land management that responds inappropriately to climatic variations (UNCCD, 2009), which are predicted to increase in frequency and severity due to climate change, increasing land degradation even further (UNCCD, 2013). Important negative effects on soil quality arise from the agricultural sector itself, creating potentially harmful changes in the ecological systems supporting the farms (Lichtenberg, 2002). For example, land use changes, irrigation, or deforestation may cause soil erosion and nutrient depletion. The use of pesticides, animal wastes, and soil siltation may contaminate surface and ground waters. Salinization of rivers may damage crop production in downstream areas, while irrigation and land clearing may lead to land loss to selenium and salt drawn up from subsoils. Degradation is often due to increased disruption of macroaggregates, reductions in microbial biomass, and loss of labile organic matter which are induced as negative externalities from aggregate economic activities. Under appropriate market conditions, farmers are able to respond to risks in an optimal manner. However, institutional deficiencies, such as incomplete land tenure systems as well as negative externalities from individual farming on aggregate soil protection, cause suboptimal development. Also, the lack of information about the impact of soil shocks and high individual rates of impatience may constitute management problems for soil. Soil degradation can be addressed by different types of corrective measures. Labrière et al. (2015) empirically estimate the impact of contour planting, no-till farming and use of vegetative buffer strips on the reduction of soil erosion and find enormous potentials. They conclude that the government or natural resource managers can help decrease soil losses on a large scale. It has been stressed in the literature that such measures can be implemented by individual farmers, by farmer cooperations, or via public policies.5 It is often stated that human pressures on soils are reaching critical limits, reducing and sometimes eliminating essential soil functions (Islam and Weil, 2000). Human pressures may entail unsustainable land use and vulnerability to land degradation, including overcultivation, overgrazing, poor irrigation practices, deforestation, and polluting industrial activities (UNCCD, 2009).6 Food production and security are global concerns but soil losses are especially acute in arid, semi-arid, mountainous, or tropical regions where the lack of protection may cause substantial deterioration in soil quality and reduction in yields. These types of environmental problems in production and food provision and their link to general economic development warrant a thorough investigation from a theoretical perspective. Model and findings ~~~~~~~~~~~~~~~~~~ We develop a dynamic model of an economy where investments increase the capital stock and the effective soil stock, defined as an index of soil quantity and quality. Investments in agricultural expansion have positive output effects but may also entail negative impacts on soil quality. For instance, clearing and fertilizing of soils are aimed at raising land productivity and profits but may, at the same time, harm the functioning of the ecosystem making it more vulnerable to degradation and natural disasters such as floods, droughts, storms, landslides, etc. The latter may typically have large economic consequences and at the same time are not easily predictable. We therefore model such calamities as random shocks. Our model refers to a planner internalizing such external effects in the economy. This is a generic approach to policy, so that we do not have to specify whether the political actors are public or private and, similarly, domestic or foreign. The paper addresses several specific research questions. We first ask about the long-run development prospects of an economy which is subject to endogenous stochastic soil degradation. A second issue concerns optimal management practices and optimal public policy in the presence of random shocks. We specifically inquire about the optimal balance between economic expansion and protection of existing soils. Our model delivers closed-from analytical solutions for the optimal growth rate of consumption, the gross saving propensity, and optimal soil preservation measures. We extend the analysis to include human stress and increasing risk exposure. Our main findings can be summarized as follows. First, we confirm that the damage-related parameters, such as the shock hazard rate, the extent of harmful agricultural practices, soil exposure, and damage intensity, have the expected negative effect on the optimal consumption growth rate. Productivity naturally fosters the optimal growth rate, while soil protection efficiency may have either a positive or a negative effect, which we explain in detail. Soil shocks are transmitted to the capital stock intensifying the economic downturn. Second, we show that the optimal soil preservation consists of devoting a constant fraction of output to soil protection. This fraction is an increasing function of the shock arrival rate, agricultural technology, proportion of harmful by-products, exposure component, and damage intensity. The role of each parameter is therefore identified precisely. Third, we show that the optimal propensity to save and invest crucially depends on the elasticity of intertemporal consumption substitution and therefore widely adopted logarithmic preferences, favored for their simplicity and tractability, may deliver misleading policy conclusions. Fourth, as long as agricultural labor force increases the marginal productivity of land, it also raises the optimal growth rate of consumption and the optimal fraction of output devoted to soil protection. Fifth, we show that higher population raises soil damages in case of a shock, reflecting the concern of human pressure on soil. Sixth, countries with a higher risk exposure turn out to experience a slower economic growth. Finally, we resort to the data on worldwide soil degradation to illustrate and discuss the consequences of increasing hazard rates of weather and ecosystem-unfriendly practices on optimal soil protection. Contribution to the literature ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Starting with the seminal contribution of Dasgupta and Heal (1979), growth possibilities in the presence of natural resource scarcity has been one of the main topics in resource and environmental economics. Within the endogenous growth literature it has been established that exhaustibility of natural resources may not pose a constraint on growth when innovations are resource-saving. On the other hand, exploitation of the resource and the associated pollution may pose environmental problems (Gradus and Smulders, 1993). In addition, resource scarcity may disrupt social processes that support innovation (Barbier, 1999). With the recent advances in endogenous growth theory, economic development under the restrictions of a finite planet has become again a topic in mainstream literature. Peretto and Valente (2015) call the sustainability of development in a limited habitat “one of the main future challenges” and investigate mechanisms linking resource availability, income, and population, thinking of the natural resource as being “land.” While their paper focuses on endogenous population and innovation dynamics, land input remains fixed. This is where our paper aims to make a contribution. Our model highlights that soil cultivation is able to increase land availability and/or land productivity even when the surface of the planet is finite. But we also stress that soil expansion comes at the cost of increasing the risk of land degradation triggered by environmental shocks. According to our setup, effective soil can be expanded by agricultural investment activities on the one hand but it is subject to stochastic and endogenous environmental shocks on the other. In this respect, our paper also builds on the literature on optimal environmental policies, where, for example, the negative impact of pollution on productivity or utility is stochastic and appropriate taxation measures are warranted (e.g. Bretschger and Vinogradova, 2017; Soretz, 2007). Many contributions in the literature deal with the complex interplay between agriculture and ecosystem services. The early contribution of McConnell (1983) uses soil depth and soil loss for aggregate production to determine the difference between the private path of erosion from the socially optimal path. Dalea and Polasky (2007) analyze the environmental impacts of agricultural practices on a wide range of ecological services such as water nutrient cycling, soil retention, quality, pollination, carbon sequestration, and biodiversity conservation. They also show that ecosystem services have a positive impact on agricultural productivity. Heal and Small (2002) study agriculture as a producer and consumer of ecosystem services and stress that the quantity and quality of ecosystem services depend on the joint actions of many dispersed resource users.7 Lanz et al. (2016b) stress that the concentration on a small number of highly productive crops has led to a significant loss of biodiversity which may have a negative feedback on agriculture. We build on these fundamental relationships in our framework by explicitly modeling the decrease of protective ecoservices as an externality of agricultural activity. Grepperud (1997) extends the previous studies by analyzing time-limited effects of optimal soil conservation measures and long-lasting effects on the soil base.8 In the study of the tropical forest ecosystem in Bangladesh, Islam and Weil (2000) argue that “human population pressures upon land resources have increased the need to assess impacts of land use change on soil quality.” Our paper departs from the deterministic framework used in these papers, introducing random environmental shocks which have been classified as “production uncertainty” and qualified as a “quintessential feature of agricultural production” (Moschini and Hennessy, 2001, p. 90). Most studies of dynamic soil decisions under risk are empirical. An early exception is Hertzler (1991), who provides an overview of the analytical tools and various exemplary applications to the field. A paper closely related to our study is Shively (2001) which also considers a dynamic stochastic model of soil preservation but does not provide analytical solutions to the theoretical problem. The focus of Shively’s analysis is on the role of farm size and liquidity constraints for the decision of subsistence-oriented households to make a one-time investment in protective installations reducing the risk of soil erosion. Thus, the dynamics of the investment process are not considered. His main conclusion is that public policy should aim at enhancing saving and insurance mechanisms for small farms which are most likely to face liquidity and subsistence constraints. Lanz et al. (2016a) introduce random shocks to total factor productivity in agriculture and show that uncertainty requires more land to be converted into agricultural use as a hedge against production shortages. Our paper aims at making a theoretical contribution to the literature seeking to understand the interactions between growth and the environment. For this purpose, we combine the essential characteristics of soil dynamics from the agricultural literature with a workhorse endogenous growth model from the macroeconomics literature. The key strength of our approach is that it allows for closed-from solutions with respect to the main variables of interest, such as preservation measures, which are of particular importance from the policy perspective. The remainder of the paper is organized as follows. Section 2 develops our baseline framework. In Section 3, we present the main results with respect to the optimal growth rate, soil protection, and saving propensity. Section 4 adds the development impacts of human stress and risk exposure and presents a quantitative application. Finally, Section 5 concludes. General model ~~~~~~~~~~~~~ Production process may involve soil management practices which are potentially harmful for effective soil in case of an environmental event, aggravating its consequences. Put differently, land clearing may weaken soil resilience or may result in by-products which lead to deterioration of soil-protecting ecosystem services. We postulate that a unit of output is accompanied by η units of harmful practices which cause damage to the ecosystem (Clarke, 1992). Exploitation of ecosystems leads to their weakening or even exhaustion, which in turn reduces their capacity to protect and preserve the quantity and/or the quality of agricultural land, making the soil vulnerable to degradation.10 Solving the model ~~~~~~~~~~~~~~~~~ It follows that the last term on the RHS is negative and it represents the downward jump in consumption every time a shock occurs. The solution of the maximization problem (1)–(7) is characterized by optimal consumption is proportional to the effective capital-soil stock; optimal protection expenditure is a constant fraction of output; consumption, effective soil stock, and protection services increase at the same constant rate, given by (11), between two subsequent shocks. Provided in the appendix. □","In this section, we provide a characterization of the optimal consumption growth and the optimal soil protection. In particular, we are interested in the effects of the key parameters such as exposure and damage intensities, protection efficiency, shock arrival probability, population size and technology. Consumption growth ~~~~~~~~~~~~~~~~~~ The expression resembles the Keynes–Ramsey formula which is widely known in standard macroeconomics. The Keynes–Ramsey growth rate is typically equal to the difference between the real interest rate (usually the marginal productivity of physical capital) and the rate of pure time preference, multiplied by the elasticity of intertemporal consumption substitution (EICS). We note that in Eq. (13) the economy’s implicit real interest rate is given by the first term inside the curly braces. It does not only include the marginal productivity of capital-soil input (ξLβ) but also the effect of harmful agricultural practices adjusted by the protection efficiency, i.e., the term η/ω. It follows that soil deterioration has an unambiguously negative growth effect. This adverse effect may be reduced by either increasing the protection efficiency, ω, or decreasing the proportion of harmful agricultural practices, η. We will discuss the role of labor for soil development separately in the section on “human stress” below. The responses of the consumption growth rate to changes in the fundamental parameters of the model are summarized in Proposition 2. It is important to distinguish between the effect of the expected frequency of disasters and the effect of the overall uncertainty. The former takes into account only the arrival rate λ. The latter includes both the arrival rate and the damage caused by a shock, as reflected in the last term in Eq. (13). The optimal growth rate of consumption is a decreasing function of shock arrival rate (λ), proportion of harmful practices (η), and exposure intensity (σ); an increasing function of the international interest rate (r) and labor force (L); either an increasing or a decreasing function of soil protection efficiency (ω) and damage intensity (δ), depending on the parameter constellation. Provided in the appendix. □ The effect of the arrival rate (λ) on the optimal growth rate is directly proportional to the negative of the elasticity of intertemporal consumption substitution. Although there is not a general consensus on the magnitude of this elasticity, the empirically plausible range of values lies between 1 and 3. This suggests that if the frequency of disasters were to rise in the future due to a weakened ecosystem, the agrarian economy may experience an important growth slowdown. The intuition behind the effects of η and of the interest rate r is rather straightforward and has already been discussed. The effect of the protection efficiency ω is also ambiguous. On the one hand, a higher ω directly improves protection efficiency and contributes to soil preservation thus enhancing the growth rate through the first term in Eq. (13). On the other hand, a higher ω also means a lesser degradation of effective soil units and a smaller jump in consumption rate at the time of a disaster (the jump-smoothing effect). The ratio of post- to pre-shock marginal utilities of consumption is reduced (see Eq. (10)) and this contributes to a growth slowdown through the last term in Eq. (13). The overall effect of ω on g is positive when the marginal productivity of effective land units is relatively high (i.e. the level of technology (ξ) and/or agricultural labor force (L) is relatively high), or the extent of ecosystem-unfriendly practices (η) is relatively large or the degradation intensity (δ) is relatively high. This suggests that economies with a relatively high damage intensity of production and with a higher degree of vulnerability to shocks, such as numerous agriculture-dependent developing economies, may enjoy substantial gains in terms of their growth rates by adopting (more) efficient protective measures. At the same time, economies with a relatively higher total factor productivity (such as advanced economies) may also experience an improvement in the growth rate of their agrarian sectors by enhancing their soil-protection technologies. Optimal soil protection ~~~~~~~~~~~~~~~~~~~~~~~ Expression (14) shows that an economy which is highly exposed and where damage intensity is sizeable has to devote a larger share of output to soil protection. When the arrival rate is lower or exposure is less pronounced, it becomes optimal to spend less effort on soil protection in equilibrium. The following proposition summarizes the effects of the fundamental parameters of the model on the optimal protection share. The optimal fraction of output devoted to soil protection is: an increasing function of the event arrival rate (λ), technology (ξ), international interest rate (r), proportion of harmful practices (η), exposure component (σ), risk aversion (ε), and damage intensity (δ); either a decreasing or an increasing function of protection efficiency (ω), depending on the parameter constellation. Provided in the appendix. □ The intuition behind the results in (i) is straightforward. A higher frequency of disasters (λ) requires more preservation measures in order to better protect land from degradation in the event of an adverse shock. If governments happen to misperceive the true arrival rate λ, the preservation policy would be sub-optimal. Specifically, if the perceived λ is lower than the true one, there is too little preservation. This might happen due to a regime switch from a low to a high shock frequency while the general expectations, if based on past experience, lag behind. If the intertemporal substitution elasticity is above (below) unity, the optimal protection share is convex (concave) in the arrival rate; the response of the protection share to a change in the arrival rate is more (less) pronounced when protection technology is more (less) efficient and when damage intensity is relatively large (small). Provided in the appendix. □ Provided in the appendix. □ Saving propensity ~~~~~~~~~~~~~~~~~ If the EICS is above (below) unity, the optimal propensity to save is a decreasing (increasing) function of the disaster arrival rate (λ) and proportion of harmful agricultural practices (η); an increasing (decreasing) function of the damage intensity (δ) and protection efficiency (ω). Provided in the appendix. □ We note that when ε approaches unity, the derived impact of all the parameters – and most importantly those characterizing adverse events, λ, η, and δ – are at the lower bound of the empirically plausible impact range. Population pressure In the context of soil management, “human stress” denotes the hypothesis that an economy with growing population and expanding activities causes more intensive land use, faster soil degradation, and hence a sharper destruction of protecting ecosystem services. Reviewing the data and the literature, Eswaran et al. (2001, p. 7) conclude that a high population density in an area that is highly vulnerable to desertification poses “a very high risk for further land degradation.” In our model we have derived the joint development of a country’s capital and soil stocks. Economic expansion involves both increasing capital and soil use causing higher soil damages when a shock occurs, see Eq. (4). Population size has also a major effect in our economy.19 An increase in labor (dL > 0) raises the productivity of each capital-soil unit which results in a higher consumption growth. The result is an example of the so-called “scale effect” of growth which is a consequence of the constant returns to capital. It says that a larger population can exploit higher capital stock more efficiently which is favorable for growth. The empirical validation of the scale effect has been challenged in the literature, see Jones (1995). As a response to the critique and focusing on market structure, Peretto and Smulders (2002) show how an innovation-driven economy achieves endogenous growth without a scale effect. Specific to our contribution is that population size induces not only a growth effect but, at the same time, a negative damage effect. Indeed, following Eq. (4), the marginal damage as a side-product of soil investments grows with labor input (because labor is an argument in the production function). With the emergence of non-monotonic growth our model adds a new perspective on the scale effect and is not directly subject to the earlier critique about its consequences. The impact of population size on damage size addresses the often-stated concern about improper soil management in a situation of human stress. Population dynamics An interesting question is what would happen in our model if population size were fully endogenized. It is clear that such an analysis requires a different model setup and, in particular, a decentralized economy where households are allowed to make fertility decisions. A model of this type is elegantly presented in Peretto and Valente (2015). We note that, differently from our initial setup, their model includes in particular (i) a more sophisticated production sector and (ii) a constant endowment of a non-exhaustible resource. Since we do not study the effect of population size on innovation, we abstract from (i) and think of the simple production structure as in our Eq. (6). Our model adds in relation to point (ii), namely we endogenize the dynamics of resource, S. It is important at this stage to define how households perceive their impact on S. If they do not internalize the negative effect through damages to ecosystem, then they do not invest anything in protective measures and some policy is required to correct for the externality. In this case, and without appropriate policy intervention, it is clear that per- capita consumption will converge to zero in the long run. On the other hand, if households do internalize the externality, at least to some extent, their fertility decision will take into account the “human stress” effect on ecosystem and will result in a lower equilibrium fertility rate, as compared to the case without internalization, or even in a stabilization of population at a constant level. The latter case would correspond to our initial framework. In the former case, we conjecture, the endogenous fertility rate is closely related to the stochastic growth rate of S: assuming that a household size can adjust instantaneously, the population size increases at a constant rate in between shocks and drops when a shock occurs as households immediately respond to the scarcity of S. Risk exposure ~~~~~~~~~~~~~ To provide empirical evidence we may note that Africa is particularly vulnerable to land degradation and desertification (UNEP, 2015). Moreover, it is the poorest of the continents with a disproportionate share of low-income countries compared to other world regions and a very low level of economic development. The Economic Commission for Africa reported that in 2013 Africa’s share of the world population was 13%, but its share of global GDP was only 1.6%. Comparing the different countries within the continent it is striking that many poor economies, in particular Burundi, DR Congo, R Congo, Gabon, Liberia, Sierra Leone and Zambia have a very high percentage of degraded land (UNEP, 2015, pp. 123–130), supporting our hypothesis of soil damages and land degradation negatively affecting economic development.20 The result also makes clear that a growing arrival frequency in the future, e.g. induced by climate change, will deepen the income gap between the vulnerable countries of the South and the less vulnerable economies of the North, which acts against the sustainable development goals. Even if deficiencies in markets and policy of the South were absent, this would indicate a need for support of the South. Realistically assuming that market failures and weak institutions in less developed countries entail higher soil mismanagement, specific policy to correct for externalities and to strengthen institutions has to be implemented to improve welfare and development conditions in the South.21 Quantitative analysis ~~~~~~~~~~~~~~~~~~~~~ In this section we attempt to quantify our theoretical predictions, in spite of the fact that reliable numbers for soil degradation and its economic impact are rather difficult to obtain. Soil degradation has different dimensions, such as soil loss, sealing, contamination, acidification, salinization, compaction, and nutrient decline, which are not directly comparable. Moreover, the situation differs depending on climatic conditions, e.g. in arid regions vs. rain forests. Gibbs and Salmon (2015) report global estimates of total degraded area varying from less than 1 billion ha to over 6 billion ha, with equally wide disagreement in their spatial distribution. Eswaran et al. (2001) state average estimates for all degraded lands as percentage of total land for the continents: Africa 73%, Asia 71%, Europe 65%, and South America 73%. Lal (2001, p. 531) reports production losses due to soil erosion in Africa between 2% and 40% with a mean loss for the whole continent of 9% which could rise to 16.5% by 2020. Soil erosion is by no means restricted to developing countries, it is also important in Europe and the U.S. The study of Kuhlman et al. (2010, p. 27) concludes that soil erosion by water and wind causes an average productivity loss per hectare in Europe of 0.16% per year. It is calculated that an effective soil erosion control programme in the EU would provide annual on-site benefits of 500 million Euros over a 20-year period equivalent to a present value of 3.25 billion euros. Comparing the figure to the costs of the programme supports our model hypothesis of external effects, motivating the need for public policy: the on-site benefits for reducing erosion on agricultural land turn out to be smaller than their costs, while adding the off-site benefits shows significants improvement in social welfare when implementing the programme. The need for public soil protection policy depends on the size of market failure in soil management. Den Biggelaar et al. (2003, p. 2) conclude: “Comparing the results of past and present erosion studies indicates that inappropriate soil management may amplify the effect of erosion on productivity by one or several orders of magnitude.” An important case for externalities is the use of water. It is reported that the common water resources are rarely used efficiently and that “environmental externalities have arisen through excessive utilization of non-recharged aquifers while, in a number of cases, the excessive application of irrigation water has resulted in rising groundwater tables, soil salinization and sodification problems” FAO (2015, p. 402). Even though the parameters of our model are not straightforward to calibrate under these preconditions we can provide some rough measures for the expected effects of accelerating climate change. We remain very cautious about choosing parameter values and limit the analysis to the impact of a few key factors on optimal soil protection. One such important factor is the frequency of natural disasters, which is predicted to increase over time due to rising global mean temperature (IPCC, 2014; Thomas and Lopez, 2015). For instance, the number of meteorological disasters has risen from about 25 per year to over a hundred during the period 1970–2014. Similarly, the number of hydrological events has risen from about 80 to almost 150 over the same time frame (Thomas and Lopez, 2015, Fig. 1). At the same time, exposure, defined as the presence of people, livelihoods, ecosystems, environmental services, resources, etc., has also increased due to societal change. IPCC (2012) documents that fatality rates and economic impact, e.g. losses as a proportion of gross domestic product (GDP), are higher in developing countries due to higher share of impoverished populations, weak infrastructure, lack of basic facilities, and limited government capacity. Residents of the developing economies are not only more vulnerable to natural disasters but at the same time they have less resources for management options and fewer coping strategies. Therefore, soil degradation and food security in general are more acute issues in those regions. When choosing parameters for calibration we shall restrict our attention mostly to those pertaining to the developing world. In addition to disaster frequencies, ecosystem-exploitative practices have also risen over time – across the globe but mostly in the developing world (e.g., shrinking of the Amazonian rainforest). We therefore consider next the effects of η and ω on the optimal preservation measures. To do so, we use again the study of UNEP (2015) which relies on the well-established WOCAT (World Overview on Conservation Approaches and Technologies) database. For the 42 African countries, the study concludes that the loss of about 105 million hectares of croplands can be prevented provided that soil erosion is managed appropriately (UNEP, 2015, p. 11). The estimation of crop losses in terms of economic value involves an econometric estimation of the loss of ecosystem services and of the marginal physical product of different environmental variables as well as the use of market prices (UNEP, 2015, p. 58). With the already-mentioned cost of land degradation of about 286 billion PPP USD each year, the costs of inaction against land degradation in Africa turn out to be very high. Conversely, implementing better soil practices is found to be much less expensive. Specifically, expenditures for sustainable land management practices were calculated to cause an annual cost of about 9.4 billion USD or 1.15 % of the GDP (UNEP, 2015, p. 91). Protection efficiency, ω, may be raised by adopting more technologically-efficient equipment and techniques, as well as by applying fertilizers with high-nutrient content. We recall from our discussion of Proposition 3 that the impact of ω on q* is ambiguous, because ω affects the leverage of the policy, the growth rate, and the downward jump in case of a shock. Using our calibrated parameter values we find that a 10% increase in protection efficiency leads to 0.1 percentage point decline in the optimal protection share, i.e. q* falls to 1.05%. In addition to the effects of rising hazard rates of climate shocks, ecosystem deterioration and changes in technological efficiency, one may be interested in the long-term projections of expenditure on soil preservation. Such projections can be useful for government budget planning when financial priorities may be shifting in light of changing climatic conditions and additional investment projects need to be undertaken. If we remain in the context of the African economies and, more specifically, the most vulnerable region of Sub-Saharan Africa, we may ask ourselves how much this region should spend on agricultural policies aimed at soil protection over the next 30 years. Such a projection undoubtedly depends on the region’s future growth rate. The World Bank data indicate that the average GDP growth rate in the region reached 1.2% and was −1.5% on a per capita basis. We shall assume for simplicity that the economies of this region will have a zero average growth rate in the next 30 years. We also posit that the protection spending derived from UNEP (2015) of 1.15% of GDP is optimal under current climate conditions. Depending on the assumptions about shock’s hazard rates and coefficient of relative risk aversion, we obtain the following projections for expenditure on soil preservation measures: As an alternative to the analysis of aggregate soil functions and overall economic linkages one may resort to the study of subsystems like soil organic carbon (SOC) which reveals some interesting characteristics of soil degradation. It has been shown that conventional agriculture leads to a loss of about fifty percent of SOC over a period of 20–30 years (Petersen and Hoyle, 2016). Conversely, increasing SOC by using alternative soil management practices is beneficial to soil functions, fertility, and buffering capacity. However, the benefits of alternative land policies strongly vary with soil quality. For example, in areas with low rainfall and sandy soils like in Western Australia, productivity improvement in agriculture is much less important compared to the sequestration value for climate policy when increasing SOC, see the numbers in Petersen and Hoyle (2016). As a consequence, the benefits of altered soil management heavily depend on the used carbon price an thus on the adopted climate policy, which is beyond the scope of the present paper. The analysis of other specific soil functions or of specific countries and regions reveals that the degree of heterogeneity is still considerable on a less aggregate level, especially when the economic valuation is a focus. Hence, different calculations and specific considerations would be needed for each of the many different cases. To concentrate on the main messages for the economics of soil degradation, we have restricted the analysis in this paper to the aggregate approach to soil degradation and protective policies, both in theory and the quantitative implications.","The present paper considers an economy which produces output employing three essential inputs: capital, soil and labor. Production process, accompanied by harmful agricultural practices, weakens the ecosystem and diminishes its protective services. The soils become vulnerable to random environmental shocks which lead to soil degradation. Such a scenario is observed in numerous developing countries with large agricultural sectors. In an attempt to raise yields and profits, farmers clear a slot to gain arable land but at the same time they lose the protective ecosystem services and make their land exposed to landslides, floods, droughts, winds and similar calamities. To ensure sustained yields, it thus becomes necessary to adopt soil preservation measures (e.g. installation of contour hedgerows to prevent landslides). Since the extent of soil degradation depends positively on the magnitude of harmful agricultural activities and negatively on the protection efforts, an optimal soil expansion and preservation policy can be designed to maximize the economy’s expected lifetime welfare. In the present article we provide a clear-cut closed- form solution to this dynamic stochastic problem. The optimal development of the economy is characterized by capital and soil stocks and the consumption rate which all grow at the same constant rate until a disaster occurs causing a downward jump in all variables. In the benchmark model we assume that shocks arrive at a constant Poisson rate. The percentage reduction in consumption, i.e. the size of the jump, is constant and depends on the arrival rate, the damage intensity, the protection efficiency and the intertemporal substitution elasticity (EICS). The optimal soil preservation strategy consists of devoting a constant fraction of output to protection measures. This fraction is an increasing function of the hazard rate, the damage intensity, the level of agricultural technology and agricultural labor force. It may be either increasing or decreasing in the protection efficiency due to three counteracting forces, the direct effect, the jump effect and the growth effect. The EICS appears to play a crucial role in determining how the economy’s propensity to save responds to changes in the key parameters, including those characterizing adverse shocks. For a relatively high value of EICS (above unity), we find that an increase in disaster frequency leads to a decline in the saving propensity, implying that soil conservation measures and current consumption increase at the expense of capital-soil stock expansion. An increase in the damage intensity of shocks leads to an increase in both protection measures and saving propensity. Consequently, an increase in the disaster frequency and in the damage intensity, while both having a positive impact on preservation measures, have diverging effects on the propensity to save and thus on how consumption possibilities are spread over time. Population size has a positive effect on growth but intensifies negative soil shocks. Finally, a high soil risk exposure causes a lower economic growth, even when capital markets are fully integrated and the world interest rate is given. Soil conservation has recently reemerged as an important issue in policy debate, especially for world food security and for developing economies where a significant share of population still relies on agriculture for subsistence. The present article provides a theoretical foundation for the analysis of optimal growth and for the formulation of management and policy prescriptions with respect to soil expansion and preservation nexus."],["We investigate the impact of land use regulation on housing vacancy rates. Using a 30-year panel dataset on land use regulation for 350 English Local Authorities (LAs) and addressing potential reverse causation and other endogeneity concerns, we find that tighter local planning constraints increase local housing vacancy rates: a one standard deviation increase in restrictiveness causes the local vacancy rate to increase by 0.9 percentage points (23%). The same increase in local restrictiveness also causes a 6.1% rise in commuting distances. The results underline the interdependence of local housing and Labour markets and the unintended adverse impact of more restrictive planning policies. --------------------------------------------------------------------------------","To an economist it might seem self-evident that vacancies in the housing stock are a natural feature of how any market must work. There even are ‘uneaten’ apples in a well- functioning fruit market. The Labour market is very much more comparable to the housing market and virtually all mainstream economists expect to observe at least frictional unemployment when the Labour market is in equilibrium (see Pissarides, 1985; Mortensen and Pissarides, 1994; Pissarides, 1994). It is the same in any normally functioning housing market. In equilibrium there must be vacant houses as people move and ‘house-hunt’, as people die or houses wait to be demolished and sellers wait to find a buyer (Han and Strange, 2015). But this view is often not shared by those who design buildings and influence urban policy or with those who plan housing supply – at least in England. Even in what was then one of the least restrictive English Regions, the East Midlands, in calculating how much land should be allocated for housing to meet their estimate of their region's ‘housing needs’, planners argued that they could allocate less land because they assumed they would reduce the number of vacant homes: ‘The annual average housing provision reflects a number of factors, transactional vacancies in new stock (about 2%) add 7,000 to the requirement, but offset against that is an assumption that vacancies in the existing stock should be reduced by a half per-cent, which will bring 8,600 dwellings back into use.’ (Government Office for the East Midlands, 2005, Appendix 4, p. 91). It is surely true that using ones stock of capital more intensively is a way of increasing efficiency. That is just how cut price airlines operate: they keep their seats full and their aircraft in the air. They, however, had an analysis of how to achieve this. They did not just assume planes would spend more of their lives in the air and seats would be fuller. Unless we understand why houses are vacant we cannot rationally hope to reduce the number of vacant houses just by being more restrictive. To help improve our understanding of the factors which determine vacancy rates in the housing market, this paper investigates the causal, albeit reduced-form, impact of regulatory restrictiveness. Moreover, since housing and Labour markets are interdependent, we also investigate the related issue of how local regulatory restrictiveness affects the average commute distance of those working in the jurisdiction. These are not the only outcomes of greater regulatory restrictiveness. We find that there are other measurable effects apart from raised house prices, all apparently responses to poorer housing market matching (discussed below); more households are in temporary homes, crowding is greater, and in-migration lower. These results stem from the insight that policy imposed restrictions on housing supply may have two opposing effects.1 The first of these we call the ‘opportunity cost effect’. Tighter restrictions on supply imply fewer available houses and therefore more demand pressure for existing homes, increasing house prices and thus the opportunity cost of keeping housing empty. This will lead to a lower vacancy rate all else equal. If this ‘opportunity cost effect’ was the only effect at work, tighter supply constraints should unambiguously lower vacancy rates. There is however a second effect, which we refer to as the ‘mismatch effect’. Tighter supply constraints not only reduce supply of new houses but also influence the composition and adaptability of the bundle of attributes of both the existing housing stock and those of new built homes. Over time the structure of households' demand for housing attributes changes because incomes rise, the demographic structure of the population changes and preferences themselves may change. For example, as real incomes rise, so does the demand for certain attributes depending on the varying income elasticity of demand for them.2 In addition there may be demographic changes such as an increase in the proportion of single adults, which mean that market preferences change. If the attributes of the housing stock, as a consequence of planning constraints, cannot, or can only more slowly adjust to these changes on the demand side, matching the demand for housing attributes with the supply of those available will inevitably become more difficult. Hence, in line with Wheaton (1990), mismatched households may have to stay longer in a less restrictive housing market while searching in a more restrictive one, implying a relatively lower vacancy rate in the less restrictive market and a higher vacancy rate in the more restrictive one. Mismatched households may also take temporary accommodation and search for longer in more restrictive markets or have to search further afield for a suitable home; they become mismatched on the locational characteristics of houses implying longer commutes. Our aim in this paper is to determine the net effect of these two opposing forces – the opportunity cost effect versus the mismatch effect – in order to identify the role that regulatory restrictiveness plays in determining the vacancy rate in local housing markets. To do so, we analyse panel data on housing vacancies from 1981 to 2011 for 350 English Local Authorities (LAs), the basic local jurisdictional unit that implements planning policies and approves or rejects individual planning applications. One key concern in this analysis is the endogeneity of local planning restrictiveness. The stylised fact that policy makers and local planners may respond to higher vacancy rates by restricting supply suggests possible reverse causation. Regulatory constraints may also be endogenous to unobserved demand factors (Hilber and Robert-Nicoud, 2013; Davidoff, 2016) and those demand factors may directly affect vacancy rates. To account for possible reverse causation and omitted variable bias and thus identify the causal effect of regulatory restrictiveness, we employ an instrumental variable strategy by exploiting specific features of the British voting system which induces a substantial ‘randomness’ of seats won (or lost) beyond the vote share. That is, we use the share of Labour seats in LAs, controlling for the share of Labour votes in a flexible way, as an instrumental variable to identify local planning restrictiveness. One could query this identification strategy because, for example, the political composition of an LA could influence local government expenditures and those, in turn, might influence house prices and vacancy rates. Based on a series of placebo regressions, we show that these alternative explanations do not plausibly invalidate the main conclusions. Our two key empirical findings are as follows. First, when we naively look at cross-sectional data, we find a negative relationship between more restrictive local planning and local vacancy rates, superficially appearing to confirm “planners' assumptions”. However, when we (i) use first differencing and so control for time-invariant unobservable characteristics, (ii) properly account for the endogeneity of restrictiveness by instrumenting for it and (iii) control for other relevant factors, more restrictive places have a significantly – and substantially – higher vacancy rate. That is, the underlying causal relationship appears to be exactly the opposite to that which planners assume. Based on our most rigorous empirical specification, a one standard deviation increase in local regulatory restrictiveness causes the average local vacancy rate to increase by about 0.9 percentage points (23%). Second, we find that regulation-induced mismatch has spatial implications for Labour markets. Workers with jobs in LAs with more restrictive planning have to search for housing they can afford and match their preferences further afield; so they are more likely to be locationally mismatched and have to commute further. Using a similar approach to that used for investigating the underlying relationship between the vacancy rate and restrictiveness we find that a one standard deviation increase in local regulatory restrictiveness causes an increase in average commuting distance of some 6.1%. We also provide additional suggestive evidence relating to other proxies for mismatch, such as the share of crowded or non-permanent properties and the share of migrants. Our findings, therefore, strongly suggest that tighter local planning restrictiveness not only leads to less efficient housing market matching but also this effect dominates the opportunity cost effect, resulting in higher local vacancy rates overall and longer average commutes. Hence, local efforts to reduce the number of vacant homes by imposing supply restrictions have three unintended effects: they increase the local vacancy rate and they increase the average commuting distance of those who work in the jurisdiction – thereby causing a welfare cost. In addition, as the literature shows, they increase local house prices (see e.g. Cheshire and Sheppard, 2002; Glaeser and Gyourko, 2003; Hilber and Vermeulen, 2016). We proceed as follows. In the next section we discuss in more depth the link between land use regulation and mismatch in the housing market and how that affects the local vacancy rate and the average commuting distance. We then describe our data and set out our main results. The final section draws conclusions.","The price of housing services is a function of both demand and supply in the relevant local markets. Various empirical studies document a positive effect of regulatory restrictiveness on house prices (Cheshire and Sheppard, 2002; Glaeser and Gyourko, 2003; Glaeser et al., 2005a, 2005b; Quigley and Raphael, 2005; Ihlanfeldt, 2007; Hilber and Vermeulen, 2016). What these studies do not consider is the fact that, on the seller's side, it takes time to sell a house and, on the buyer's side, search for a new house is costly too. These search frictions lead to housing vacancies (Merlo and Ortalo-Magné, 2004; Han and Strange, 2015). It has been documented – and our data also suggest – that housing vacancies are not constant across space and time and depend on the characteristics and preferences of households living in a housing market, as well as on those of the location such as characteristics that are systematic of persistently weak housing demand (Rosen and Smith, 1983; Gabriel and Nothaft, 1988; Gabriel and Nothaft, 2001; Deng et al., 2003; Molloy, 2016). However, the impact of land use restrictions on housing vacancies has not yet been studied. In the context of this paper we use data on local jurisdictions – in Britain, LAs – which we refer to as local housing markets.3 On the demand side, households often search in a local housing market while still living in another local market, for example due to changes in where they work (Mulalic et al., 2014; Koster and Van Ommeren, 2017). On the supply side, the characteristics of housing are the result of both the characteristics of new build housing and the adaptation of the characteristics of the existing stock. The degree of regulatory restrictiveness influences the characteristics of new construction and of the existing stock in very great detail. Both new construction and significant changes to the characteristics of existing houses – converting loft to living space, for example – likely require ‘development control’ permission. This is the responsibility of the LA's Planning Committee made up of locally elected politicians. This decision making process tends to be politicised and unlike a Zoning or Master Planning system, such as in force in the US or in most of Continental Europe, decisions are not very predictable. As noted in the introduction, planning induced housing supply restrictions will have two opposing effects on the housing vacancy rate: an ‘opportunity cost effect’ and a ‘mismatch effect’. The opportunity cost effect works via restrictions of supply reducing the availability of land for development (see for example Cheshire and Sheppard, 2005, or Hilber and Vermeulen, 2016). This reduces the rate of new building and so over time the size of the stock of housing relative to demand within the market. This, all else equal, increases prices and thus the opportunity cost of keeping housing vacant. The effect of this is unambiguously to reduce vacancy rates. It will also be likely to increase price volatility. However, more restrictive planning policies will also change the bundle of attributes on offer and, other things equal, slow the rate of adaptation of housing characteristics to changes in the structure of demand with respect to them – the mismatch effect. The latter effect is expected to increase vacancy rates. This will come about via two separate forces, one working on the characteristics of new build and the other on the adaptation of the characteristics of the existing stock of houses. The first force may imply that new build houses become smaller, more distant from jobs and are more likely to be in the form of flats or terraced houses, because there is less land available for dwellings. The second force arises because the structure of demand for housing characteristics changes over time and to accommodate this, the characteristics of the existing stock of housing need to be constantly adjusted. For example, entry to the best state schools in Britain is determined by the exact location of houses. As the relative standing of different schools changes over time, people seeking to ‘buy’ entry to better state schools will want more bedroom space in the best schools' catchment areas. However, the more restrictive is the LA, the more difficult it will be to adapt existing houses to provide more space or for developers to build additional family housing near better schools. Another example is that as more cars have been bought (car ownership has increased 13-fold since the current form of land use planning in England was introduced in 1947 and doubled since our vacancy data starts in 1981, Department of Transport, 2013), the demand for garages and off street parking has increased. Such examples of ways in which the demand for housing attributes changes over time could be increased almost indefinitely. However, what it means is that if the supply and demand for the structural characteristics of housing are to be efficiently matched to each other, there will need to be constant adaptation of the characteristics of the existing stock of houses. So more restrictive LAs will slow the adaptation of the existing stock to (changes in) the structure of demand for housing attributes. Over time, in more restrictive LAs the characteristics of new and existing housing available will be less adapted to preferences of households. Hence, other things equal, if people have a (strong) idiosyncratic preference for locations and house type (e.g. a double-earner household with children that needs at least two-bedrooms and garden space), they will spend more time searching for housing that matches their preferences. When households live in a less restrictive housing market while searching in the more restrictive local market, this will imply a decrease in the vacancy rate in the former and an increase in the latter housing market.4 In other words, given idiosyncratic preferences, households stay longer in the ‘wrong’ places. This may imply that younger people live longer with their parents in ‘crowded’ properties, or that households are induced to stay in temporary accommodation while searching. Because the housing stock does not match their current preferences, this implies a higher vacancy rate, other things equal, in the more restrictive housing market. We provide some evidence on some of these different symptoms of mismatch in Section 3.6 where we look at how regulation influences the share of non-permanent homes and ‘crowded’ properties. We discuss also a more obvious measure of mismatch in Section 3.4: commuting distances from the workplace in the LA. We do indeed find that for workers in more restrictive LAs commuting distances increase significantly. This result is consistent with house hunters finding it more difficult to match their preferences in more restrictive local housing markets so becoming ‘mismatched’ locationally. This has interesting implications for the boundaries of local Labour markets – they appear to be determined not just by transport costs but also by local planning policies – and how these affect the total supply of housing and the supply of individual housing characteristics. Because households may decide not to move to the desired more restricted place, the share of in-migrants is expected to be lower in the more restrictive housing market. We provide evidence for the latter in Section 3.6. The well-documented fact that tighter local regulation leads to higher prices is indicative that the opportunity cost effect may be important in determining local vacancy rates. However, we lack evidence on the importance of the offsetting mismatch effect. Thus the net effect of local regulatory restrictiveness on local vacancy rates is ambiguous. The empirical analysis that follows aims to identify this net effect while eliminating alternative explanations. One may question whether changes in vacancy rates are a sufficient statistic when one is interested in the welfare effects of land use regulation. We do not argue that vacancy rates in general are a sufficient statistic. However, when one is specifically interested in the change in welfare due to an increase in mismatch caused by more restrictive local planning, an increase in the vacancy rate (beyond the natural rate) is a sufficient statistic for the former. In line with a large Labour market literature on matching, and as demonstrated by Koster and Van Ommeren (2017), from a welfare perspective housing search may be either too low or too high. However, when search is too low, that is caused by an externality for sellers: if buyers search more they will find a suitable property sooner, thereby reducing the time on the market and the associated costs for the seller. Buyers do not take this into account when increasing search effort. However, a regulation-induced increase in vacancy rates while increasing search effort, neither reduces sales times, nor increases matching quality, so the welfare effects are unambiguously negative. This is important because planning policies that aim to reduce vacancies by reducing new construction but end up leaving more houses empty, cause an under-utilisation of a major capital asset. According to ONS by the end of 2013 houses accounted for 61% of the UK's net worth: up from 48.7% 20 years previously (ONS, 2016). So the capital stock represented by housing is very significant indeed and so its underutilisation represents a significant economic inefficiency. The effects on commuting are important in their own right as again they represent a welfare loss resulting from increased difficulty of matching. Of course planning policies per se have the potential to increase social welfare via correcting market failures and we are not claiming here that our evidence on the effect of land use regulation on vacancy rates and commuting distances in isolation suggests that local planning restrictions reduce net welfare. However, there is evidence at least for the UK and the US that an increase in the restrictiveness of planning policy (from current levels) has a net negative effect on welfare (see Cheshire and Sheppard, 2002 for the UK and Turner et al. (2014) for the US). In this context, our finding that more restrictive local planning increases the local vacancy rate and commuting distance via raising mismatch in the local housing market, adds to the alleged net negative effect on welfare. Both effects (on vacancy rates and commuting distances) have been ignored in the literature so far. Data and descriptive statistics ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our data come from several sources. The vacancy rates are from the UK Census for the years 1981, 1991, 2001 and 2011.5 For the first three Census years we have information on the number of vacant dwellings and we are able to distinguish between primary dwellings and second homes.6 The 2011 Census reported only information on the number of unoccupied dwellings including second homes. To estimate vacancies for 2011 in the most consistent way possible, therefore, we assume that the share of second homes remained constant between 2001 and 2011. In a robustness check we use an alternative dataset for vacancy rates (available for 2001 and 2011 only) to test whether our findings are sensitive to this adjustment. The latter dataset is provided by the Department of Communities and Local Government (DCLG) using the LA returns for the Council Tax.7 Our measures of regulatory restrictiveness come from the DCLG's Planning Statistics. Following the literature, our key measure is the refusal rate for major residential projects available for each LA on an annual basis. The refusal rate for ‘major’ projects is defined as the share of applications for residential developments of ten or more dwellings that is refused by an LA in any year during the process of ‘development control’. We calculate this for each LA using data on all applications and refused applications of major developments for the Census year itself plus the two years preceding it.8 In what follows we call this variable the refusal rate. As a proxy for local (housing) demand we use LA-level male weekly earnings for the period from 1981 to 2011. Our earnings data come from the Annual Survey of Hours and Earnings (ASHE) for 2001 and 2011 and from the New Earnings Survey (NES) for 1981 and 1991. We obtained the ASHE data at the LA-level but the NES data for earlier years are only available at the county and London borough level. We then geographically matched all earnings data to the LA-level and deflated the nominal earnings figures by the Retail Price Index to obtain real earnings. For more details on the data and procedures used, see Hilber and Vermeulen (2016). A number of other factors may influence vacancy rates, in particular housing tenure, demographics and socio economic characteristics. We obtain these control variables from the Population Censuses. Our list of controls includes the local homeownership rate. Homeowners tend to move less often than renters, and this is likely to be reflected in higher vacancy rates for rental housing. We also control for the share of council housing. Because rents of council houses are usually below market value, there are waiting lists for them. This is likely to imply a shorter duration of vacancies (Pawson and Kintrea, 2002). However, this effect could be offset to the extent councils have less efficient housing management. The Population Censuses also provide data on the share of people between 30 and 64 and the share of elderly, 65 and over. Young people may be more flexible in their housing choices than older people, and they may be less selective because they are more income constrained or have lower search costs (perhaps because of lower opportunity costs of time) leading to lower vacancy rates in LAs where there are proportionately more young adults. On the other hand, younger people tend to have a higher mobility rate, leading to higher vacancy rates. The mortality rate is of course highly correlated to the share of elderly. Death frequently implies that houses become vacant and, moreover, because of probate and perhaps other reasons (the new owner may not be a local resident or the house has suffered a period of neglect so is more likely to need refurbishment) houses that become vacant on the death of their owner are likely to remain vacant for longer. Other control variables derived from the Population Censuses are the share unemployed, the share of highly educated, and the share of residents with permanent illnesses. As a proxy for mismatch and as a significant focus of interest in its own right, we gather data on the average commuting distance from the workplace for all the Census years. The data provide us with the share of people per commuting distance band (0–2 km, 2–5 km, etc.). We then calculate the average commuting distance by taking the midpoint of each category and weighting it by the number of persons in each category. We further gather data on other variables that may relate to spatial mismatch from the Census, such as the share of crowded properties, the share of shared properties, the share of migrants and the share of non-permanent dwellings. Our instrumental variable strategy employs information on the political composition of the LA and local vote shares. We obtained the local election data from various sources: (i) the British Local Election Database (1889–2003) compiled by Rallings and Thrasher (2004), (ii) the Local Election Handbooks (1999 to 2008), (iii) the Local Elections Archive Project (LEAP) (2006 to 2010) and (iv) the BBC (2009 to 2011). We do not have data on four LAs, so these are excluded from the analysis, leaving us with a sample of 350 LAs and four Census years (1981, 1991, 2001 and 2011).9 Since it might be argued that turnout is unrepresentatively low at local elections in the sensitivity analysis, we also use data on general elections, by matching each Census year to the nearest general election year (i.e., 1983, 1992, 2001 and 2010). The LA-level share of votes for the Labour party in the general elections is derived from the British Election Studies Information System. For more information on the election data, see Appendix 1. We also gather data on net local expenditures from the Chartered Institute for Public Finance and Accountancy (CIPFA) annual reports on finance and general estimates available for each LA. We choose spending categories that remain robust over time, such as spending on education, personal and social services (such as social care), highways, housing services, local planning and the total local net expenditures. Because these are net expenditures, they may be negative in certain instances. We express the local expenditures in £ per head of the population. We note that for education, personal and social services, and highways, the largest share of the spending is done at the county level. Although LAs have some freedom to spend extra money, we add the net spending per head at the county level to the local expenditures in these categories (otherwise most values would be zero). This also explains why the total local expenditures of an LA are lower than, for example, the net spending on education: the total expenditures only refer to expenditures by the LA itself. In a few instances data are missing for individual LAs (in particular for a dozen LAs in Greater London in 1981). In cases such as this we impute the missing values from the average spending in a county, implying (a small) measurement error. However, in the placebo regressions in Section 3.4 the spending is the dependent variable. As long as this measurement error is random, it does not affect the estimated coefficients. We obtained data on house prices from the Land Registry (1995–2011) and the Council of Mortgage Lenders (CML) (1974–1995). We do so by taking account of the composition of sales in terms of housing types by adopting a mix-adjustment approach (see Wall, 1998). The real price index is obtained by again deflating the nominal series with the Retail Price Index. We then use the price index to create a measure of local price volatility; for more information see Hilber and Vermeulen (2016). Table 1 presents the descriptive statistics. The average overall vacancy rate is about 4%. The vacancy rate in 2011 was 3.6%. This is only slightly lower than in the United States, where it was 4.5% in 2012. This might seem surprising when one takes into account the enormous excess supply of housing in the wake of the Great Recession that made housing extremely affordable in the US. In Fig. 1 we plot the cross-sectional relationship between the vacancy rate and house prices. Vacancy rates are somewhat lower in areas with high prices (ρ = − 0.246), consistent with the opportunity cost argument discussed above. There is little response to the housing market cycle; the correlation between the change in the vacancy rate and the change in house prices is very low with ρ = − 0.069. We map the average local vacancy rates over the sample period in Fig. 2. There is meaningful variation in vacancy rates over space. They are generally higher in the less prosperous north. Cities like Liverpool and Bradford, which respectively relied on traditional port and port-related manufacturing or textiles, experienced decline from the 1950s. Apart from high unemployment and lower earnings there was outward migration tending to generate a more obsolete housing stock and higher housing vacancy rates. Also in areas where mining was historically important (in County Durham and Lancashire for example), vacancy rates tend to be higher. We implicitly control for these geographical differences in the industry composition by first differencing our empirical specification, thus capturing all time-invariant characteristics that vary over space. The inclusion of the first difference in the local unemployment rate as a further control should effectively control for any relevant influence of changes in industrial structure on housing vacancy rates. Refusal rates over the last 30 years have been clearly highest in the Greater London Area and in the south of England and lowest in the north of the country (Fig. 3). The south of England has not only been economically considerably more successful than northern regions over the period, but it has (perhaps relatedly) had much tighter planning restrictiveness. This – despite strong housing demand – has constrained the growth of housing supply in southern England relative to the north. Econometric framework and identification ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ This specification only partly addresses the first endogeneity concern because there might still be correlation with unobserved shocks. For example, in locations with increasing demand, house prices and regulatory restrictiveness may increase simultaneously. Anecdotal evidence suggests that in England regulatory restrictiveness is strongly pro-cyclical. In times of high demand, planners reject more proposals in attractive areas, perhaps to avoid what they perceive as a threatened ‘oversupply’ and perhaps because the system cannot cope with the workload. Because housing supply takes time to adjust, this will lead to lower local vacancy rates during boom periods. This again implies that α is likely strongly downward biased if we estimate (2) by OLS. We therefore have to find an instrumental variable to identify refusal rates that is uncorrelated with local unobserved shocks. Bertrand and Kramarz (2002) exploit the cumulative representation of each political party at regional level as an instrument for how restrictive French départments are likely to be towards new retail entrants to document that stronger deterrence of entry by regional zoning boards increased retailer concentration and slowed down employment concentration. In a similar vein, Cheshire et al. (2015), Sadun (2015) and Hilber and Vermeulen (2016) use the share of party representation at LA-level as an instrument to identify the impact of local regulatory restrictiveness on, respectively, retail store-level output, the entry of large retail stores, and house prices. Our study, to the best of our knowledge, is the first to (i) exploit the fact that the particular details of the electoral system of local government in England generates random variation in local party influence relative to vote share and (ii) use this exogenous variation to identify the impact of local regulatory restrictiveness on housing and Labour market outcomes. Our specific instrument is the change in the number of seats for the Labour party between election years close to the Census years. Traditionally Labour voters and politicians have been less opposed to new residential construction than their Conservative counterparts. Labour councillors typically represent a part of the population that has less housing equity and so is less subject to NIMBY pressures aiming to protect house values. Labour councillors are also likely to be more interested in the job generating effects of construction. Thus, we can expect that an increase in the share of Labour seats may induce LAs to become less restrictive, yet, a change in the share of Labour seats should not directly affect the local vacancy rate other than through any effects it has on planning restrictiveness. To make the reasons for our choice of instrument clearer it may help to have a brief explanation of the English local government system and of the mechanics of the local electoral system.13 There is a recent, succinct account available in Sandford (2016). The English system is very heterogeneous but has certain common features. It is highly centralised. As is discussed below, for most purposes, LAs are hardly more than agents of central government with legal obligations but little fiscal autonomy; property taxes, for example, are essentially national taxes (see Section 3.5) and there is no local income tax. The major area of autonomy is with respect to planning decisions. There are a total of 354 LAs of three different types: County Councils; District Councils; and Unitary Authorities. The 125 Unitary Authorities – which have the fullest range of functions – include the London and Metropolitan Boroughs. Members of the Councils of all LAs are elected for ‘wards’ – geographical subdivisions of the council's area – and all elections are conducted on a ‘first-past-the post’ system. Some LAs have single-member wards, others multi-member wards. Voters vote for as many councillors as there are vacant seats at any date there is an election. So if, for example, all members of a three-member ward face re- election on the same date, the elector will have three votes. To complicate matters further some councils elect all their members every three years; others elect one third of their members at any given election while a few elect half their members each year, so political control can change rapidly. Over the period of our analysis there were three main political groups: The Labour, Conservative and Liberal-Democrat parties. As with any first-past-the post system the party winning a seat contested by three parties may have a minority of votes; in wards where one party is dominant, their candidate may have only token opposition or even none at all. Equally, councils may be quite evenly split in terms of vote shares for the different parties. Thus there are two independent reasons why the share of votes at any election and the share of seats on the council may differ. The first is just the way that the first-past-the-post voting system works when there are three parties all gaining significant vote shares but those shares are highly variable between constituencies. The second is that in many councils only one third or a half of the elected members are voted for at any election. So the composition of the council is a moving average of past votes. And, of course, the share cast for any party may change significantly over the course of even a year. The result is that the share of votes and the number of members on a council is not perfectly correlated—for the purposes of our identification strategy important: the discrepancy between the two can be considered random. The correlation between the share of Labour votes and seats, for example, is 0.77. As is explained in Appendix 1, the variable we use for ‘seats’ is the closest measure we can find for ‘seats controlled on the council’ so allows for the fact that in many councils only a third or a half of members are elected in any given election. To illustrate the random element of seats won beyond vote shares, in Fig. 4, we provide a scatterplot of the share of Labour votes and share of Labour seats from 1978 to 2011 for each year and LA. Not surprisingly there is a strong positive correlation between seats and votes. However, below a vote share of about one third, any vote share translates into a less than proportional number of seats (denoted by the dashed line). In a number of cases, votes did not translate into any seats at all. Furthermore, because of the first- past-the-post feature of the system, above a vote share of about one third, an increase in the vote share leads to a more than proportional number of seats. If the Labour party has a vote share of > 70%, this usually implies that all seats are assigned to the Labour party. By including LA fixed effects ηℓ, we control for all linear trends caused by unobservable factors, which increases the likelihood that changes in the instruments are uncorrelated with Δϵℓ, t. Despite the fact that we identify changes in regulatory restrictiveness from the random component generated by the particular features of the English local government system, one might still be concerned that greater Labour representation does not only affect regulatory restrictiveness but may also affect other local variables that may separately affect local vacancy rates. That is, the exclusion restriction may be violated. To address this crucial concern we first argue and provide evidence to support the claim that — unlike in countries with decentralised government structures — LAs in England, especially since 1972, have very little fiscal discretion or power other than making planning decisions.14 Next, we show that even for those LAs — Unitary Authorities — that provide more local services than others, the effect of a random increase in the local Labour representation has a very similar effect on local restrictiveness. This suggests that the relation between the share of Labour seats and local restrictiveness may not be significantly biased by other local policies and services that may be correlated with both regulatory restrictiveness and local vacancy rates. Most reassuringly, when we run a battery of placebo (first-stage) specifications, in which we replace the change in local refusal rates with changes in local expenditures — our placebo variables — we find no significant relationship between share Labour seats and these placebo measures (see Section 3.5). This is in contrast to a strong and statistically significant negative relationship between the random change in the share of Labour seats and local refusal rates. Overall, these results provide a strong indication that the exclusion restriction is not violated and our identification strategy is valid. Results for housing vacancies ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We start by ignoring any potential endogeneity issues and simply regress the vacancy rate on the refusal rate of major residential projects (Eq. (1)). From Fig. 5 we can see that the cross-sectional relationship between the major refusal and the vacancy rates is negative. The regression line implies that a one standard deviation increase in refusal rates is associated with a 0.23 percentage point decrease in the vacancy rate (s.e. 0.040). This naïve correlation provides ‘common sense’ evidence supporting the view that vacant houses can be ‘regulated away’. However, the quantitative impact is not very large. Table 2 reports estimates for Eqs. (2) to (6). In the cases of Eqs. (3) to (6) these are the second stage results of our IV-estimates. In column (1) we regress the change in the vacancy rate on the change in the refusal rate still ignoring potential endogeneity issues (Eq. 2). We first difference controls to offset for any time invariant omitted characteristics such as differences in income levels across LAs. We see that even without instrumenting for the refusal rate or adding control variables, the relationship between (the change in) planning restrictiveness and (the change in) the vacancy rate is no longer negative and statistically significant. However, because of the endogeneity concerns discussed above, the coefficient on the refusal rate cannot be interpreted as a causal effect. So in column (2) we include LA fixed effects. The coefficient on the change in the refusal rate variable now becomes positive and statistically significant at the 10% level. In column (3) of Table 2 we add further controls as discussed in Section 3.1 above. The estimated coefficient for the change in the refusal rate is hardly affected, although it is not statistically significant at conventional levels anymore. The control variables often have a statistically significant impact on the change in the vacancy rate with the anticipated sign. For example, areas with an increasing share of elderly people or of council housing experience an increase in the vacancy rate. Also, areas with an increasing unemployment rate, from which people may have been tending to move away, experience an increase in the vacancy rate. In areas with a rising share of highly educated people, vacancies tend to decrease. Still, however, regulatory restrictiveness is likely measured with error (because developers may not apply in the first place in more restrictive places). It may also be correlated with unobserved shocks. Moreover, we should address the potential reverse causality issue that higher vacancy rates may induce policy makers to be more restrictive. We therefore instrument for the change in the refusal rate with the change in the share of Labour seats in column (4). This specification corresponds to Eq. (3) above. Kleibergen-Paap F-statistics indicate that there are no issues of weak identification of regulatory restrictiveness. The results suggest that a one standard deviation increase in the refusal rate leads to an increase in the vacancy rate of 0.82 percentage points. As noted in the previous subsection, one objection to the instrument is that it may be correlated with unobserved characteristics of the area. To control for this, we include LA fixed effects in column (5) – corresponding to Eq. (4). The coefficient on the refusal rate hardly changes and remains statistically significant at the 5% level. Column (6), corresponding to Eq. (5), includes the same range of control variables as in column (3). This makes almost no difference to the estimated coefficient of primary interest. One might still be worried that changes in the share of Labour seats are correlated with unobservable shocks (e.g. gentrification) that simultaneously have an impact on voting behaviour and vacancy rates. So in column (7) we estimate our final model (6). That is, we additionally include a flexible function of changes in the share of Labour votes in local elections, approximated by a fifth-order polynomial to isolate the impact of voting behaviour caused by any change in the demographic and socio-economic composition of the LA from political power (measured by seats). In the sensitivity checks, discussed below, we report results for different orders of polynomials. Reassuringly, the estimated effect of regulatory restrictiveness in column (7) is very similar to the previous specifications. The instrument is somewhat less strong (with a Kleibergen-Paap F-statistic of 8.2). Still, we find a positive and economically meaningful effect of regulatory restrictiveness on the vacancy rate: a one standard deviation increase in the refusal rate increases the vacancy rate by 0.90 percentage points. Due to the correlation between changes in the Labour vote shares and changes in the share of Labour seats, it is no surprise that the coefficient is now only statistically significant at the 10% level.15 In Table 3 we report the corresponding first-stage estimates: a standard deviation increase in the share of Labour seats leads to a decrease in the refusal rate of 0.26–0.34 standard deviations. It is notable that the first-stage coefficients of the change in the share of Labour seats instrument are highly statistically significant and are hardly affected by the inclusion of LA fixed effects and other control variables. If we include vote share controls, the coefficient on change in Labour seats becomes slightly lower, but it is still statistically significant at the 5% level. Results for commuting distance ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In Section 2 we hypothesised that a positive relationship between restrictions and vacancy rates might be explained by increased mismatch. There are few obvious measures of mismatch but for reasons discussed earlier we think that the ‘average commuting distance from the workplace’ does not only provide a useful measure but can potentially illuminate in a useful way the underlying interrelationship between spatial housing and Labour markets. Moreover longer commutes unambiguously signal a welfare loss. One of the most important characteristics of a house is its location with respect to jobs. It seems reasonable therefore that the average commuting distance from the workplace should capture mismatch in this dimension of housing characteristics for any given housing market. In principle, households have a preference to live close to their workplaces. If regulatory restrictions make it more difficult for people to find a home ‘matched’ to their preferences on other characteristics close to work, their search takes longer and will extend further. This adaptation of search behaviour implies, other things equal, vacancies will tend to be higher in the more restrictive LAs and lower in neighbouring, less restrictive ones, as workers become more locationally mismatched.16 We provide evidence that regulation also has an impact on other proxies for mismatch in Section 3.6. Table 4 replicates Table 2 except that the log of average commuting distance replaces the vacancy rate as the dependent variable. The results very closely parallel those for vacancies. Those in the first column suggest that commuting distance is not influenced by regulatory restrictiveness in the LA of the workplace. When we include LA fixed effects and demographic control variables in columns (2) and (3), the results are still statistically insignificant. This is not too surprising as the refusal rate is highly endogenous and correlated with other factors that might explain commuting distances. For example, places that have become denser might have tended to become more restrictive (Hilber and Robert- Nicoud, 2013), but denser places also might have shorter commutes because jobs and households are located closer to each other. In column (4) we therefore control for other factors that might be correlated with the refusal rate by instrumenting for the change in the refusal rate with the change in the share of Labour seats, as in Table 2. This reveals a positive and significant effect. As restrictiveness in the LA in which a worker is employed increases so does the average commuting distance: a one standard deviation increase in the refusal rate in the workplace LA increases the commuting distance of its employees by 8.5%, a non-negligible effect. The effect becomes somewhat smaller (5.8%) when we include in column (5) LA fixed effects. The effect continues to be essentially the same when we add further control variables in column (6) and a flexible function of the share of Labour votes in column (7), with 5.9 and 6.1% respectively. In the last column the effect is somewhat imprecisely estimated and only statistically significant at the 14% level.17 On the other hand, the results are consistent in pointing towards a meaningful effect of regulatory restrictiveness on commuting distance, and therefore increasing the spatial mismatch between home and work locations. Evidence in support of the identification strategy ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The central assumption of our identification strategy is that a random increase in the local representation of the Labour party only influences local regulatory restrictiveness and does not separately affect other local decisions that may themselves be correlated with local vacancy rates. If this assumption does not hold, then the exclusion restriction is violated. In this context it is important to first re-emphasise that LAs in England have considerable discretion over local planning decisions. However, unlike, for example, in the US, they have almost no local fiscal resources and very limited ability to determine local public service levels since these are largely set by central government. To a large extent local jurisdictions – LAs – are simply agencies of central government in charge of delivering services locally to regulated national criteria using a dedicated budget stream for each purpose (such as education, social services, local roads and street cleaning or refuse collection). These direct grants of revenue from central government have over the past 75 years or so accounted for some 80% of LA spending. Revenues from taxes on residential property are subject to revenue equalisation across LAs so that local resources are ‘needs’ based and revenues to LAs become, with only a short delay, independent of the local property tax base: commercial property taxes since 1990 have been a national tax. Over the first half of the 20th Century LAs did gain an increasing role in the provision of social housing – ‘council housing’. This was not initially designed to be ‘safety net’ housing for the very poor and deprived but rather public housing for those not wanting or not able to become owner occupiers. At the start of the 20th Century owner occupation as a tenure accounted for only about 10% of all housing. Council houses were built to common nationally set design guidelines but delivery was under local control as was the setting of rents (subject to the wider financial rules governing LAs' budgets). This housing role of local government peaked in the 1950s and 1960s and declined thereafter. By the late 1980s the construction of LA (council) housing had almost stopped and local powers to set rents had been all but abolished. The two key changes were the Housing Finance Act of 1972, which required LAs to set nationally defined ‘fair’ rents, and the introduction of the Right to Buy for LA tenants, introduced in 1980.18 Tables 5 and 6 provide empirical evidence to support the above account of how limited the powers of LAs in England are except for their powers over development. In Table 5 we test whether the provision of public services is affected by the political composition of an LA more directly by looking at expenditures on services, such as education, personal services (e.g. social care), highways, (social) housing and planning. We then run a series of placebo tests by replicating the preferred first-stage results in column (7) of Table 3, but using different dependent variables. Instead of the change in the refusal rate we use the log change in the following variables: education expenses, social services expenses, highway expenses, local housing expenses, local planning expenses, and total local expenses.19 Our identification strategy stipulates that a random change in the share of Labour seats should only significantly reduce the refusal rate but should not affect the expenditures on various public services. Indeed, that is what Table 5 reveals. The coefficients for the change in the share of Labour seats on expenses are in all cases highly statistically insignificant. The only exception is a weakly significant effect on the total local expenditures (p-value = 0.15). The coefficient seems to suggest that a one a standard deviation increase in the share of Labour seats increases local spending by 8.75%. Although the effect is not statistically particularly strong, one may be worried that the share of Labour seats has a direct impact on vacancy rates via the total expenditures of local planning authorities, which would call into question the exclusion restriction. In Appendix 2 (Table A2.1) we therefore re-estimate the preferred specification in Table 2 (column (7)), but include expenditure by category, one by one, as additional controls (columns (1) to (5)). In column (6), Table A2.1, we instead include total local expenditures as an additional control. The results show that the impact of the refusal rate on vacancy rates is essentially unaffected, even when we control for total local expenditures or for each category of expenditure individually. Planning expenses seem to have some direct negative impact on vacancy rates (column (5)): more planning expenses lead to lower vacancy rates. In column (6) we show that, even if the share of Labour seats might have a positive effect on total local expenditures, it seems that local expenditures do not have a direct impact on vacancy rates. Hence, this provides additional evidence that the exclusion restriction holds. We repeat this exercise in Table A2.2, Appendix 2 with commuting distance as dependent variable and show that the estimated coefficients are also very similar to the baseline specification. Table 6 provides further evidence supporting this narrative. First we replicate our baseline results but we allow for the coefficient on the share Labour seats to vary between Unitary LAs and all other LAs. Unitary Authorities have a wider remit in terms of the types of services they can provide. We therefore use as an instrument the change in the share Labour seats interacted with a dummy indicating whether an LA is not a Unitary Authority and control directly for the change in share Labour seats interacted with a dummy indicating whether an LA is a Unitary Authority. The first-stage results, reported in Panel A of Table 6, reveal that the coefficient on the share Labour seats is similar for the two types of LAs: the coefficient for Unitary Authorities is somewhat lower but not statistically significantly different from the coefficient for other LAs. This also suggests that the local share of Labour seats does not significantly affect the nature and quality of local service provision (which in turn might be correlated with regulatory restrictiveness and local vacancy rates). Indeed, Panel B of Table 6 shows that the baseline results for vacancy rates are essentially unaffected. Finally if we replace the vacancy rate by the log commuting distance in Panel C of Table 6, the results are again similar. This is indicative that LAs with a greater remit of service provision are not fundamentally different from those that have a more limited remit—presumably this is because even Unitary Authorities have very little discretion over services other than planning decisions. Other evidence for the importance of mismatch ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As noted in Section 2 we also investigate the impact of local restrictiveness on other symptoms or measures potentially capturing mismatch, such as the rate of new construction, the share of properties which are crowded, the share of non-permanent properties and migrants. The results are shown in Table 7. We replicate the preferred specification where we instrument for the refusal rate and include fixed effects and control variables (as in column (7) in Tables 2 and 4).20 We first investigate a symptom of restrictiveness which would be expected to induce greater mismatch: whether in more restrictive areas, despite the effect on prices, it is more difficult to build additional houses. Because the refusal rate should have an impact on the absolute number of dwellings, we regress the change of the number of dwellings (rather than logs) in an LA on the refusal rate, while additionally controlling for the number of dwellings (in levels).21 Column (1) of Table 7 shows that on average over a ten year period, a one standard deviation increase in regulatory restrictiveness in an LA reduces the number of additional dwellings by more than half a standard deviation of the growth in the number of dwellings over this time period (about 1800 dwellings per LA). To make sure that this effect is not entirely explained by larger LAs we include a fifth-order polynomial of the number of dwellings and control for house prices in column (2). In line with expectations, a higher price is associated with a slower growth in the number of dwellings, most likely because higher prices are predominantly found in already developed areas with fewer possibilities to extend the building stock (see Hilber and Vermeulen, 2016). More relevantly for present purposes, the effect of regulation on new dwelling supply is essentially unaffected once we control for prices. Another symptom of mismatch is the share of officially classified ‘crowded’ properties – properties with more than one person per room: some 2% of all properties. If households cannot find a property to their liking and cannot afford larger properties, they are likely to end up in smaller properties, perhaps staying with their partner in the parental home. Column (3) in Table 7 shows that there is indeed a positive effect of regulation on the share of crowded properties. One may argue that this results from the fact that tighter controls make housing more expensive so that households will occupy smaller homes with fewer rooms. However, when we control for house prices in column (4) the effect of regulation on the share of crowded properties is very similar, even somewhat stronger. House prices are, as expected, positively associated with crowding. The coefficient indicates that a standard deviation increase in the refusal rate leads to a 0.58 percentage points increase in the share of crowded properties (about one-third of a standard deviation). Hence, although the results are only significant at the 10% level, the implied magnitude of the point estimates is non-negligible. A further symptom or proxy for mismatch is explored in columns (5) and (6) of Table 7, the share of non-permanent dwellings. Our underlying explanation for why this measure should proxy for housing market mismatch is that in more restrictive markets there is an incentive to accept even less optimally matched housing characteristics in the short term in order to intensify and increase the efficiency of search. Living in temporary accommodation – mainly caravans or trailers - has a low switching cost associated with it and is a cheaper strategy than buying a suboptimal place to live and then reselling it when a more suitable house is found. The ease with which search can be undertaken and its effectiveness will increase if the house-hunter can be physically present in the local market (Ha and Hilber, 2013). Moreover since the chances of finding a better match in the housing market will improve with length of time spent searching then there will be a payoff to having temporary accommodation available for searchers in markets where matching is more difficult. Thus in more restrictive LAs, other things equal, matching is more difficult and the share of temporary dwellings is greater. Although by improving the efficiency of search, temporary housing may itself reduce vacancies of permanent dwellings, it is such a relatively sub- optimal form of housing we would expect house-hunters to resort to it only when there is extreme difficulty in matching their preferences to available housing supply. So we would expect that the net share of temporary housing would be positively correlated with local restrictiveness. In column (5), Table 7, we find weak but consistent evidence that more restrictive markets have a higher rate of non-permanent homes. Given that the change in non-permanent homes is a somewhat noisy and indirect proxy for mismatch, it may not be too surprising that the results are not statistically significant at conventional levels.22 Column (6) addresses the issue that this may be entirely the result of higher house prices. This does not seem to be the case; indeed the effect becomes somewhat stronger. The coefficient implies that a standard deviation increase in the refusal rate leads to an increase in the share of non-permanent homes of 0.213 percentage points (about one-third of a standard deviation). This positive relationship is certainly consistent with the proposition that more restrictive local planning increases the costs of matching would-be house buyers in the local housing market to the available permanent housing, inducing people to live temporarily in caravans and mobile homes – clearly inferior substitutes to houses. In column (7) of Table 7 we show the results for yet another symptom or proxy for mismatch. We expect that tighter regulation in an LA makes it harder for people to move into the preferred area and find a new, satisfactory property in it. Hence, we expect that the share of in-migrants, defined as households that had a different address in the previous year, is lower. We initially find only very weak evidence supporting this proposition with a p-value of 0.15. But migration responds not just to spatial differences in wages but also to house price differences across space. So in column (8) we control for house prices. Although house prices are indeed negatively correlated with the share of migrants, controlling for house prices increases the effect of regulation. Not only is the relevant coefficient larger but it is now significant at the 10% level. The coefficient implies that a standard deviation increase in the refusal rate leads to a reduction in the share of migrants of 1.53 percentage points (almost half of a standard deviation), a substantial effect in economic terms. Other potential explanations ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Are there other explanations for the positive relationship between restrictions and vacancy rates? One might, for example, expect tighter policy restricting the elasticity of supply to be associated with greater price volatility and for that to be associated with higher vacancy rates. This is because price volatility might create a (real) ‘option to wait’ (McDonald and Siegel, 1986; Grenadier, 1995, 1996). The greater the uncertainty (price volatility) the more valuable is a property owner's option to delay selling or renting out the property. Especially in markets with lengthy leases – such as office markets – this can generate high and persistent (“sticky”) vacancy rates; landlords are better off keeping their units empty. We would expect the real options argument to be less important in the British residential property markets compared to office markets, however, since demand volatility tends to be lower (people have to live somewhere) and, for rentals, the length of tenancy agreements is typically quite short (a year or even less). The real options argument, nevertheless, could be relevant since tight regulation (more inelastic supply) amplifies demand shocks, so we would expect price volatility to respond more strongly to demand volatility in places with tighter regulatory constraints. We proxy house price volatility by calculating the coefficient of variation of house prices in the year of observation and the three years that follow. Demand volatility is proxied by the coefficient of variation of earnings based on the year of observation and the three years that follow. The results (reported in Panel A of Table 8) show that at least according to the instrumental variable specifications (columns (4) to (7)), the sensitivity of price volatility with respect to earnings volatility increases with regulatory restrictiveness. This provides some evidence that regulatory constraints potentially could increase vacancy rates by increasing price volatility (i.e. increased price volatility raises the value of the real option to keep properties empty), but when we investigate this possibility further (see results in Appendix 2, Table A2.3 and discussion below) the evidence does not appear to support it. In Appendix 2 we estimate regressions where we directly control for commuting distance from the workplace, price volatility and room diversity. In Table A2.3 we show that the coefficient on the (instrumented) refusal rate variable decreases by about 15% and ceases to be statistically significant at conventional levels once we control for commuting time (column (1)). This suggests that at least a part of the positive effect of a change in regulatory restrictiveness on the change in vacancy rates is driven by mismatch in the housing market, as proxied by commuting time. When we instead control for price volatility (column (2)), interestingly, the effect of the refusal rate on vacancy rates increases somewhat in magnitude and is statistically significant. This finding could be interpreted as suggesting that the real options argument does not play an important role in explaining the positive link between regulatory restrictiveness and vacancy rates. When we include room diversity in column (3) the effect of regulation on vacancy rates again becomes somewhat stronger, although it is close to that estimated in the baseline specification. Room diversity is, as expected, positively correlated with vacancy rates; when diversity is one standard deviation higher, vacancy rates are 0.65 percentage points higher. We finally control for all variables, including the change in the share of crowded properties, the share of non-permanent properties and the share of in-migrants (column (4)). The coefficient related to the refusal rate is very similar to the previous specifications. We should note two caveats. First, our proxies, for mismatch in particular, are at best partial and, in the case particularly of the share of crowded and temporary properties, as well as the share of in-migrants, more symptoms than measures, so we would not expect these variables to fully account for the positive impact of the (change in the) refusal rate on the (change in the) vacancy rate. Second, the findings reported in Table A2.3 should generally be interpreted with caution. This is because both the commuting distance and the price volatility are likely highly endogenous, as are the alternative proxies for mismatch (some of them have the opposite sign from what might have been expected). Hence, the reported coefficients are likely biased. Still, overall, we interpret the results reported in Table A2.3 as indicating that the increased mismatch between the preferences of the local residents and the characteristics of the available local housing stock causally contributes to our key finding of a positive impact of tighter local restrictiveness on local vacancy rates. This finding is certainly plausible in the context of the extraordinarily rigid British planning system. Whether similar effects can be observed in other countries is an interesting question. Sensitivity analysis ~~~~~~~~~~~~~~~~~~~~ Finally, we conducted a set of robustness checks. We tested 1) the sensitivity of our choice of polynomial in relation to changes in the share of Labour votes; 2) the possible endogeneity of earnings; 3) whether results are driven by idiosyncratic factors associated with the Greater London Area; 4) whether the results are sensitive to including all refusals of residential development, not just refusals of applications for major developments; 5) whether a fixed effects approach rather than first differencing makes any difference to the main findings; and 6) whether the change of the Census definition of vacancies in 2011 affects our findings. The main results survive all these tests and alternative specifications essentially unchanged. The results are set out in detail in Appendix 2.","This is the first attempt to rigorously analyse the impact of land use regulation on the spatial and temporal variation in housing vacancy rates. It would come as no surprise to economists to observe that in well-functioning Labour markets there was unemployment. Workers search for jobs and employers seek (better) qualified workers. Attempting to regulate unemployment away makes no sense. Vacant houses are equivalent to unemployed workers yet, at least in Britain, policy does try to ‘regulate’ vacant homes away by using their existence to justify being more restrictive in the control of the supply of new homes and the structural adaptation of existing ones. In this paper we argue that such restrictions have two main opposing effects on housing vacancies. The ‘opportunity cost effect’ leads to a lower vacancy rate in the more restrictive housing markets because supply constraints lead to higher prices, and thus to higher opportunity costs of keeping housing vacant. The ‘mismatch effect’, however, implies higher vacancy rates in the more restrictive housing markets because households will find it more difficult to match their preferences to the characteristics of the local housing supply given their budget constraints. So search becomes more prolonged and costly and households choose to search for properties further afield and commute longer. We do indeed confirm that there is a simple negative correlation between local planning restrictiveness and local housing vacancies. Superficially this appears to support the planners' ‘common sense’ that the existence of empty houses means they can – even should – plan to be more restrictive in supply. This unconditional correlation, however, is the result of a form of joint causation. When we effectively control for unobserved characteristics at the local level by using first differencing, the negative correlation turns positive. When we further control for local linear trends, for other potential explanatory variables and account for the endogeneity of local regulatory restrictiveness, the causal effect of restrictiveness on vacancy rates is firmly positive, statistically significant and economically substantial. Our empirical analysis does not fully unpack the box of explanations. It is a reduced form telling us what the net impact of increased planning restrictiveness is on housing vacancies. If an LA becomes more restrictive, signalled by an increase in the rate of refusal of residential development proposals, then all else equal, vacancy rates increase in that LA. We also provide direct evidence that mismatch may be an important reason for higher vacancy rates in more restrictive housing markets: The more restrictive in planning terms a local jurisdiction is, the longer is the average commuting distance of workers within it. We subject these findings to an extensive range of sensitivity analyses. They survive remarkably unaltered. Welfare implications in markets with search frictions are not easy to derive. This is because in such a second best world, households may search less or more than would be welfare optimal (Koster and Van Ommeren, 2017). This paper does not set out to assess the net welfare impact of tightening the supply of housing. However, because regulatory constraints seem to worsen matching for households, the effect of tighter restrictions on vacancies and commute times unequivocally generates a welfare cost. It is the mismatch between the preferences of households and the housing stock on offer that leads, other things equal, to higher vacancy rates in the more regulated – typically more desirable – places. This is not to say that tighter regulatory constraints in some places necessarily increase vacancy rates at the aggregate level (for the whole country). However, even if on aggregate vacancy rates were unaffected by regulatory restrictiveness, such constraints will still likely cause a significant welfare loss. This is because too much housing stays empty in the most regulated, most desirable and, by implication, most productive places with the strongest demand and highest valuations for living space. So people are induced to commute further, while living in the “wrong” places. There are important implications for policy, particularly for the UK because of its extraordinarily restrictive planning system. Crucially planners should not allocate less land for development on the grounds that there are empty houses. Some vacancies are integral to the well-functioning of any market and, as our results show, trying to ‘regulate housing vacancies away’ is counterproductive. There is moreover a nice irony for advocates of the ‘compact city’. The most common policy to attempt to implement this ideal is to impose growth boundaries (make land scarcer) and be more restrictive to adaptations of the existing stock or – in the US – make it more difficult to obtain zoning ordinance waivers. In summary aiming for a compact city makes planning policy more restrictive. Our results show this will have exactly the opposite to the intended effect because average commuting distances will lengthen as residents search further afield for housing they can afford and matches their preferences."],["This paper examines the impact of exchange rate uncertainty on different components of net portfolio flows, namely net equity and net bond flows, as well as their dynamic linkages. Specifically, a bivariate VAR GARCH-BEKK-in-mean model is estimated using bilateral monthly data for the US vis-à-vis Australia, Canada, the euro area, Japan, Sweden, and the UK over the period 1988:01-2011:12. The results indicate that the effect of exchange rate uncertainty on net equity flows is negative in the euro area, the UK and Sweden, and positive in Australia. The impact on net bond flows is also negative in all countries except Canada, where it is positive. Under the assumption of risk aversion, the findings suggest that exchange rate uncertainty induces a home bias and causes investors to reduce their financial activities to maximise returns and minimise exposure to uncertainty, this effect being stronger in the UK, the euro area and Sweden compared to Canada, Australia and Japan. Overall, the results indicate that exchange rate or credit controls on these flows can be used as a policy tool in countries with strong uncertainty effects to pursue economic and financial stability. --------------------------------------------------------------------------------","The macroeconomic effects of exchange rate uncertainty, especially on trade flows, have received considerable attention since the collapse of the Bretton Woods system in 1971 and the adoption of floating exchange rates in March 1973, both in the theoretical and empirical literature (see McKenzie, 1999, for a comprehensive review). By contrast, such effects on financing activities, in particular on equity and bond portfolio flows, have not been thoroughly examined. In addition, there is a substantial literature examining the determinants of international asset transactions, but there are very few empirical papers analysing the impact of exchange rate uncertainty. For example, Bohn and Tesar (1996) found that investors tend to move to markets where returns are expected to be high. The validity of this ‘return chasing’ hypothesis has been confirmed by Bekaert et al. (2002), who found that positive return shocks lead to an increase in short-term equity capital flows using data from 20 emerging countries. Portes et al. (2001) and Portes and Rey (2005), by contrast, showed that financial transactions are explained by the gravity model at least as well as goods trade. More recently, Fratzscher (2012) found evidence that push factors were important drivers of net capital flows during the recent financial crisis, but not during the recovery period (2009–2010), when domestic pull factors were dominant, especially for emerging countries in Asia and Latin America. The underlying idea is that exchange rate volatility increases the costs of international financial transactions and reduces potential gains from international diversification by making the acquisition of foreign securities such as bonds and equities more risky, which in turn affects negatively portfolio flows across borders. Indeed, Eun and Resnick (1988) had previously shown that exchange rate risk is non-diversifiable and has an adverse impact on the performance of international portfolios. This finding is also consistent with the evidence presented in the study by Levich et al. (1999), who found, by surveying 298 US institutional investors, that foreign exchange risk hedging constitutes only 8% of total foreign equity investment. Further, Choi and Rajan (1997) reported that foreign exchange risk has a significant effect on asset returns in seven major developed countries other than the US, and that ignoring such a factor results in misspecification when analysing the integration or segmentation of international capital markets. By considering a wide range of developed and emerging market economies, Fidora et al. (2007) and Borensztein and Loungani (2011) also found that exchange rate volatility is an essential factor for bilateral equity and bond portfolio home bias. Eun and Resnick (1988) suggested that hedging through forward exchange contracts and multicurrency diversification are effective ways to reduce exchange rate risk. Glen and Jorion (1993) and Eun and Resnick (1994) provided further evidence that hedging in the forward exchange markets improves the performance of diversified portfolios of equities and bonds. Jorion (1991) also found that the exchange rate risk is diversifiable. In particular, his empirical findings provided little evidence that US investors require compensation for bearing the exchange rate risk. Gehrig (1993), instead, argued that exchange rate risk, purchasing power risk, and capital market restrictions are insufficient factors for explaining equity portfolio home bias, whilst informational segmentation plays a key role. Finally, Hau and Rey (2006) provided a theoretical framework for analysing the implications of incomplete foreign exchange risk trading for the correlation structure of exchange rate changes and equity returns, as well as exchange rate changes and net portfolio flows1; however, they did not carry out any statistical tests for the impact of exchange rate uncertainty on portfolio flows across borders. The present paper makes a fourfold contribution to the existing literature. First, it analyses empirically whether exchange rate uncertainty affects international portfolio flows and their variability using a bivariate VAR GARCH (1, 1)-in-mean framework. It is in fact the first empirical investigation of this kind, based on bilateral monthly data for the US vis-à-vis six developed economies, namely Australia, Canada, the euro area, Japan, Sweden, and the UK over the period 1988:01-2011:12. The analysis is based on longer monthly time series and differs from previous studies which focus on the determinants of home bias and international financial transactions using panel and cross-sectional techniques (see, for examples, Portes and Rey, 2005; Fidora et al., 2007; Bekaert and Wang, 2009; Batten and Vo, 2010; Borensztein and Loungani, 2011; Mishra, 2011; Mercado, 2013; Daly and Vo, 2013). We use the most common time series measure of uncertainty found in the literature, i.e. the conditional variance modelled as a GARCH (1, 1) process2 (others are the continuous volatility measure in Portes and Rey (2005), the stochastic deviation from purchasing power parity (PPP) in Fidora et al. (2007) and Mishra (2011), the standard deviation of exchange rate changes in Bekaert and Wang (2009) and the coefficient of variation of the real exchange rate in Mercado (2013)). This approach is flexible enough to allow for joint estimation of the relationship between uncertainty and portfolio flows taking into account past information on perceived uncertainty.3 Second, unlike Hau and Rey (2006), who assumed that the supply of bonds is infinitely elastic, thereby simplifying the dynamics of bond acquisitions in their model, we examine the impact of exchange rate uncertainty on the individual components of portfolio flows across borders, i.e. on net bond and equity flows (as well as their variability) in turn. According to Hau and Rey (2006), exchange rate uncertainty should affect equity, but not bond flows. Fidora et al. (2007) and Borensztein and Loungani (2011), by contrast, found evidence that bond flows exhibit stronger home bias compared to equity flows. We provide some relevant empirical evidence on this issue. Third, existing empirical studies on the relationship between exchange rate changes and portfolio flows investigate short-run dynamic interactions only with linear dependence techniques (i.e., first moment analysis). For example, Brooks et al. (2004) and Hau and Rey (2006) use simple correlations and regression analysis for the US vis-à-vis the euro area and Japan, and 17 OECD countries, respectively; Siourounis (2004), Chaban (2009), and Kodongo and Ojah (2012) estimate VAR models respectively for four developed countries (the UK, Japan, Germany, and Switzerland), three commodity-exporting countries (Canada, Australia, and New Zealand), and four African countries (Egypt, Morocco, Nigeria, and South Africa) vis-à-vis the US. Their results are characterised by significant deviations from normality and conditional heteroscedasticity, i.e., volatility clustering or so- called ARCH effects (see Engle, 1982) that are not captured by their setup. By contrast, we model the first and second moments simultaneously to analyse the dynamic interactions between exchange rate changes and portfolio flows. In this way, we capture the volatility in the flows and exchange rate changes which is well documented in the economics and finance literature, and address some of the potential pitfalls of earlier studies. Fourth, since volatility is a measure of the information flow (see Ross, 1989), it is of paramount importance to understand how the stochastic information arrivals in the form of simple portfolio investment shifts in bonds and equities are transmitted to the foreign exchange market, and vice versa. Furthermore, knowledge of the response of investors to exchange rate uncertainty provides important information to policy-makers and regulators to formulate appropriate policies based on imposing or relaxing credit controls on these flows depending on the state of the economy, with the aim of achieving economic and financial stability. For example, if exchange rate uncertainty dampens net inflows, expansionary policies to boost the economy during recessionary periods could be unsuccessful if exchange rates are too volatile. In such circumstances, credit controls on inflows may be relaxed, and such a policy is likely to increase inflows, thereby boosting the economy. In addition, financial and economic stability can be pursued by reducing exchange rate volatility. The remainder of the paper is organised as follows. Section 2 describes the data and carries out some preliminary analysis. Section 3 outlines the econometric model and the hypotheses tested. Section 4 discusses the empirical results, and finally Section 5 concludes.","We examine the impact of exchange rate uncertainty on different components of net portfolio flows, namely net equity and net bond flows, as well as their dynamic linkages for the US vis-à-vis Australia, Canada, the euro area, Japan, Sweden, and the UK. Throughout, the US is considered the domestic or home economy. The data on portfolio investment flows, obtained from the US Treasury International Capital (TIC) System,4 are monthly and cover the period from 1988:01 to 2011:12 for all series. The reason for selecting this start date is that cross-border portfolio flows for the period preceding 1988 are known to be negligible (see Brooks et al., 2004). Net equity (bond) flows are calculated as equity (bond) inflows minus outflows. While inflows are measured as net purchases and sales of domestic (US) assets (equities and bonds) by foreign residents, outflows are the net purchases and sales of foreign assets (equities and bonds) by domestic residents (US). In the case of the euro area, we aggregate the data for the individual EMU countries (Austria, Belgium–Luxembourg,5 Finland, France, Germany, Ireland, Italy, the Netherlands, Portugal, and Spain) to extract cross-border bond and equity flows between the US and this region. Positive numbers imply net equity and bond inflows (in millions of US dollars) towards the US or outflows from its counterparties. Moreover, since without scaling model convergence is difficult to achieve, we use monthly averages to adjust these flows, specifically the average of their absolute values over the previous 12 months as in Brennan and Cao (1997), Hau and Rey (2006), and Chaban (2009) among others. Following more recent papers in the literature (e.g., Chaban, 2009; Kodongo and Ojah, 2012), the exchange rates are end-of-period data, defined as US dollars per unit of foreign currency6; the source is the IMF's International Financial Statistics (IFS). Exchange rate changes are calculated as Et = 100 × (PE,t/PE,t−1) where PE,t stands for the log of the exchange rate at time t. For the period preceding the introduction of the euro, i.e. before 1999, we use US dollar per ECU as the euro area's exchange rate. Descriptive statistics are displayed in Table 1. The mean of monthly exchange rate changes is positive (a US dollar depreciation) for Japan and Canada, and negative (a US dollar appreciation) for the rest of the countries. On the other hand, the monthly mean of net equity flows is positive for Sweden and Canada and negative for the remaining countries, indicating equity inflows from Sweden and Canada towards the US and outflows from the US towards the other countries. The monthly mean of net bond flows is negative for Australia and positive for the other countries. This indicates the existence of bond inflows from all countries except Australia (for which there is evidence of bond outflows) vis-à-vis the US. Exchange rate changes are found to exhibit higher volatility than the two types of flows. Furthermore, equity flows appear to be characterised by higher volatility than bond flows (although their volume is very small). As for the third and fourth moments, exchange rate changes, net equity flows, and net bond flows all exhibit skewness and excess kurtosis in most cases. The Jarque–Bera (JB) test statistics imply a rejection at the 1% level of the null hypothesis that exchange rate changes and the two types of flows are normally distributed in all countries in question. Table 2 reports the LM ARCH test statistics for the residuals of the fitted bivariate VAR model εi,t, where i = 1 corresponds to exchange rate changes, and i = 2 to net flows (equity and bond flows in turn, since the analysis is bivariate). At the 10% significance level, the results indicate the presence of ARCH effects up to order 6 in all variables, except for net equity flows for Australia and net bond flows for the euro area.","The objective of our analysis is to establish whether exchange rate uncertainty affects net equity and bond flows across borders, and also whether there is volatility transmission (hence information flows) between these flows and exchange rate changes and, if so, in what direction causality runs.11 The QML estimates of the bivariate VAR GARCH (1, 1)–BEKK-in-mean parameters as well as the associated multivariate Q-statistics (Hosking, 1981) are displayed in Tables A1–A6 (see Appendix A) for Australia, Canada, the euro area, Japan, Sweden, and the UK, respectively.12 Panels A and B in each Table concern the bivariate regression of exchange rate changes against equity and bond flows respectively. The Hosking multivariate Q-statistics for (6) and (12) lag orders for the standardised residuals in the exchange rate changes-equity flows cases indicate no serial correlation at the 5% level, when the corresponding conditional mean equations are specified with p = 1 for Japan, p = 2 for Sweden and p = 3 for the other countries (the insignificant parameters in the mean equations have been dropped13). With regard to the exchange rate changes-bond flows relationships, whilst no dynamic terms appear to be necessary for Sweden, setting p = 1 for the UK, p = 2 for the euro area, p = 3 for Australia and Canada and p = 5 for Japan is required to capture adequately the dynamic structure in these cases. Also, the Hosking multivariate Q-statistics for (6) and (12) lag orders for the squared standardised residuals suggest that the multivariate GARCH (1, 1) structure is sufficient to capture the volatility in the series. Hence, the estimated models are shown to be well specified. Table 3 reports a summary of the estimated results displayed in Tables A1–A6 (Appendix A). These suggest that there are limited dynamic linkages between the first moments compared to the second ones. The results in the conditional mean equations indicate the existence of bidirectional mean spillovers between exchange rate changes and net bond flows in Japan, as well as spillovers from net bond flows to exchange rate changes in Canada and the UK, and from net equity flows to exchange rate changes in the euro area. The results also suggest that exchange rate uncertainty affects net equity flows negatively in the euro area, Sweden, and the UK, and positively in Australia, and has no effect in Canada and Japan. Its impact on net bond flows, on the other hand, appears to be negative in all countries except Canada for which it is positive. In Fig. 1 we plot net equity and net bond flows, exchange rate changes, and the conditional variances of exchange rate changes for all countries over the sample period. It can be seen that the conditional variances of exchange rate changes, measured using the bivariate VAR GARCH-in-mean parameterisation, were high during the recent global financial crisis in most countries.14 The impact of the 1992 crisis is also apparent in the cases of the UK and the euro area. Interestingly, net inflows towards the US were low during the recent crisis period in most cases, and also during the 1992 crisis, especially for the UK. Overall, these plots support the econometric results implying that periods of high exchange rate uncertainty were associated with declines in net flows. This holds in most cases, except Australia and Japan for net equity flows and Canada for both net equity and net bond flows. The estimated negative impact of exchange rate uncertainty on net equity as well as net bond flows has important implications. First, it indicates that risk-averse market participants, especially those of the counterparties to the US, respond to exchange rate uncertainty by reducing their financial activities, and favouring domestic rather than foreign securities in their portfolios to minimise their exposure to uncertainty. This finding is broadly consistent with the evidence in Bayoumi (1990), Iwamoto and van Wincoop (2000), and Bacchetta and van Wincoop (1998, 2000). While Bayoumi (1990) showed that net capital flows as a share of GDP are lower during the floating exchange rate period (1965–1986) than during the gold standard (1880–1913), Iwamoto and van Wincoop (2000) reported that net capital flows as a fraction of GDP are much larger across regions of a country, which use the same currency, than across countries. Bacchetta and van Wincoop (1998, 2000), on the other hand, showed that exchange rate uncertainty should dampen net international capital flows in the context of a two-period general equilibrium model. Second, in contrast to Hau and Rey (2006), who assumed that bonds are hedged instruments not affected by exchange rate uncertainty, it appears that uncertainty in fact affects bond as well as equity flows, and the former more widely, since a negative impact is found in five of the six countries examined (see also Fig. 1). This is consistent with the results of Fidora et al. (2007), who found that exchange rate volatility is an important factor for bilateral portfolio home bias, this being higher for bonds than for equities. This finding has recently been confirmed by Bekaert and Wang (2009) and Borensztein and Loungani (2011), although in the former study it is not found to be economically significant. The rationalisation of Fidora et al. (2007) of the higher home bias for bonds compared to equities is that it is consistent with Markowitz-type international CAPM specifications in which less volatile financial assets should be characterised by a larger home bias. However, the results indicate that exchange rate uncertainty does not induce home bias in Australia and Japan for equity flows and in Canada for both equity and bond flows (see also Fig. 1). The finding that exchange rate uncertainty has a positive effect on net equity flows in Australia is consistent with the evidence in Batten and Vo (2010) and Daly and Vo (2013), whilst Mishra (2011) found a negative effect. A possible explanation for the findings of Australia and Canada may be that they are commodity-exporting countries and developments in their financial markets are driven by terms-of-trade shocks. Chaban (2009) and Ferreira Filipe (2012) indeed found that the portfolio-rebalancing motive of Hau and Rey (2006) in these countries is weak. Chaban (2009) argued that commodity prices play a significant role in the transmission of shocks in these countries, and Ferreira Filipe (2012) found that differences in the volatility of country-specific shocks also do so. Japan is a special case: as highlighted by Hau and Rey (2006), bond flows represent most of the international portfolio flows for this country, even though a high percentage of Japanese debt is financed internally. The estimates of the conditional variance equations indicate that the conditional variances exhibit persistence in all cases except for net equity flows in Canada (see Tables A1–A6 in Appendix A). While the persistence of the conditional variance of exchange rate changes ranges from 0.54 (Japan) to 0.98 (euro area), that of the corresponding flows ranges from 0.38 (Sweden) to 0.91 (euro area) for net equity flows and from 0.43 (Japan) to 0.98 (Canada) for net bond flows. The ARCH, a11, and GARCH, b11, parameter estimates for exchange rate changes in the bivariate GARCH–BEKK models are rather similar, regardless of whether the relationship with net bond or equity flows is considered (see Panels A and B respectively in all Tables). More specifically, a11 changes by a magnitude of less than 0.10 and this also applies to b11, except for Japan where it is around 0.26. Furthermore, the off-diagonal elements of the ARCH and GARCH matrices suggest the following: statistically significant volatility spillovers running from net equity flows to exchange rate changes (measured by b21) in the cases of Canada and the UK; both shock (measured by a21) and volatility (measured by b21) spillovers from net equity flows to exchange rate changes in the case of Sweden, and bidirectional volatility spillovers (measured by b21 and b12) for Japan. The results also show that net bond flows' shocks affect the volatility of exchange rate changes (a21) in the case of Australia; shock and volatility spillovers running from exchange rate changes to net bond flows are present for Canada, whereas for Sweden they run in the opposite direction; in the case of Japan there are volatility spillovers from exchange rate changes to net bond flows. In the euro area and the UK the volatility spillovers between net bond flows and exchange rate changes are bidirectional. Finally, exchange rates' shocks are found to affect the volatility of net bond flows in the UK, whereas in the euro area shock spillovers run in the opposite direction. Broadly speaking, these findings are consistent with those of the causality-in- variance (i.e., the information flow) test results (Table 3). More specifically, the Wald test statistics (see Tables A1–A6) provide evidence of strong causality-in-variance from net equity flows to exchange rate changes in the case of the UK and Sweden, and bidirectional causality-in-variance in the case of Japan. There is also causality-in- variance from net bond flows to exchange rate changes in Australia and Sweden, and causality in the reverse direction in Canada, as well as bidirectional causality in the rest of the countries. A possible explanation for the existence of stronger dynamic linkages between exchange rate changes and bond flows instead of equity flows is that foreign exchange dealers usually follow bond yields in their trading behaviour; these yields, in turn, drive cross-border bond acquisitions, which results in volatile exchange rates. Spillovers from the exchange rates may be due to the fact that investors adjust their portfolios on the basis of their volatility. Finally, Fig. 2 displays the evolution of the dynamic conditional correlation between exchange rate changes and net flows to provide further insights into the dependence between these variables. The graphical analysis indicates that these correlations are time-varying in most cases. Furthermore, there are clear shifts during turbulent periods such as the 1992 crisis in the case of the UK, and the recent global financial crisis of 2007–2009 in most cases. These plots confirm the existence of strong linkages between exchange rate changes and net flows during such periods, as found before.","In this paper, we have analysed the impact of exchange rate uncertainty on net bond and net equity flows, as well as the dynamic linkages between exchange rate volatility and the variability of these flows, using monthly data for the US vis-à-vis six advanced economies, namely Australia, Canada, the euro area, Japan, Sweden, and the UK over the period 1988:01-2011:12. By estimating bivariate VAR GARCH-BEKK-in-mean models, we find evidence that exchange rate uncertainty impacts on net equity flows negatively in the euro area, Sweden, and the UK and positively in Australia. Furthermore, in contrast to the assumption of Hau and Rey (2006), it also affects net bond flows negatively in all countries except Canada, where the effect is positive. The general conclusion that can be drawn from these results is that exchange rate uncertainty induces risk-averse investors, especially those of the counterparties to the US, to reduce their financial activities and to favour domestic rather than foreign assets in their portfolios in order to minimise their exposure to uncertainty. This evidence is stronger in the case of the UK, the euro area and Sweden compared to Canada, Australia and Japan. The results for Australia, Canada and Japan may be due to the specific characteristics of these economies, as documented in other studies (e.g., Hau and Rey, 2006; Chaban, 2009; Ferreira Filipe, 2012). The causality-in-variance analysis suggests the existence of strong spillovers from net equity flows to exchange rate changes in the UK and Sweden, and bidirectional causality-in- variance in Japan. Causality-in-variance is also found to run from net bond flows to exchange rate changes in Australia and Sweden, in the opposite direction in Canada, and in both directions in the other countries. Overall, our findings have important policy implications that are country-specific. In particular, they suggest that policy-makers and economic and financial regulators in countries with strong uncertainty effects can use exchange rate or credit controls on equity as well as bond flows as instruments to achieve economic and financial stability."],["In most equilibrium sorting models (ESMs) of residential choice across neighborhoods, the question of whether households rent or buy their home is either ignored or else tenure status is treated as exogenous. Of course, tenure status is not exogenous and households' tenure choices may have important public policy implications, particularly since higher levels of homeownership have been shown to correlate strongly with various indicators of improved neighborhood quality. Indeed, numerous policies including that of mortgage interest deduction (MID) have been implemented with the express purpose of promoting homeownership. This paper presents an ESM with simultaneous rental and purchase markets in which tenure choice is endogenized and neighborhood quality is partly determined by neighborhood composition. The public policy relevance of the model is shown through a calibration exercise for Boston, Massachusetts, which explores the impacts of various reforms to the MID policy. The simulations confirm some of the arguments made about reforming MID but also demonstrate how the complex patterns of behavioral change induced by policy reform can lead to unanticipated effects. The results suggest that it may be possible to reform MID while maintaining the prevailing rates of homeownership and reducing the federal budget deficit. --------------------------------------------------------------------------------","The promotion of homeownership has been a widespread and long-term focus of public policy (Andrews and Sanchez, 2011). Support for such policies derives both from political ideology and from a belief that homeownership delivers positive spillovers. Homeowners, it is argued, have greater incentives to invest in the physical and social capital of their communities, thus providing private and public benefits. There is a substantial body of empirical evidence that lends credence to this view. Homeownership is strongly correlated with property condition and maintenance (Galster, 1983), neighborhood stability (Dietz and Haurin, 2003; Rohe and Stewart, 1996), child attainment (Bramley and Karley, 2007; Green and White, 1997; Haurin et al., 2002), citizenship (DiPasquale and Kahn, 1999) and lower crime rates (Glaeser and Sacerdote, 1996; Sampson and Raudenbush, 1997).1 A wide variety of policy measures have been implemented to promote homeownership. Attempts have been made to encourage the supply of mortgage lending; for example, in the U.S. through the establishment of Government Sponsored Entities providing liquidity and security for mortgage lenders. Policies have also been implemented to encourage particular groups into homeownership; for example, in the U.K. through the Right to Buy scheme for social housing tenants and more recently the Help to Buy schemes for equity loans, mortgage guarantees and new buyers (NewBuy). Homeownership has also been promoted through the tax system e.g. through exemptions from capital gains tax on property sales and mortgage interest deduction (MID). MID, the focus of this paper, allows taxpayers to subtract interest paid on a residential mortgage from their taxable income. MID is present in the tax laws of many countries including the U.S., Belgium, Ireland, the Netherlands, Switzerland and Sweden and was previously offered in the U.K. and Canada. It was introduced in the U.S. in 1913 when the homeownership rate was 45.9%. Under MID and numerous other initiatives, homeownership rose after the Second World War reaching a peak of 69% in 20042. Currently, MID constitutes the second largest US tax expenditure3 with the cost estimated to be some $104.5 billion dollars in foregone tax revenue in 2011 (Office of Management and Budget, 2011). In the context of a large US fiscal deficit, MID has come under increased scrutiny. It has been argued that rather than encouraging homeownership the tax subsidy is simply capitalized into property values making properties no more and potentially less affordable than without the policy (Glaeser and Shapiro, 2002; Hilber and Turner, 2010). Furthermore, critics contend that MID most greatly benefits high-income taxpayers who would likely be homeowners irrespective of the tax incentives (Shapiro and Glaeser, 2003). Certainly higher income households are more likely to own their homes, hold larger mortgages and itemize mortgage interest payments on their tax returns (Poterba and Sinai, 2008). Of course, courtesy of their higher incomes, they also itemize at a higher rate (Glaeser and Shapiro, 2002). As a result, in 2004 the government paid an average $5459 in MIDs to households earning over $250,000 compared to $91 for households earning below $40,000 (Poterba and Sinai, 2008). In the face of strong opposition, particularly on the part of financial services interests and housing lobbyists, repeated efforts to reform MID in the U.S. have borne little fruit Ventry (2010)4. Over the last three budget cycles the U.S. administration proposed reforms to MID, but on each occasion those initiatives have failed to pass into law. The key element of those proposals was to limit MID for households paying the top marginal rates of income tax. Other proposals for reform include; replacing MID with a system of tax credits (Dreier, 1997; Follain et al., 1993; Green and Vandell, 1999), scrapping MID in order to fund cuts in federal income taxes (Stansel, 2011) and replacing MID with a fiscal incentive open only to first time buyers (Gale et al., 2007). The debate is fueled by a lack of clarity with regard to how such reforms will play out. Clearly, eliminating the MID will increase the cost of borrowing for the purposes of buying property and, ceteris paribus, cause demand for owned properties to fall. This reasoning underpins the National Association of Realtors claim that “eliminating the MID will lower the homeownership rate in the U.S.”5. Of course, it is recognized that the impact of eliminating the MID also depends on supply conditions in the property market. The extent to which falling demand translates into reductions in homeownership as opposed to falling prices depends on the price elasticity of housing supply. Bourassa and Yin (2008) estimate that for some groups the negative effect of losing MID may be more than outweighed by the positive effect of falling property prices; homeownership amongst such groups could actually rise as a result of eliminating the MID. What is less widely recognized is that changing market conditions in the property market will have ramifications in the closely associated rental market. Falling demand for homeownership can translate into rising demand for rental housing. More complex still is the interplay between homeownership and the desirability of residential locations. Since residential location choice is endogenous to the problem, eliminating MID may not only encourage the movement of individuals between ownership and rental but also the migration of households between neighborhoods. While numerous attempts have been made to identify the impacts of eliminating the MID (e.g. Bourassa and Yin, 2008; Hilber and Turner, 2010; Toder, 2010) those studies have been based on a partial characterization of the problem. This paper develops a model that more completely describes the complex adjustments in spatially defined and interrelated property markets and uses that model to explore some of the possible ramifications of MID reform. The model developed in this paper is an equilibrium sorting model (ESM) (Kuminoff et al., 2010). ESMs provide a framework within which it is possible to examine how households choose their residential location from a set of discrete neighborhoods. As reviewed in Section 2, ESMs have been developed to examine a number of economic issues relating to choice of residential location. As far as we are aware, however, our model is the first ESM to simultaneously model purchase and rental markets while endogenizing tenure choice. In Section 3 the innovations of the model are outlined in detail; particularly the specification of a neighborhood level of public good provision whose value depends, in part, on endogenous levels of homeownership and the development of an adjustment process to policy reform that accommodates capital gains. To elucidate the pathways of adjustment that MID reform may initiate in property markets, Section 4 presents a simple two-jurisdiction calibration of the model based on the 2000 census data for Boston, Massachusetts. The calibrated model is used to simulate four different MID reform proposals; capping MID at a rate of 28%, replacing MID with refundable tax credits, scrapping MID and reducing income taxes and replacing MID with a lump sum payment to new owners. The simulations allow us to examine several important questions with regard to MID reform. In particular, to explore how reforms may impact property prices, levels of homeownership, the distribution of welfare across income groups and the mixing of income groups within and across jurisdictions. Our analysis suggests that, contrary to existing claims, with the right policy design it may be possible to reform MID while maintaining the prevailing rates of homeownership, increasing public goods provision and contributing to a reduction in the federal deficit.","In essence, equilibrium sorting models (ESMs) provide a stylized representation of the interactions of households, landlords and government within a property market. Originally developed to explain observed patterns of socio-economic stratification and segmentation in urban areas (e.g. Ellickson, 1971; Epple and Romer, 1991; Oates, 1969; Schelling, 1969; Tiebout, 1956), ESMs provide a formal account of the process whereby heterogeneous households sort themselves across the set of neighborhoods within a property market. Neighborhoods, it is assumed, differ in quality according to the level of public goods each provides. Those public goods may reflect purely physical attributes of a location (for example, a neighborhood's proximity to commercial centers) or the levels of provision of local amenities (for example, the quality of local schools). An important distinguishing feature of ESMs is in allowing local amenity provision to be shaped by endogenous peer effects; that is to say, by the characteristics of the set of households that choose to locate in a neighborhood. Epple and Platt (1998), for example, present a model in which local taxes and lump sum payments are determined by the voting preferences of the residents in a neighborhood; these computationally complex models often have no closed form solution and are instead solved using numerical computation. Similarly, Ferreyra (2007) and Nesheim (2002), present models in which school quality is related to measures of the average income of households in a locality. In an ESM, the mapping of households to quality-differentiated neighborhoods is mediated through property prices. Indeed, a solution to an ESM is taken to be a set of property prices that support a Nash equilibrium allocation of households to neighborhoods such that the supply and demand for properties are equated in all neighborhoods. While some simple ESMs have closed form solutions equilibria for more complex models, particularly those including endogenous neighborhood quality, are usually calculated using techniques of numerical simulation (Bayer et al., 2004; Ferreyra, 2007). Over the last decade ESMs have increased in popularity and complexity. Recent modeling extensions allow for moving costs (Bayer et al., 2009; Ferreira, 2010; Kuminoff, 2009), overlapping generations (Epple et al., 2010) and simultaneous decisions in a parallel labor market (Kuminoff, 2010). In addition, the ESM framework has been used to explore empirical data on the distribution of households and property prices in order to derive estimates of the value air pollution (Smith et al., 2004), school quality (Bayer et al., 2004; Fernandez and Rogerson, 1998) and the provision of open space (Walsh, 2007). ESMs have also been used to explore policy issues such as school voucher schemes (Ferreyra, 2007), open space conservation (Klaiber and Phaneuf, 2010; Klaiber, 2009; Walsh, 2007) and hazardous waste site clean ups (Smith and Klaiber, 2009). A comprehensive review can be found in Kuminoff et al. (2010). One area that has received relatively little attention in the ESM literature is that of tenure. Indeed, the vast majority of ESM applications make the assumption that households rent their properties from absentee landlords. Where different tenure statuses have been considered, those applications have treated tenure status as a fixed characteristic rather than a choice variable (Bayer et al., 2004; Epple and Platt, 1998). In reality, of course, households choose from a number of tenure options, with the key distinction being between ownership and renting. The joint decision of tenure and housing consumption has been examined in the real estate literature. For example, King (1980) estimated preferences for the UK housing market developing an econometric model of joint tenure and housing demand. Similarly, Henderson and Ioannides (1986) consider joint housing decisions in the US and Elder and Zumpano (1991) developed a simultaneous equations model of housing and tenure demand. For a number of issues, such as the reform of MID policy, the choice of tenure is the central consideration of the policy debate. Accordingly, one of the key contributions of this paper is to describe an ESM in which tenure choice is endogenized. In our model, household choices whether to rent or purchase property are a function of market conditions, including the endogenous provision of local public goods. When policy reforms result in price changes in the property market, homeowners and renters are affected differently. In particular, homeowners will be impacted by capital gains (or losses) that are not experienced by renters. The modeling framework developed in the next section outlines a method for incorporating such distinctions. The economy ~~~~~~~~~~~ Consider a closed spatial economy consisting of a continuum of households. The model is closed insomuch as households may not migrate in or out of the economy. Households differ in their incomes, y. They also differ in terms of their preferences over the amount of housing they consume, β, and the value they attach to owning a property, θ. Ownership preferences represent the private returns to homeownership that are not realized when renting. Such private returns are motivated by numerous considerations including i) freedom to modify housing, ii) satisfaction from homeownership status and iii) anticipated financial returns from capital gains. The distribution of household types in the population is defined by the joint multivariate density function f(y, β, θ). The economy is divided into a set of spatially discrete neighborhoods, j = 1, …, J. In our model, each neighborhood is assumed to have its own local government. As such, we refer to these areas as jurisdictions. Each jurisdiction is characterized by a vector of local public goods, gj = {zj,1 …, zj,U, qj,1, …, qj,V}, comprising U exogenous elements, zj,u, and V endogenous elements, qj,v. The level of provision of endogenous elements is determined by the composition of the set of households that choose to reside within a jurisdiction. The provision of public goods is assumed to be homogenous within a jurisdiction. The demand side ~~~~~~~~~~~~~~~ To reside in jurisdiction j household i must buy housing there. The decision to rent, R, or own, O, housing is referred to as tenure choice. We describe the set of tenure options as T = {R, O} Accordingly, our model is characterized by households choosing to participate in one of a number of property markets each defined by a jurisdiction and tenure bundle, {j, t}. Households also choose a quantity of housing; a decision approximating real life choices over the size and quality of home to buy or rent.6 Housing is defined as a homogenous good that can be owned or rented from absentee landlords at a constant per unit cost, pj, within a jurisdiction7 (Epple and Romer, 1991; Epple and Sieg, 1998; Epple and Platt, 1998; Bayer et al., 2004; Ferreyra, 2007). The quantity of housing demanded by a household in market {j, t} is denoted hj,t = h(p, g, τp; y, β, θ, m, δ). The two arguments in that function yet to be explained (m and δ) concern the borrowing a household must assume in order to purchase a property. In particular, to become a homeowner a household must take out a mortgage8 and pay mortgage interest, m, to the lender. Mortgage interest is paid only on the amount borrowed, where that borrowing is determined by the value of the housing purchased, pjhj,t multiplied by a loan-to-value ratio, δi. Differences in δi can be interpreted as representing the varying abilities of households to make a down payment. Property taxes τp are paid on both rented and purchased housing. Accordingly, the implicit subsidy a household, i, receives by itemizing mortgage interest and property tax payments on their tax return, MID(p, h, τp, y, m, δ), is endogenous to the household's decision and depends upon the purchase price of property (not including tax), p, the quantity of housing demanded, h, the property tax rate, τp, and household income, y (which in our model also determines the loan-to-value ratio, δ, and the probability that a household itemizes). Local public goods ~~~~~~~~~~~~~~~~~~ Notice that our specification assumes that public good provision is increasing in the homeownership rate: a relationship that might imply homeownership has a direct effect on local public good provision or that homeownership simply proxies for unobserved inputs that themselves have a direct effect. The presence of xj in the public good production function defines a peer effect whereby community characteristics, perhaps median household income, affect the provision of public goods. Such peer effects have considerable empirical support (Nechyba, 2003) and have been incorporated in a number of existing ESM specifications (Nesheim, 2002; Ferreyra, 2007). The household optimization problem ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Households derive utility from local public goods, g, consumption of housing, h, and other consumption, c. Preferences for local public goods are determined by parameter α, which is assumed to be constant across households. Meanwhile, housing quantity preferences are determined by parameter β, which is assumed to vary across households. Our model also allows for the fact that households can derive more utility from housing when they own their home than when they rent it (or vice versa). Each household is characterized by values for the preference parameter set θ, which scales the utility derived from housing in the utility function for home owning. Finally, households select the jurisdiction and tenure combination that maximizes their level of utility. Simulating responses to exogenous policy change ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In reality policy changes occur in a world in which households already rent or own existing properties. That reality influences the outcome of a policy change in at least two ways. First, changes take place in the context of an existing housing stock whose quantity and location has been determined by households' initial choices. Second, a household's current tenure status determines whether their choices following the policy change are influenced by capital gains.10 To see that more clearly, it is instructive to briefly contemplate how market changes impact differently on renters and owners. Consider a change that leads to increased house prices. When prices go up existing renters are unable to afford their current consumption bundle. Households can respond in a number of different ways. They can alter their tenure choice, they can move to another location where property prices are lower or they can reduce their demand for housing and consumption. Indeed, they could do any combination of these. In contrast, owning a property prevents rises in prices from making the current consumption bundle unaffordable; homeowners are shielded against price increases. Instead, a rise in prices presents homeowners with the opportunity to sell-up and use the capital gains to increase consumption or relocate to a jurisdiction that provides more desirable public goods. To simulate the process of adjustment in the property market within the context of what is essentially a static model requires some careful consideration. We first assume that the market is in a state of long-term equilibrium, an equilibrium achieved under the baseline policy. Households have optimally chosen where to live, whether to rent or own and how much housing to consume. To reflect that state of the world, we imagine a property market in which all the housing units demanded under that baseline policy have been constructed and that these existing housing units cannot be demolished in the face of a policy change (though they can be repackaged and new units may be constructed). The policy change is introduced to this world at a point after homeowners have paid for their current properties at the pre-change prices but before rent has changed hands, consumption goods have been bought and taxes and mortgage interest have been paid. As a result of the policy change, households reconsider their choices of housing quantity, location and tenure status and the model is solved for the set of property prices that bring the market back to equilibrium under the changed conditions. For, households that were previously renting, things are relatively simple: they either buy or rent at the prices determined by the new equilibrium. In contrast, having bought at the prices characterizing the old equilibrium, households that previously owned must make their new housing decisions in light of the fact that the new price conditions may present them with capital gains or losses.","The model developed above provides a rich environment in which to explore the general equilibrium consequences of reforming MID policy. Within that environment the impact on government expenditure, patterns of community composition, homeownership rates and the levels and distribution of household welfare can be considered simultaneously. To undertake this exercise it is preferable to examine a model that replicates the real world. Such a model requires reasonable but tractable functional forms that can be calibrated to produce a model resembling a real world property market. Following the convention of Epple and Platt (1998) we specifically model Boston, using updated data for 2000. To provide a clear and accessible illustration of the pathways of change that operate in light of a policy reform it is prudent to consider a simple two-jurisdiction version of the model. While it is eminently possible to investigate problems with many more jurisdictions, this simplification enables us to most clearly characterize the chain of reactions that occur within property markets in response to policies that reform MID. The model is coded in Matlab11 and uses simulation and iterative numerical techniques to solve for market clearing prices and provision of endogenous public goods (Lagarias et al., 1998)12. The proposed policy reforms ~~~~~~~~~~~~~~~~~~~~~~~~~~~ The current debate regarding reform of MID policy is motivated in part by the large U.S. deficit. Indeed, as part of plans to reduce that deficit, President Obama submitted federal budget proposals in 2011, 2012 and 2013 that advised capping itemized deductions, including MID, at 28%. Each time Congress has rejected the recommended tax reforms. All the same, we take the proposal of capping MID at 28% as our first potential policy reform. To be clear, under the current tax system homeowners are permitted to deduct mortgage interest and property tax payments from their taxable income when calculating their income tax bill. Those in the top three income tax brackets, therefore, are entitled to an implicit rebate on those expenditures at their marginal tax rates of 31%, 36% and 39.6% respectively. The cap limits that implicit rebate to 28%. Three alternative MID-reform policies are also considered: a refundable flat-rate tax credit; an income tax reduction; and a New Owner Scheme. To compare the various proposed policies, we make the assumption that the central motivation for reform to the MID is reduction of the budget deficit. Accordingly, we calculate the reduction in deficit brought about by our baseline reform of a 28% cap on MID. The three alternative MID-reform policies are tailored to ensure that they facilitate the same reduction in the budget deficit as the cap.13 Let us briefly review the alternative MID-reform policies. First, replacing MID with a refundable14 flat- rate tax credit has been advocated by both the Center for American Progress, who propose a 15% refundable tax credit, and the National Commission on Fiscal Responsibility (Moment of Truth, 2010), who propose a 12% non-refundable mortgage interest tax credit. For the purposes of our simulations, we model this reform as being a policy change in which MID is abandoned and, instead, all households who are owners can claim back a flat-rate percentage of their mortgage interest and property tax costs. As explained previously, the flat-rate we apply in our subsequent simulations is chosen such that federal budget savings achieved by this policy are identical to capping MID at 28%. Our second alternative MID-reform policy follows the proposal made by the Reason Foundation (Stansel, 2011) to scrap MID and instead introduce a revenue neutral reduction in federal income tax for all households. Here, we consider a policy in which MID is abandoned and a portion of the savings in government expenditure are used to fund an equal percentage reduction in income tax for all households. Again, to ensure comparability across our policy simulations the level of income tax reduction is chosen such that the policy achieves the same reduction in the federal government budget deficit as the other proposed reforms. Our final alternative MID reform policy takes motivation from the First Time Buyers scheme proposal made by Gale & Gruber (2007), which suggests scrapping MID and introducing a refundable payment to first-time buyers in the first year after a property is purchased. In the model this is achieved through a New Owner Scheme, which makes an equal lump sum payment to new homeowners. Again in our simulations the level of payments to these first time buyers is chosen so as to ensure comparability in the reduction of the federal budget deficit across reforms. Calibration ~~~~~~~~~~~ To conduct the simulations, specific functional forms are selected for the structural equations of the model. Following Epple and Platt (1998), parameter values for the functions were calibrated such that our model approximates the reality of the Boston Metropolitan (PSMA) area; though in our application we take data for Boston from 2000 and not 1980. Table 1 presents a summary of important statistics for Boston in 2000 and Table 2 summarizes the parameters obtained by calibrating the model to that reality. The assumptions and methods used in deriving those parameters are explained in the following sections. Jurisdictions To allow the pathways of response to MID reform to be studied with reasonable clarity, we explore a simple two-jurisdiction version of the model. Extensions to multiple-jurisdiction models are relatively easy to implement, but greatly reduce the tractability of interpretation. Again following Epple and Platt (1998), we begin by dividing the Boston Metropolitan area into two jurisdictions, labeled A and B. Jurisdiction A provides a higher level of exogenous local public goods provision, consequently attracting a larger share of the population. Competition for access to public goods increases the price of housing in A relative to B, which leads to some income segregation. As a result the median household income in jurisdiction A is higher than that of B. Households Households in the model are characterized by three parameters; income, y, housing preference, β, and ownership preferences, θ. The first step in calibrating this model, therefore, is to establish the joint distribution of those parameters amongst the residents of Boston in 2000. As made explicit shortly, a Cobb–Douglas utility function is assumed such that a household's housing preference, βi, is related to the proportion of their income that they spend on housing. That data along with information on household income, yi, is available from the census which provides a cross-tabulation of the share of income spent on rent (and equivalently on monthly owner costs for owner–occupiers) by income. Accordingly, to establish the joint distribution of y and β we use maximum likelihood estimation to fit a bivariate-normal distribution flny, β) ∼ N(μf, Σf) to 2000 census data for Boston.15 Parameter values from that estimation are recorded in the first row of Table 2. Notice that β is negatively correlated with y, such that while high- income households spend absolutely more on housing than low-income households, that expenditure constitutes a smaller proportion of their income. For simplicity, and due to a lack of existing empirical evidence, it is assumed that household ownership preferences, θ, are independent of income and housing preferences. Accordingly, values were drawn from a lognormal distribution ln(θ) ∼ N(μθ, σθ2) with mean and variance chosen such that the baseline model predicts homeownership rates comparable to those observed in Boston in 2000.16 The parameters selected through that procedure are also recorded in Table 2. For the purposes of simulating our model, we create a sample of 2000 households with income, y, and preference parameters, β and θ, drawn from the calibrated distributions.17 Taxes The property tax rate, τp, was set at the average level for Boston in 2000 using data supplied by the Massachusetts State Government. To capture the correlation between income and itemization rates, the probability of a household itemizing was calibrated using data on itemization rates by income from Poterba and Sinai (2008) which is reproduced in Table 2.18 In this calibration we model property taxes as if they are fully passed on to renters. This is aligned with a simplified interpretation of the economy in which housing is constructed and supplied at marginal cost including property taxes. In this situation, a competitive market would lead to the full incidence of the tax being levied upon renters. This assumption could be relaxed through further refinement of the property market model and the development of a buy-to-let market to expressly model differences in the tax burdens of owner–occupiers, landlords and renters.19 Housing supply The assumption of a single housing supply function covering rental and purchase properties is motivated by considering the direct sale or rent of housing from a zero-profit housing constructor. If property is supplied in a single market, constructors must be indifferent between renting and selling properties. If they sell a property at marginal cost, they are no longer responsible for maintenance, depreciation and foregone interest. If the property is rented the constructor must include these costs in the rental price in order to break even. As a result, if housing for rent and purchase is produced for sale in a single market, both renters and purchasers must face the same per unit user cost of housing, although technically rents exceed purchase prices since homeowners pay the remaining user costs (e.g. maintenance costs) directly rather than to the constructor.21 Local public goods For simplicity, the calibrated model considers one exogenously determined local public good and one endogenously determined local public good. The extension to multiple local public goods is trivial, but adds complexity to the interpretation of the simulation results. We use air quality to act as a representative exogenous local public good. In our simulation, air quality is defined in units of concentration of nitrogen oxides (measured in pphm) below the highest level observed in Boston in the Massachusetts Air Quality Report. Using that measure, the mean level for air quality in Boston in 2000 was 3. Accordingly we set air quality in jurisdiction B to that level but assume that jurisdiction A offers a slightly higher level of provision: 4. To test the sensitivity of our simulated equilibria to the assumption that homeownership directly impacts on local school quality, we explore two versions of the model. In the first version, a direct affect is assumed away. Rather homeownership is taken to proxy for a set of omitted factors that impact directly on school quality through channels that are independent of property market decisions. Since those omitted factors are taken to be unchanging, we progress by exogenously fixing the homeownership argument in the school quality production function at the observed Boston state average. In our simulations, the argument maintains that initial level despite adjustments in rates of homeownership that result from changes in MID policy. In the second version of the model the assumption of a direct affect is maintained. In our simulations, the homeownership argument in the school production function updates in response to changes in property market decisions brought about by reforms of MID policy. The results reported in Section 4 are derived from the first model version, assuming local homeownership has no direct impact on school quality (although local endogeneity is still present through expenditure per pupil and median incomes). Comparable results from the second model version, with an endogenous homeownership feedback, are presented in Section 5.2. Additional results exploring the sensitivity of the results to sample size, the strength of preferences for public goods and alternative housing supply specifications are available from the authors upon request. As a general comment, the patterns of relocation suggested by the model are the same for both versions under each reform. It is notable, however, that when local homeownership rates are assumed to directly affect the provision of local public goods the magnitude of welfare gains associated with reforms that increase homeownership rates are sensitive to this assumption. Results ~~~~~~~ The long-run equilibrium under current policy conditions was calculated for our simulated sample of 2000 households.25 The impact of MID-reform was then investigated using an iterative solution algorithm to calculate the new equilibrium characterizing the property market when each of our four proposed policy reforms was instituted from that baseline. Tables 4 and 5 describe important features of the equilibrium in the baseline and for each policy-reform scenario. Table 4 presents a characterization of those equilibria in terms of the composition and characteristics of the households in each jurisdiction. Table 5 characterizes the equilibria from the perspective of households in each of the six tax brackets. Throughout our discussion of the results we will use the term “price” to refer to the price inclusive of property tax since this is the effective price faced by households. Baseline with MID Consider the results displayed in the first rows of Tables 4 and 5 that describe the equilibrium that evolves under the current system of MID. In the baseline, jurisdictions A and B differ initially only in their exogenous provision of public goods (Column 1 of Table 4). The higher exogenous public good provision in jurisdiction A shapes the resulting equilibrium. Households prefer a greater provision of public goods which increases demand for housing in A relative to B. Consequently, the population of A is higher than the population of B, with 65% of all households residing there (Column 2 of Table 4). As the supply of housing in A is not infinitely elastic, relatively stronger demand in A drives the prices of housing in A above the prices in B. The purchase price of housing (including property tax) is $7379 in A compared to $5496 in B (Column 3 of Table 4). Price differences between jurisdictions and tenure options precipitate the stratification of households. Column 4 of Table 4 confirms that households with relatively strong ownership preference, θ1, choose to purchase housing while those with relatively weak ownership preference rent housing. Similarly, as can be seen from column 5 of Table 4, households that spend a relatively large proportion of their income on housing, e.g. those with relatively high β, prefer lower housing prices and tend to choose to reside in jurisdiction B. Since β is negatively correlated with income, this reinforces segregation by income. As shown in Columns 1 and 2 of Table 5, only 19% of households in the lowest income tax bracket (1st) choose to live in A compared to 96% in the highest tax bracket (6th). Consequently, the median income of households in A is almost 3 times that of B (Column 7 of Table 4). Within each jurisdiction some households rent while others own. Recall from the calibration that households with higher incomes face relatively lower loan-to-value ratios and, under the existing MID policy, can itemize their mortgage interest and property tax costs at a relatively higher marginal rate. Accordingly, the marginal cost of purchasing housing is lower for higher-income households and, ceteris paribus, households with high incomes are more likely to become homeowners. As shown in Columns 7 of Table 5, only 52% of households with incomes below the standard deduction choose to own compared to 74% of households in the highest income tax bracket. This result is consistent with observed homeownership rates in Boston in 2000. Returning to Table 4, the concentration of higher income households in A leads the homeownership rate to be higher than in B (Column 8). Recall from Eq. (5) that local property tax revenues depend on both purchase prices and the total quantity of housing demanded in a jurisdiction. In the baseline equilibrium, higher property prices in A are slightly offset by larger property sizes in B such that tax revenue per household in A is marginally lower than in B; $22,266 and $22,344 respectively (Column 10 of Table 4). Larger local tax revenues translate directly into higher levels of local government expenditure on the endogenous public good. However, since median income is higher in A than B, jurisdiction A benefits from relatively larger provision of the public good through a stronger peer effect (Column 11 of Table 4). Overall, provision of the endogenous public good is higher in A, with a school quality score of 498, than it is in B, at 379. Combined with the exogenous public good provision this indicates that the index of public goods provision is 32% higher in jurisdiction A. That difference in provision of the public good acts to further exaggerate the patterns of sorting initiated by the initial difference in public goods provision. 28% cap Now let us consider how things change when MID is capped at a rate of 28%. Under the new policy those households in the top three tax brackets who had previously been able to itemize their expenditures on mortgage interest and property tax at 31, 36 and 39.6% respectively, would now be limited to itemizing at 28%. In the absence of other adjustments, the cap raises the per-unit cost of housing for the 86% of households in the top three income tax brackets that itemize MID on their tax returns in the baseline. Characteristics of the new equilibrium are presented in the second row of results in Tables 4 and 5. The equilibrium outcome is the product of a number of forces: the immediate impact of the reform is to reduce demand amongst existing homeowners with a relatively strong housing preference. This reduction threatens to lower local expenditure on public goods through a reduction in property tax revenues, which precipitates the relocation of some renters from jurisdiction A to B, increasing housing demand in B and pushing up property prices there. In turn, this redistribution of the population leads to higher median incomes in both A and B (column 7, Table 4), and increases local expenditure on public goods in both jurisdictions. Those affects combine in precipitating a rise in the endogenous public good (proxied here by school quality) of 1% in A (to a score of 499) and 0.3% in B (to a score of 380). The slightly larger increase in public good provision in A makes it more desirable relative to B, this stimulates a rise in prices in order to avoid relocation of households from B to A and maintain the balance of supply and demand. Overall, property prices rise by roughly 1.04% in A and 1.23% in B despite the removal of MID. Comparing results in column 9 of Table 4 we can see that mean rental property sizes fall; this is a consequence of the relocation of households with relatively small housing consumption from A to B and the rise in prices26. However, as anticipated we also see a contraction in mean housing consumption across tax brackets four to six (see Table 5), however homeownership rates in A and B are left almost entirely unchanged (identical to 2 s.f.). It is worth taking a moment here to reflect on the adjustment in prices. Intuitively, one might assume that the cap on MID increases the marginal cost of housing for individuals in the top tax brackets, which in turn leads to a contraction in their demand. It follows that from this partial equilibrium perspective one would expect property prices to fall. The equilibrium sorting model, however, shows that that logic is incomplete and results in an erroneous conclusion regarding the price impact of the policy. There are two key factors at play. First, households can move between jurisdictions. Accordingly, while a reduction in demand from the residents of a desirable area has the immediate effect of pushing down prices, those same price falls will encourage households from other jurisdictions to move into the area driving up demand and, as a result, prices. Second, the endogenous public good responds to changes in neighborhood composition and tax revenues. As a result, a policy that initially stimulates relocation can alter the relative provision of public goods across neighborhoods, a change that in turn can drive secondary demand and price adjustments. As demonstrated by our simulations, in an area where public good provision rises, housing becomes more desirable and those demand increases act so as to drive prices upwards. Perhaps unexpectedly, despite policy reform constituting a significant (20%) reduction in federal government spending, within our simulation the knock-on effects of a policy capping MID at 28% actually precipitates general welfare increases for households in our simulated population. The endogenous tenure choice aspect of our model allows us to explore the impact of the policy change on patterns of renting and owning. Our model suggests that the key driver of the welfare increase provided by the 28% cap is the fact that a policy which increases mortgage costs for high-income households has very little impact on their demand for housing or homeownership. Consider the final column of Table 5, the rise in public good provision leads to utility gains for almost half of the households in the lowest tax bracket and the majority of households in the second to fourth income brackets. Moreover, despite being most directly and adversely affected by the cap, 78% of households in the 5th and 6th income tax brackets experience gains in utility, primarily through the increased local public good provision. Table 6 presents a summary of the monetized changes of the policy in terms of the federal budget deficit, mortgage interest payments, landlords' rents and households' willingness to pay. Willingness to pay is calculated as the sum of the amounts of money that each household would pay, or would require in compensation, in order to leave them indifferent between their original position under MID and that after the policy change. In addition to providing a 20% reduction in the deficit, the 28% cap leads to a rise in mortgage interest payments and substantial gains in household welfare. 20.3% refundable flat-rate tax credit A seemingly more progressive reform of MID would be to replace the current system with a refundable flat-rate tax credit. Under this policy, rather than being able to claim MID against income tax, the federal government reimburses all homeowners a fixed percentage of their mortgage interest payments. To maintain comparability with the MID cap reform discussed in the last section, we consider a refundable tax credit of 20.3% which leads to the same overall deficit reduction as the cap. In contrast to capping MID, the introduction of a tax credit has immediate implications for all households. In the absence of any other adjustments, the marginal cost of purchasing housing reduces for households in the lowest two tax brackets and all non-itemizers. For itemizers in the top four tax brackets the marginal cost rises. For the top tax brackets, the MID cut is more severe than under the cap (down to 20.3% compared to 28%). Accordingly, as in the case of the cap, the reduction in MID leads to a contraction in housing demand amongst previous owners in the top tax brackets. At the same time, demand from households in the 2nd and 3rd tax brackets expands. Most interestingly, in jurisdiction B expanding demand from households in the 2nd and 3rd tax brackets forces lower income households out of homeownership, leading to a reduction in the homeownership rate to 66% and a rise in the number of house units per homeowner in B (see Table 5). This effect is partially a product of the minimum house sizes that must be purchased to reap the gains from homeownership under the chosen functional form, in addition to the higher cost of owning relative to renting per unit of housing. Lower income households taking advantage of the tax credit also stimulates demand for homeownership in A, offsetting the contraction in demand from higher income households. The homeownership rate remains stable at 75% and the average number of housing units per homeowner falling only slightly from 1.95 to 1.86. Moderate rises in tax revenues combined with an increase in the median income in jurisdiction A lead to small gains in public good provision with school quality rising by 1% in both A and B. This increased public goods provision provides some compensation for households in light of the higher cost of housing. The flat-rate tax credit stimulates changes in the tenure and location choice of almost 6% of the population. The progressive nature of the policy makes it unsurprising that the majority of the benefits are focused upon the lowest three tax brackets. What is surprising, however, is that a smaller proportion of households in the second and third income tax brackets benefit from the tax credit in comparison to the 28% cap. In addition, more substantial increases in the cost of housing result in losses for 66% of households in the top three income tax brackets. As can be seen in Table 6, the flat-rate tax credit produces large reductions in welfare through higher costs of housing for renters and previous itemizers in the top four tax brackets; however the reform also leads to a reduction in total mortgage interest payments as overall homeownership rates decline. In comparison to the 28% cap, the flat-rate tax credit provides utility gains to a smaller proportion of households in every tax bracket except the 2nd and from a Kaldor–Hicks perspective the flat-rate policy is inferior to the 28% cap. 4.2% income tax rebate Consider next a policy that removes MID and uses the money saved first to reduce federal expenditure (by 20% to maintain comparability with the 28% cap) and second to reduce income taxes by cutting all household's tax bills by 4.2%. For non- itemizers and renters, the design of this reform is potentially positive. Their lower income tax liability opens up the possibility of consuming larger properties or relocating to A to enjoy relatively higher levels of public good provision. For homeowners, the immediate impact of the reform depends on their taxable income. Households in the lowest tax bracket do not pay income tax and, as such, are not immediately affected by the policy reform. The characteristics of the new equilibrium are presented in Tables 4 and 5. Overall, mean owned property size in B falls slightly as new owners purchase smaller properties than existing owners. The reform motivates a number of households to relocate from A to B in order to benefit from the relatively cheaper housing. This leads to an increase in the median income in both A and B and higher local tax revenues, supporting an increase in local public good provision. As a result, purchase prices rise by 1.5% to $7482 in A and 1.26% in B to $5564 as low income households (those in the first and second income tax brackets) enter the purchase markets. In jurisdiction B, higher housing demand increases homeownership by 1% (column 8, Table 4). However, lower demand from existing homeowners is not offset by the rise in price and increase in rental demand, as a result lower tax revenues contribute to a slightly decreased provision of endogenous public good in both jurisdictions. Despite its seemingly regressive design, this policy increases utility for the majority of lower income households. For many higher income households, the utility benefits of the income tax cut are outweighed by the loss in MID. Indeed, the proportion of households in the top four income tax brackets who gain from this policy reform is lower than under the 28% cap.27 Finally, considering Table 6, the income tax rebate reform increases mortgage interest payments and landlord's rental revenues, but creates moderate net reductions in welfare, relative to the other reforms, through the higher housing costs that are inflicted on wealthier households. $3620 new owner scheme This final policy reform replaces MID with a new owner scheme that pays a lump sum of $3620 to new homeowners. Again, this is revenue equivalent to the 28% cap. The characteristics of the new equilibrium under this policy appear in the final rows of Tables 4 and 5. As with the other reforms, the removal of MID has the immediate effect of contracting housing demand amongst existing homeowners. The introduction of a New Owner Scheme, however, stimulates entry into homeownership amongst previous renters in the lowest tax bracket (as can been seen in columns 2 and 4 of Table 5). Simultaneously, the New Owner Scheme increases total housing demand and leads to a rise in purchase prices in both jurisdictions, rising by 15.2% in A and by 17.1% in B (column 3, Table 4). In turn, we observe decreases in the average units of housing demanded in the purchase markets, an increase in the average units of housing in the rental markets and substantially higher tax revenues in both jurisdictions. Despite relocation between A and B reducing median incomes in both jurisdictions, the provision of endogenous public goods rises by 2% in A and B (column 11, Table 4). Accordingly, previous homeowners who lose the MID are compensated in two ways: first, since property prices rise, they benefit from capital gains and second, they benefit from increased levels of public good provision. While focusing on new owners, this policy reform results in welfare gains for households in the first three income tax brackets. The key pathways through which those gains are delivered is by supporting the movement of lower income households into homeownership and increasing property values and local tax revenues, thus facilitating a greater provision of local public goods. Despite this, the policy represents the greatest reductions in utility for homeowners in the top two income tax brackets: these households face the complete removal of MID and are ineligible for the New Owner Scheme. In addition, persisting renters face welfare losses as a result of higher house (rental) prices. Returning to Table 6, we observe that the New Owner Scheme produces large reductions in landlord's rental revenues and household welfare.","This paper contributes methodologically to the existing equilibrium sorting literature by developing a model that incorporates an explicit endogenous tenure decision as well as endogenous local public goods. These innovations extend the range of policy problems to which ESMs can be applied to include those where tenure choice and the impact of policy reform on rates of homeownership are central. Moreover these innovations allow us to account for the influence of capital gains and housing stock constraints on the distribution of welfare changes. A simplified model was calibrated using real world data to examine the possible consequences of reforms to the policy of MID in the U.S. This exploration begins to shed some light on the complex patterns of change that such reforms may precipitate in the property market and provides insights that help to inform some of the more acrimonious disputes surrounding the debate over MID reform. With regard to that debate, the calibrated simulations show that the impact of removing MID depends crucially on the nature of the policy that takes its place. First, consider the argument that MID inflates property prices making homeownership less affordable (Glaeser and Shapiro, 2002). The simulation results suggest that while MID disproportionately reduces the cost of purchasing housing for higher income households, we do not find evidence to suggest that reforming MID would necessarily lead to reductions in house prices. To the contrary, our simulations indicate that entry into homeownership and greater public good provision could lead to rising property prices. Second, supporters of MID argue that removing the policy would damage homeownership rates. This is where the key innovation of our model, the introduction of endogenous tenure choice, enables us to make a significant contribution to the policy debate. Our simulations suggest that the impact of reform on homeownership may be positive or negative. Indeed, for the cap, income tax reduction and New Owner Scheme we predict increased homeownership rates as new incentives for homeownership are introduced. Furthermore, despite the fact that the model accommodates changes in tenure as the relative costs of renting and owning change, we predict quite small changes in homeownership rates for the cap, flat-rate credit and the income tax rebate. Third, critics of MID argue that it subsidizes excessive housing consumption amongst wealthy households, suggesting that the removal of MID would lead to a contraction in the average property size of owners in the top tax brackets. As with the homeownership rate, our simulations suggest that the nature of the policy reform has a strong influence on the mean property sizes demanded by homeowners in each income tax bracket. Contrary to previous predictions, however, in the cases of the 28% cap and the income tax rebate the mean property size demanded by homeowners in the top income tax bracket remains the same. In the cases of the flat-rate tax credit and New Owner Scheme, our conclusions are consistent with a reduction in the average purchased property size for households in the top income tax bracket. Model sensitivity ~~~~~~~~~~~~~~~~~ The model presented in this paper provides a tool for analyzing policy reforms and exploring the types of market adjustments that would characterize the resulting equilibria. It is important to acknowledge that the calibrated model approximates many dimensions of the joint housing, tenure and location decision that are not well understood. To test the robustness of the simulation results, and to identify the most influential parameters we explored a variety of parameter values for i) the endogeneity of local public good provision, ii) the specification of housing supply, iii) the degree to which rental and owned property markets are connected, iv) household preferences, and v) the tax incidence of property taxes for renters. For the purpose of conciseness a selection of these results, exploring the first two points, is presented in Table 7, further results are available from the authors upon request. Table 7 compares three specifications of local public good provision; i) calibrated feedback via local homeownership, ii) state level homeownership as a proxy and iii) an inverted calibration of the feedback effects, alongside two specifications of the housing supply function; i) fixed short term housing supply and ii) Cobb–Douglas housing supply function with elasticity of one. Across the range of calibrated parameter values, including those not presented in Table 7, we find consistent patterns of change in homeownership rates, predominantly with increases in homeownership rates being achievable through a deficit reducing policy reform. This is consistent with Shapiro and Glaeser's (2003) assertion that MID subsidizes households who would be homeowners even in the absence of the policy. Likewise, our simulations consistently demonstrate that increases in public goods provision under several of the reforms serve to compensate households for the reduction in federal spending; the magnitude of this compensation is sensitive to the value of preferences for public goods, α. Nonetheless, despite this sensitivity, our results consistently suggest that the proposed 28% cap leads to utility gains for the majority of households and would be supported, from a utility perspective, by the majority of households. Moreover, our results suggest that the benefits of the policy would be quite broadly distributed across the income tax distribution (see columns 8–13 of Table 7). In contrast, the impacts of MID reform on property prices and the average property size of homeowners are sensitive to the calibration of the model. In the simulation results presented in Section 4 we find that property prices do not decrease and the average property size of owners in the top tax brackets decreases. This finding is quite robust for the cap, income tax rebate and New Owner Scheme. However, for the flat-rate tax credit and in calibrations where either rental and purchase markets have been defined independently or renters do not face the full property tax burden, the model predicts that property prices could fall (consistent with the argument made by Bourassa and Yin, 2008) and the average property size amongst homeowners in the top tax brackets could increase under the 28% cap, flat-rate tax credit and income tax rebate reforms. Concluding remarks ~~~~~~~~~~~~~~~~~~ Examining a range of alternative policy reforms demonstrates the importance of policy design and the role of path dependency in shaping the outcome of those reforms. With regard to the latter, there are three key mechanisms at work. First, owning a property shields high-income households against rises in property prices and subsequently enables them to channel benefits through capital gains. Second, housing stock constraints fix the current capital stock of housing making it unresponsive to price changes, these constraints act to suppress price falls and stabilize homeownership in the face of contracting demand. Third, endogenous public goods can act as a mechanism for compensating households. As a result, the complex patterns of change precipitated by policy reforms in the property market can have quite unanticipated results. Policies designed to be progressive, such as the tax credit reform, may do less to benefit poorer households than those that appear to be regressive, such as the income tax reduction reform. Likewise, policies that economists would normally assume to have excellent efficiency improving qualities, such as the income tax reduction reform, may lead to significant net welfare losses. Taken as a whole, our investigation suggests that several reforms to MID could maintain the prevailing levels of homeownership while delivering more public goods and contributing to a reduction in the federal deficit. Of course, these results relate to the calibration of a simplified two-community problem. Given our results, it would be interesting to see future work directed towards the estimation of a large-scale model with more formally quantified social returns to homeownership. With these extensions it would be possible to simulate economy-wide responses to the proposed reforms. Nonetheless, our results demonstrate the usefulness of the modeling framework and provide important insights into the broader implications of reforming MID."],["This paper analyzes the interaction between credit constraints and trading behavior, decomposing trade in extensive and intensive margins. I construct a unique dataset containing firm-level trade transaction data, balance sheets and credit scores from an independent credit insurance company for Belgian manufacturing firms between 1999 and 2007. Firms are more likely to be exporting or importing if they enjoy lower credit constraints. Also, firms that have better credit rating export and import more. Importing and exporting behaviors differ in how both the level and growth of the various margins of trade are related to credit constraints in one important dimension. In the case of exports, it is the intensive and extensive margins of exports in terms of both product and destinations that are significantly associated with credit constraints whereas for imports it is the extensive margin in terms of products only. --------------------------------------------------------------------------------","When the economy enters a recession and a credit crunch shakes the financial sector, or simply when they suffer other types of shocks, firms might find it harder to access credit. This is likely to affect their operations, the investments and R&D they conduct and the way they develop and expand. There is however little empirical evidence on how important financial considerations are for the international expansion and activities of firms and on how they adjust their trading activities. This paper considers the determinants of firm trading patterns by matching firm-level trade transaction data with individual, time-varying credit ratings. In particular, it seeks to analyze the interactions between financial or credit constraints on the one hand and exports and imports margins on the other. While the analysis of firm-level trade has mostly focused on the relationship between trade and productivity, this paper contributes to a recent literature studying another critical determinant of trading decisions: the financial situation of the firm, and in particular the credit constraints it faces. On the one hand, a bad financial situation might make its suppliers based abroad less willing to risk trading with the firm, hence affecting its imports. On the other hand, being credit- constrained would prevent the firm from overcoming any fixed costs associated with either exporting or importing. Based on a unique and detailed dataset, I find that firms that enjoy lower credit constraints and bankruptcy risk are more likely to be exporting. Firms that have better credit rating also export more. They have higher extensive margins: they export more products to more destinations. Their intensive margin, the average export value per firm-product-country observation for all combinations with positive exports, is higher too. The same patterns hold for imports except that the country extensive margin and the import intensive margin, are not correlated with credit constraints. Finally, most of these results are shown to hold over time, when estimated with the growth of the various trade margins. This brings novel insights on the differences between import and export choices, and these correlations are useful to guide future theoretical work. The detail of the datasets used is particularly suitable for the questions addressed. First, the trade and balance sheet data cover the full sample of Belgian manufacturing, at the firm level, with detailed information on trade participation, but also on the destinations or origins and products traded. This allows for a full decomposition of trade flows into the intensive margin as well as firm, country and product extensive margins and the density of trade. Using firm-level analysis in this paper allows a better understanding of how firms vary within a given sector. Second, the measure of credit constraints used is unique in its kind: a yearly measure of the creditworthiness of firms established by an institution external to the firm. By assessing financial constraints with a continuous credit rating rather than single and extreme default payment episodes implies that the effects identified do not only hold in the case of extreme credit constraints. The paper's main contribution is to be using a measure of credit constraint that exhibits sufficient within-firm variation over time to relate it to the firms' importing and exporting outcomes after controlling for firm fixed effects. This paper contributes to two related areas of the literature. First, the relation between liquidity constraints and exports has been studied both in theory models and empirically. The next section will discuss relevant theoretical papers. Empirically, the question has also been studied using different datasets. Several papers use the sector-level Rajan and Zingales (1998) measure of “external finance dependence” to examine how it affects exports. Manova (2008) shows how financial frictions and credit market development explain cross-country patterns of trade at the sector level. Export growth is proven to be slower in external finance dependent sectors in Iacovone and Zavacka (2009). Considering Belgian exporters as in this paper, Behrens et al. (2013) find that imports fell more during the recent recession for firms with above median reliance on external finance. Analyzing transaction-level data and the financing terms used by firms, Antràs and Foley (2015) show how trade finance, and hence financial constraints, influence the impact of macroeconomic and financial crises. Analyzing monthly US imports from countries with varying degrees of credit market tightness, Chor and Manova (2012) demonstrate that exports are less sensitive to the cost of external capital in industries relying less on external finance or less financially- vulnerable according to other measures, and that this sensitivity rose during the financial crisis.1 Albornoz et al. (2012) analyze the sequential pattern of exports expansion by successful new exporters for Argentina and find that credit constraints do not explain these patterns.2 The firm-level dimension of the dataset I use allows me to go beyond this type of sectoral analysis and exploit intra-sector variations. Others have used firm-level measures to capture credit constraints. Minetti and Zhu (2011) analyze a cross-sectional survey of Italian manufacturers and find that “credit rationed” firms are less likely to export and are likely to export less. They focus on the firm extensive margin of exports. Berman and Héricourt (2010) find similar results for developing countries, using financial ratios as measures of constraints. They find that financial constraints are not correlated with export values or export survival in those countries. Other authors also explore these questions by deriving measures of financial health and constraints from balance sheet financial ratios. Greenaway et al. (2007) find no ex-ante effect on the probability of becoming an exporter while Bellone et al. (2010) do. Askenazy et al. (2011) find that credit constraints negatively affect the entry into a new destination and increase the probability of exiting a market.3 Exploiting data available from the international firm-level data from the World Bank Enterprise Surveys, Wang (2011) reports that the probability of exporting and the export volume increase with age, which is consistent with the hypothesis that firms need to accumulate sufficient collateral before they can borrow enough funds to profitably export. This paper extends this literature by analyzing the extensive margins at the level of destinations and products as well as considering financial constraints as a continuous measure. A second area of the trade literature has empirically analyzed firm-level imports. The fact that the import of new varieties leads to higher productivity and growth has been shown empirically both at the country level (Broda and Weinstein, 2006) and at the firm level, with imports of intermediates or reductions in input tariffs being associated with productivity gains or higher productivity levels (see Antràs et al. (2014); Amiti and Konings (2007); Kasahara and Rodrigue (2008) or Goldberg et al. (2010)). While my results do not contradict these findings, the question I address is whether financial constraints might prevent some firms from reaping the benefits of importing intermediates. Although it may appear surprising that financial constraints shape the country-level extensive margin differently on the export side and on the import side, these results are consistent with the findings of Antràs et al. (2014). These authors show that the determination of the country-level extensive margin of importing is much more complicated than in models of selection into exporting, mainly because the marginal increase in profits from adding a country to a firm's set of potential sourcing locations depends on the number and characteristics of other countries in the set. Imports and exports have often been compared in the recent analysis of firm-level trade data. Descriptive evidence by Bernard et al. (2007), Muûls and Pisu (2009) or Halpern et al. (2005) shows that importing firms share many attributes with exporters: they are both larger and more productive and product and country-level patterns of trade at the firm level are similar for both. Based on a theoretical model estimated with Chilean data, Kasahara and Lapham (2013) analyze the complementarities of exports and imports for the productivity and welfare gains of trade. This paper shows that imports and exports are different in some important dimensions when put in relation with credit constraints. The remainder of the paper is organized as follows. Section 2 presents the conceptual framework of how financial constraints can affect exporting and importing patterns. Section 3 describes the data, and demonstrates in particular why the Coface score is an appropriate measure of credit constraints. Section 4 contains the empirical analysis of the links between export and import patterns and credit constraints at a given point in time, while Section 5 takes a closer look at trade growth. Section 6 concludes.","Why are credit constraints analyzed here specifically in relation to trade transactions and not both domestic and international activities? External finance is indeed used by firms to finance investments in capital, product development and R&D or advertising among others. However, in addition, trading firms face incremental fixed and variable expenses thus needing even more access to external capital. In order to frame the empirical analysis below, one can decompose total trade flows. What are the impacts of credit constraints on these different export margins that are predicted by the existing models? Focussing mainly on a framework with only a firm level extensive margin as in Eq. (1), Chaney (2013) and Manova (2013) introduce credit constraints in a theoretical model of trade with heterogeneous firms à la Melitz (2003) and yield several predictions on the equilibrium relationships between productivity, credit constraints, exports and export margins. In both models, firms must pay up-front a fixed cost of entry into foreign markets and hence need sufficient liquidity to do so. In Chaney's model, firms finance these costs with cash flows from their domestic operations. Once a firm has entered foreign markets, financial constraints do not impact the marginal cost of exporting: the firm will finance an increase in the scale of its exports through its internal cash-flow and foreign trade credit. In equilibrium, financial constraints impact the extensive but not the intensive margin of exports. In Manova (2013), firms need to borrow to cover both the fixed and the variable costs of exporting. This follows from the imperfect enforceability of international transaction contracts together with imperfect information on the potential returns from foreign markets. In equilibrium, total exports will increase with lower credit constraints. More productive firms and less credit-constrained firms will be more likely to export. Credit constraints will decrease the firm extensive margin and the overall intensive margin vt (average exports per firm). To summarize, these models illustrate that in equilibrium, credit constraints affect the intensive (respectively, extensive) margins of exports if financial constraints are assumed to affect the variable (respectively, fixed) costs of exporting. There is to the best of my knowledge no model that explicitly considers firm exporting decisions in relation to credit constraints with a distinction of export margins according to Eqs. (2) and (3).5 In a stylized dynamic model, Besedeš et al. (2014) focus on changes in the intensive margin and show that credit constraints can have an important role in the start of exporting activity, but not on the growth of exports in later stages. There is also limited empirical firm-level evidence on these issues, with existing studies focussing on export margins as defined by Eq. (1), or considering whether or not the firm exports to more than one market as in Minetti and Zhu (2011). By illustrating the correlations between these constraints on the one side and the different extensive margins and the intensive margin of exports on the other, the results of the empirical analysis below should serve as a motivation for future theoretical work in this area. The Belgian balance sheet and trade transaction data ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ This dataset provided by the National Bank of Belgium has been used in several papers analyzing export and import patterns and behavior (see Muûls and Pisu (2009); Behrens et al. (2013) and Araujo et al. (2012) among others). It merges firm-level balance sheet and trade data for Belgium. The balance sheet part of the BBSTTD is used to extract firm-level annual characteristics, including employment, value added, profitability, sector of activity and to compute total factor productivity. The trade data includes the value, destinations and origins as well as products at the 8-digit Combined Nomenclature (or CN8) level of the EU.6 Manufacturing firms only are selected as belonging to sectors 15 to 36 of the NACE-BEL classification.7 The data is then merged into the Coface database, described in the following subsection. Only firms for which a Coface score is given for each year a balance sheet was available are included in the dataset. All observations are kept in the resulting dataset, which is described in Table 3. The Coface score As a measure of credit constraints, I use the Coface Services Belgium Global Score for more than 9000 Belgian manufacturing firms between 1999 and 2007. Credit insurance firms offer insurance policies to businesses that wish to protect their accounts receivable from loss due to commercial and political credit risks such as protracted default, overdue accounts, insolvency or bankruptcy. Established in France in 1945 as a credit insurance company, Coface is now an international firm, providing a range of services to facilitate business-to-business trade. Among these, it also provides credit worthiness information: through a worldwide network of credit information entities, it has constructed an international buyer's risk database on 44 million companies. Data from public and private sources are added to Coface's internal data in order to manage each company's rating and Coface's risk exposure on a continuous basis. Fig. 1 provides an example of who the parties would be in a credit insurance contract. Belgian firm A is a client of Coface. It wants to sell a product to Belgian firm B and protects itself against the risk that B might default on paying for that product. As it would do for any transaction, it asks Coface whether it will be covered. Coface, if it hasn't already done so, will compute a score for B, and if it is high enough will insure the transaction. A will send the product to B in exchange for payment. If B defaults, Coface will pay A and seek to recover the amount from B. Imagine that Belgian firm B also exports goods to firm C and imports inputs from firm D. Neither B, C nor D are clients of Coface, but because of its transaction with firm A, the credit rating score of firm B will be in the dataset. There is a large academic literature on bankruptcy prediction models such as that used to construct Coface's score (see for example the review by Balcaen and Ooghe (2006)). However, privately-computed probability of default or credit scores such as Coface's are naturally less well-known. Various datasets are compiled to construct the Coface score: the firm's financial statements (leverage, liquidity, profitability, size, etc.), its legal form, age and life cycle, location and information on its commercial premises as well as industry specific information. Data on payment incidents both with other firms and to the social security (ONSS) are also used. Finally, legal judgments and the board structure are taken into account. For example, if a firm goes bankrupt it will negatively affect the score of all companies that have common board members. These various inputs are combined using several statistical methods and trial-and-error. The result is a score ranging from 3/20 to 19/20. Although the model predicts continuous scores they are rounded to unity in the obtained data. The score therefore contains information about the firm's quality, performance and productivity. However, two firms with equally valuable projects, and identical profitability and productivity can be very different in terms of financial health, board structure, and other elements that will determine their score and their access to credit. The empirical analysis will therefore control for a number of variables that could potentially influence both the Coface score and the trading activity, such as size, profitability and productivity of firms.8 A firm's score varies from year to year: the average yearly change is 2 points (or 12.5% of the largest possible variation from 3/20 to 19/20) with a standard deviation of 2. The average difference between a firm's largest score over the time sample and its lowest is 6 points. It is this variation that I exploit in my analysis below. Importantly for the purpose of this paper, Coface's score does not include in its determination model any information on the firm's exports or imports. Firm-level trade data is not public information in Belgium and even if Coface would have such information for some firms through it's network of international clients, it does not enter directly in the computation of the score. However, trade performance might affect the score indirectly, through its impact on profitability for example or other variables that are included in the construction of the score. While it does not remove all potential endogeneity in the Coface score, great care is taken in the remainder of the paper to include crucial firm-level controls and to exploit the panel dimension of the data. Constructed as a bankruptcy risk measure, the score is highly correlated with how credit-constrained a firm is, reflecting the same type of information that a bank would use to decide whether it lends to a firm. Being determined independently by a private firm, it is unusual for such data to be available and has a great advantage on measures of credit constraints used in the literature so far: it is firm-specific, varies through time on a yearly basis9 and allows for a continuous measure of the degree of credit constraint rather than classifying firms between two constrained or unconstrained categories. Compared to datasets on payment incidents that would identify a small subset of firms as being credit-constrained,10 the Coface score ranks firms along the whole spectrum of ratings. Payment incidents would only be one of the elements affecting the score, in combination with many others. Overall, the Coface score is a well-suited direct measure of creditworthiness used by other firms and by banks when extending loans, and I therefore use it in my empirical analysis to measure how credit-constrained firms are. External validation This section presents the correlation between the score and firm fundamentals. It also relates it to the important literature on credit constraints, in particular in corporate finance. Given the methodology used to construct the score is not available publicly, it is shown here how correlated the score is with the firm's financial situation and productivity. A selection of financial ratios11 measures each firm's solvency and investment. Table 1 shows how strongly the Coface score is correlated with the financial situation of the company, in particular its solvency and investment intensity. Firm and year fixed effects are included in the OLS regression, thus also controlling for possible differences in, for example, risk premie across industries and years which might affect the Coface score and other financial measures differentially. Solvency is measured with two ratios: financial independence and coverage of borrowings by cash flow. The strong correlation between these and the score shows that firms that are more able to meet their short- and long-term financial liabilities have a higher score. Financial independence, the ratio between equity capital and total liabilities, reflects how independent the firm is of borrowings. The coverage of borrowings by cash flow measures the firm's repayment capability, and its converse specifies the number of years it would take to repay its debts assuming that its cash flow was constant. Higher scores are also associated with larger investment ratios, the acquisitions of tangible fixed assets over value added. EBITDA (Earnings before Interests, Taxes, Depreciation and Amortization) is a commonly used financial measure of the operational profitability and performance of the firm. It appears as being positively associated with the Coface score. I will include it as a control in the regression analysis below, in order to control for the effects of the profitability of a company. Productivity has been shown to be an important determinant of trade patterns. It is measured here as in Levinsohn and Petrin (2003).12 Column (5) of Table 1 reports a positive but not perfect correlation of the Coface score with productivity, confirming that credit constraints and productivity are two different issues to be considered when analyzing export behavior. The effects of financial constraints on firm behavior are an important area of research in corporate finance. Compared with the existing literature, the Coface score provides many advantages, as described above. One of the many approaches in the literature consists of sorting firms into financially- constrained and unconstrained types on a yearly basis by ranking firms according to different measures. In Almeida et al. (2004), firms in the top three deciles of their payout dividend ratio would be considered as less financially-constrained than firms in the bottom three. Allayannis and Mozumdar (2004) use total assets. I test in Table 2 whether the score is consistent with such classifications. The mean Coface score, its standard error, maximum and minimum observations are reported separately for constrained and unconstrained firms. Firms whose dividend payout is in the top 30 percentiles are considered as financially unconstrained, whereas those in the bottom 30 percentiles are financially constrained. The same is done with total assets. The mean test is passed, meaning that constrained firms have a lower score than unconstrained firms, in both criteria. It is robust to using only one cross-section of the data, or taking out observations within the top and bottom percentiles of each measure. This confirms that the Coface score offers a creditworthiness measure that is consistent with other measures used in the literature. Export or import status ~~~~~~~~~~~~~~~~~~~~~~~ I begin the empirical analysis by exploring the variation in credit scores between exporters and importers on the one hand and non-traders on the other. It appears that less credit-constrained firms are more likely to be trading. This is shown at first in the descriptive statistics presented in Table 3: on average, traders are not only significantly larger and more productive, they also have a significantly higher score, meaning they are more creditworthy and less liquidity-constrained. Where Ex/Importer(0/1)i,t is a dummy that takes the value 1 if firm i is an exporter/importer at time t and zero otherwise. CSi,t − 1 is the Coface credit score13 for firm i at time t − 1 and additional firm characteristics are added: productivity, operational profitability, wage, age, employment, MNE status and financial ratios proxying firm access to finance. Of course, many other factors might affect a firm's export status such as the current economic situation, and other characteristics of the firm. Other potential factors such as exchange rates, factor endowments, factor prices or industry demand will be common to all exporters of a sector in a given year. This is why I include firm and sector-year fixed effects in my specifications, denoted by {FE}, thus eliminating any bias that they could cause. The results are also presented with an alternative set of fixed effects – year and sector – for comparability with the previous literature. Each firm only belongs to one sector so in the specification where firm fixed effects are not included, sector fixed effects are used to control for non-time-varying sector-specific idiosyncrasies. Including fixed effects, controlling for firm-level observables and given the composition of the score described above, the residual effect of the Coface score is a good measure of credit constraints faced by a firm. Given the number of fixed effects to be included in the specification, using a linear probability model in levels addresses the incidental parameter problem that affects non-linear fixed effect estimates. This specification is used in Bernard and Jensen (2004) for a very similar binary choice problem despite the problems this might provoke (e.g. predicted probabilities outside the 0–1 range). The first four columns of Table 4 only include sector and year fixed effects, as in Bernard and Jensen (2004). In columns (5) and (6), the full set of sector-year and firm fixed effects are included. Finally, as a robustness check, the results using a conditional logit estimator with year and firm fixed effects are presented in column (7). Other firm characteristics are also included as controls in columns (3) to (7): operational profitability, wage levels, age of the firm, multinational status and the financial ratios presented in Table (1). As would be expected, the Table shows that more productive firms are more likely to export, although the coefficient becomes insignificant once firm and sector-year fixed effects are included. The coefficient on the lagged credit score is positive and significant in all specifications, confirming that firms which are less credit-constrained have a higher probability of being exporters. In column (2), the coefficient on productivity is not reduced compared to the first column, indicating that the score captures the additional effect of credit constraints. Column (3) shows that the effect of productivity decreases while that of the credit score increases when more controls are added. When including the lagged export status variable, as in Bernard and Jensen (2004), the coefficients on TFP and the Coface score are strongly reduced as shown in column (4). Columns (6) and (7) show that the sign and significance of the score coefficient remain robust to including firm as well as sector-year fixed effects. Conditional logit with firm and year fixed effects yields similar results although the significance of the coefficient on the Coface score is lower, as shown in Column (7). These results are consistent with the literature showing that the firm extensive margin of exports (fm in Eq. (1)) is positively correlated with lower credit constraints. Very similar results are obtained when estimating the effect of the Coface score on import status using exactly the same specifications. As shown in Table 5, one notable difference with export status is that the positive effect of productivity on the probability of being an importer remains strong and significant across the different specifications. Less financially-constrained firms have a higher probability of being importers. Also, the strongly significant coefficient on the Coface score is robust to including sector-year and firm fixed effects as well as the lagged importer status in the linear probability model as shown in column (6) or the conditional logit specification in column (7). The coefficient of the credit constraint score is larger than in the case of exports. Import and export status, and hence the firm extensive margins for both trade flows, are therefore correlated to credit worthiness in a very similar way. These first results show that there are factors of credit worthiness that the Coface score integrates, over and above the observed balance sheet firm-level variables that are included as controls and to which the score is correlated. The regressions here and in the next sections are highlighting the effects of these additional elements of credit worthiness. Value, destinations, origins and products ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In order to obtain a sense of the magnitude of these effects, one can compute that a one standard deviation increase in the log of the credit score corresponds to a 4.4%14 increase in total exports. This can be compared to the average annual increase in total Belgian exports between 1995 and 2008 which was 5.4% (Baugnet et al., 2010). Comparing the different margins, a 10% increase in the score (and not the logarithm of the score) would correspond to a 0.4% increase of both the number of products exported and destinations served and a 0.45% increase of the intensive margin. This is a small figure in the case of products given that the mean annual increase in the number of products exported by a firm in the dataset is 6.5%. It is however a much larger result for countries given that on average firms do not increase the number of countries they export to. In the case of imports, the lagged score is positively and significantly associated with total imports as well as the number of products imported, the product extensive margin. But in contrast with the corresponding result for exports, the coefficient on the number of origins is insignificant. A potential explanation behind this result is that firms rarely source goods from multiple countries. The mean number of countries from which a firm imports a given product, whether defined at the 8-digit or 6-digit level, is 1.5, while it is 3.68 countries per exported CN8 product. This has also been found using U.S. data in Antràs et al. (2014) and it could be due to the fixed cost of expanding on the product extensive margin of imports being smaller than the fixed cost of expanding on the country extensive margin. The fact that the country-level extensive margin of imports is not significantly shaped by the credit score is also consistent with the model in Antràs et al. (2014). Another channel to take into account is that if a firm does not have a good financial health, foreign firms will be less willing to risk a payment default, limiting the range of inputs it can import from abroad. There appears to be no relationship between the import intensive margin and credit constraints, which would indicate that variable costs for imports are less related to the access to external finance. One standard deviation increase in the logarithm of the score would correspond to a 3.4% increase in total imports, suggesting that access to credit matters less than in the case of exports on aggregate, although the coefficient for the product extensive margin is slightly higher than for exports. These various results will be analyzed in their growth dimension in Section 5. It is also interesting to note that for exports, only total value and the intensive margin are positively related to productivity while for imports, it is also the case of the number of origins. A higher EBITDA is positively associated with all margins for both imports and exports. This suggests that firms that reach a certain maturity make bigger margins on their products thus obtaining higher profitability and exports. A higher EBITDA might also reflect lower input costs which could be the result of a more intense importing behavior. The negative and significant coefficients for the financial independence ratio that appear in most specifications of Tables 6 and 7 could be due to the fact that a higher ratio at t − 1 will facilitate the firm's potential to increase its liabilities which in itself will decrease the ratio at t. This decrease has a negative impact on trade values and its margins. Finally, the strongly significant and large coefficient for the employment variable in the case of the extensive margins confirms that it is important to control for firm size when considering the trading decisions of firms. These results clearly establish the relationship that exists between credit constraints and exporting and importing patterns, even if once productivity, size, profitability, access to finance and other firm characteristics are controlled for. Extensive and intensive margin for exports ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ I found in Section 4.2 that firms with a higher credit score are also likely to display higher values of total exports, to be exporting more products, serving more destinations and to have larger average exports per non-zero product-country. This result also applies to the growth in total exports and its components, except for the intensive margin. In this section, I run the same specification as in Eq. (5), but taking as dependent variables the various elements of Eq. (6). I include firm fixed effects, sector-year dummies and control for firm characteristics in an OLS specification as above. Levels of the dependent variable in the previous year are included as explanatory variable in each case, to capture together with the firm's age the existing exporting activity level effects on each margin, although the results are robust to excluding them from the regression. The first column of Table 7 shows that the increase in total exports relative to the previous year is positively related to creditworthiness. One standard deviation increase in the logarithm of the score can be associated with a 2.5% increase in the growth of exports. This can be decomposed into a positive relation with the increase in the number of destinations served as shown in column (2) and the increase in the number of products (column (3)). It confirms that credit constraints can be associated with the fixed costs of exporting to more countries or more products. Variations in the intensive margin of trade, the dependent variable of column (5), are not correlated with credit constraints. The coefficients on productivity are not significant when considering the product and destination extensive margins suggesting that when looking at changes over time, it is important to also consider other determinants of trading patterns. Besides, being part of a multinational is positively associated to the increase in the number of destinations and products exported but not the intensive margin. Finally, the export levels, whether in value, products or destinations in t-1 are correlated to the total growth as well as to all three margins negatively: firms are less likely to grow in all dimensions if they are already strong exporters. Extensive and intensive margin for imports ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Identical specifications are run for imports. There is a strong statistical relationship between the growth of imports and the Coface score at t-1: an increase of one standard deviation in the logarithm of the score is associated with a 2.6% increase in the growth of imports, a figure very similar as that for exports. However, in contrast with the results for exports when decomposing into different margins, credit worthiness appears to be strongly significantly correlated with the imported product extensive margin but not with the growth in the number of countries imported from, as shown in columns (7) and (8) of Table 7. This suggests that liquidity is not correlated with importing from more origins, but is positively so with expanding the range of imported inputs. As discussed above, this correlation suggests the presence of sunk costs of importing an additional product from abroad. In addition, it could also be related to the default risk reflected in the score: foreign companies will be less willing to take this risk if the importer has a lower credit rating. As in the case of exports, changes in average import values per active product-country are uncorrelated with credit constraints. The coefficients on operational profitability are positive and significant across all import margins, including the intensive one. It is also noteworthy that MNE status is not correlated with import margins. The results in Table 7 reinforce the findings from Section 4.2 that considered the levels of the trade margins. The lack of association between number of countries imported from and credit constraints could reflect the fact that contrary to exports where more destinations imply larger markets, the primary rationale for imports is the necessity to source inputs that are either not available domestically, or not at the same levels of price or quality. The novel finding is that at the firm level, the relationship between credit constraints and the country extensive margin is different for imports and exports.","In this paper, I show that credit constraints are related to export and import volumes and patterns. A precise and complete dataset on trade transactions at the firm level for the Belgian manufacturing sector is combined to a unique and very useful yearly measure of credit constraints faced by firms, a creditworthiness score constructed independently by a credit insurance company. These allow me to examine the relationship between credit constraints and trade in a new way. My main contribution is to show that credit constraints are important across the spectrum, not only in cases of payment defaults and that they are correlated differently with export and import margins. It is shown that firms which are less credit-constrained, more productive and profitable have a higher probability of being exporters or importers. Such firms are also likely to report larger total trade values. While credit constraints are positively associated with both the country and product extensive margins and the intensive margin of exports, this is not true in the case of imports where it is only the case for the product extensive margin of imports. Finally, this result also holds when decomposing the growth of exports and imports. I find that the growth in the number of products exported and destinations served is positively correlated with the Coface score measure. This supports the hypothesis that entering a new market or exporting a new product imply fixed costs for exporters. On the other hand, for imports, a rise in its credit score is associated with an increase in the number of inputs imported by a firm. This might reflect the impact of the firm's financial situation on both a firm's potential to pay the fixed cost of sourcing an additional input from abroad, as well as the willingness of its potential suppliers to take the risk it might default. Related to the fact that a firm rarely imports a good from multiple countries, I find that a decrease in credit constraints is not positively associated with an increase in the number of countries that a firm sources its imports from. These results confirm the link between credit constraints and export and import margins. They also highlight the potential role of government agencies in reducing the fixed costs of entry to new markets or of importing new inputs. Exploring further the relationship between financial constraints and trading behavior, by using firm level information on specific products' domestic sales versus exports, could shed further light on the links between the dynamics of trade and financial constraints."],["This paper introduces a new ‘supply-push’ instrument for foreign aid, to be used together with an instrumental variable estimator that filters out unobserved common factors. We use this instrument to study the effects of aid on macroeconomic ratios, and especially the ratios of consumption, investment, imports and exports to GDP. We cannot reject the hypothesis that aid is fully absorbed rather than used to build foreign reserves or exiting as capital flight, nor do we find evidence of Dutch Disease effects. Aid leads to higher consumption, while the evidence that it promotes investment is less robust. --------------------------------------------------------------------------------","Although a large literature studies the effects of aid using cross-country regressions, its problems are well-known. Researchers must contend with the endogeneity of aid, the high persistence of output, the uncertain determinants of growth rates, nonlinear effects of aid, biases from measurement error, and the likelihood of substantial heterogeneity in the effects of aid. Moreover, since aid is given in many different forms and with a variety of motives, these regressions invite concerns that are not purely statistical. For its detractors, this literature uses unreliable data to arrive at fragile answers to the wrong question. These criticisms may seem decisive, but some important questions are hard to answer without cross-country data. In this paper, we seek to advance the literature in two ways. First, we introduce a new ‘supply-push’ instrument for aid, to be used together with an estimator that filters out unobserved common factors, even when their effects differ across countries. In principle, this combination of instrument and estimator will identify the causal effect of aid under more general conditions than existing approaches. It could be applied to a wide range of aid-related questions in future research. Second, we shift the focus to whether and how foreign aid is absorbed by the domestic economy. Aid, as a capital transfer, is not part of measured GDP. The aid could be absorbed, by allowing increased domestic expenditure, but this is not the only possibility. It might be offset by a corresponding capital outflow, or used to accumulate foreign exchange reserves. Some of the aid flows recorded by donors will not correspond to international transfers: for example, some forms of donor-sponsored technical assistance will not have a direct effect on the recipient's domestic expenditure. In all these cases, aid is not absorbed by the domestic economy. For absorption to take place, domestic expenditure must increase relative to domestic production, implying an increase in net imports. Hence, we begin by examining the causal effect of aid on net imports. We are also interested in how absorption takes place. Absorption requires an increase in at least one of the components of domestic final expenditure: household consumption, government consumption, and gross investment. We study the effects of aid on the ratios of these components to GDP. This should help us to understand the potential effects of aid. For example, if aid improves the investment climate, we would expect to see an increase in investment relative to GDP, and this has been a central concern of the empirical aid literature since its inception. We will argue that the effects of aid on investment, and other macroeconomic ratios, are easier to study than the effects on growth. Further, in our empirical work, these effects will be identified from relatively persistent changes in aid receipts, rather than transitory and endogenous changes. The well-known identification problem in the cross- country literature is that aid is not randomly assigned. To address this problem, we introduce a supply-push instrument. It is based on the idea that the exposure of recipients to changes in donor budgets varies across recipients. Consider two aid recipients, A and B, and a single donor. Country A accounts for a larger share of aid from the donor, and this greater exposure persists over time. In that case, when the donor's budget increases for some exogenous reason, the movement in aid is larger for country A than for country B, driven solely by the changing supply of aid. This suggests the following instrument: we can construct a synthetic measure of aid at each date t, based on each country's share of aid in a donor budget at some initial date t0, multiplied by the current donor budget at date t. As an example, consider what happens if the British aid budget increases relative to the French aid budget. Former British colonies are likely to see an increase in aid received, relative to former French colonies. More generally, there will often be long-term connections between particular donors and recipients, so that recipients are more exposed to variation in some donor budgets than others. Our instrument uses changes in total donor budgets, weighted by the initial shares of recipients in those budgets, to isolate exogenous changes in aid receipts that are not driven by the conditions of individual aid recipients. We call this a supply-push instrument; it is related to the work of Bartik (1991) on regional economics and Card (2001) on the labor market effects of immigration. As in the immigration setting, the origins and destinations of flows of aid are large in number, and this makes it unlikely that the instrument — as a weighted average of many donor budgets — will be correlated with recipient-specific conditions. We investigate this further below. A possible objection is that donor budgets may be influenced by forces common to many recipients. For example, world economic conditions are likely to affect donor generosity, and also the outcomes of poor countries. Drawing on recent work in the panel time series literature, global forces can be seen as unobserved common factors with loadings that differ across countries. We filter these out using an instrumental-variable version of a common correlated effects (CCE) estimator. This class of estimators was introduced by Pesaran (2006) and extended to instrumental variables by Harding and Lamarche (2011). Our paper is the first to apply this approach to the study of foreign aid. The combination of the new instrument and estimator should mean that we identify causal effects of aid under more general conditions than the existing literature. We find that aid is at least partially absorbed, reflected in net imports. We cannot reject the hypothesis that aid leads to a one-for-one increase in net imports, corresponding to full absorption. This occurs mainly through an increase in imports rather than a decline in exports, and hence we do not find symptoms of Dutch Disease. The findings hold across a range of estimators and robustness checks. There is similarly robust evidence that aid leads to increases in total consumption. This appears to be driven by increases in household consumption, but those estimates are less precise unless we exclude outliers. The evidence that aid promotes investment is weaker. In some models, aid has a delayed effect on investment, but these results are sensitive to the estimation method and the exclusion of outliers. The next section will sketch possible relationships between aid and macroeconomic ratios. Section 3 explains the approach to estimation and its relation to the literature. Section 4 describes the data. In Section 5, we analyze whether and how aid is absorbed, and the possibility of Dutch Disease. Section 6 presents robustness checks, before Section 7 concludes. An appendix describes the CCE IV estimator.","From a national accounts perspective, foreign aid is a capital transfer which does not contribute directly to GDP, but in principle allows an increase in domestic expenditure on final goods and services, relative to domestic production. Alternatively, aid might be used to accumulate foreign reserves, or lead to a capital outflow. Some aid may be spent on consultants who work exclusively in the donor country, with no direct effect on the aid recipient's domestic expenditure. It is therefore interesting to ask whether aid is absorbed. Domestic absorption is typically defined as the sum of household consumption, gross investment, and government consumption. We are interested in (1) whether aid is reflected in higher domestic expenditure on final goods and services, and (2) which expenditure components are most affected. This helps to clarify what is at stake in the paper. We show that aid is generally absorbed — it increases expenditure relative to output — but also find that consumption responds more strongly than investment. We do not uncover any symptoms of Dutch Disease. These results do not establish whether aid is ‘effective’, a hard task for a single paper, but they do contribute new evidence to the relevant debates. We first consider what it means for aid to be fully absorbed. Take Y as GDP, equal to the sum of household consumption C, gross investment I, and government consumption G, minus net imports M − X. For aid to be absorbed, at least one of C, I or G must increase, along with their total. If they increase relative to GDP, the GDP identity implies that the ratio of net imports to GDP, (M − X)/Y, must also increase. There is nothing problematic about this; it is what must happen if aid permits greater domestic expenditure relative to domestic production.1 In the short run, if aid is devoted to higher domestic expenditure on final goods and services, net imports will rise one-for-one with aid. If the response of net imports is smaller than this, aid absorption is only partial. To study absorption, we take the ratios C/Y,I/Y,G/Y and (M − X)/Y as our dependent variables. The way aid is absorbed might differ between the short and the long run. In the short run, aid might be used to build foreign exchange reserves which are used to finance higher expenditure only later, so that full absorption is temporarily postponed.2 More generally, the relationships between macroeconomic ratios and aid could be complicated over longer time horizons. If aid is spent in ways that improve the investment climate, the long-run effect of aid on investment could be much larger than the short-run effect. Or consider what happens when donor funds are spent on consultants working in the donor country: short-run absorption will be zero, but technical advice may later be reflected in economic policies and hence in macroeconomic ratios. Our empirical analysis will distinguish between short-run and long-run effects, by estimating dynamic models, sometimes with a role for lagged aid. The models we estimate can be related to macroeconomic theories of the aggregate effects of aid. In the one-sector Ramsey model, a permanent increase in aid raises the investment ratio in the short run, but not the long run. Aid promotes faster convergence to the steady-state, but the long-run levels of the capital stock and GDP are invariant to aid (Obstfeld, 1999). Along the balanced growth path, all aid is consumed. From a national accounts perspective, consumption is higher while investment and GDP are unchanged, and the increase in steady-state consumption is permitted by imports of the final good. When the ratio of aid to GDP increases permanently, the long-run C/Y and (M − X)/Y ratios increase to the same extent, leaving the other ratios unchanged. In a two-sector model of a small open economy, with traded and non-traded goods, the effects are more complicated. Aid may increase or decrease the long- run capital stock and gross investment, depending on whether traded production is relatively capital-intensive. Section B.1 of the online appendix discusses this in more detail. Models with balanced growth paths typically imply that the long-run consumption and investment ratios are stable functions of structural parameters. In a model with shocks, there would be a common stochastic trend in consumption, investment and output, while the long-run ratios would be mean stationary. The long-run ratio of consumption to output would be linear in the ratio of aid to GDP, where the intercept depends on structural parameters and can be treated as a fixed effect. By working in terms of ratios to GDP, we stay close to these predictions and avoid the non-stationarity that would arise with alternative explanatory variables, such as aid per capita. One drawback is that scaling our variables by GDP risks inducing correlations between variables that were originally unrelated. We examine this possibility in detail, including the pattern of results across the different ratios. We consistently find effects of aid on some ratios and not others, where the pattern conforms with the predictions of theoretical models, and where the effect sizes have plausible magnitudes. To generate this pattern, a story based on spurious correlations might have to be somewhat contrived. Other reasons to be wary of that explanation include additional results from first-differenced models, and models with the dependent variable in logarithms and which also include the logarithm of GDP as an explanatory variable. There are some advantages to studying absorption rather than growth. First, the relationships between macroeconomic ratios and aid intensity are more likely to be linear, for the reasons just discussed. Second, if aid improves the conditions for domestic investment, this effect might be relatively easy to detect. If we see consumption and investment as jump variables, they can respond quickly to changes in aid. In contrast, GDP is a function of state variables such as the capital stock: the relevant effects of aid could take time to emerge, and empirical researchers have to contend with the high degree of persistence of GDP. With these points in mind, it should be easier to establish reliable findings for absorption than for growth. Our approach allows us to make progress on some fronts, but not others. Although our instrument has some major advantages, it would be difficult for us to adapt the current approach to allow for parameter heterogeneity or non-linearities. And given the limitations of the available data, we cannot disaggregate aid while retaining a sample large enough for the methods that we adopt. The estimates we obtain might be best interpreted as the average effects of typical or business-as-usual aid. The paper is therefore complementary to previous work, given its narrower focus and the trade-offs that inevitably arise in addressing some econometric issues and not others.3","As we noted earlier, when studying the effects of aid, a central problem is that aid is not randomly assigned across countries. Even in a model that controls for country and time fixed effects, it is likely that aid flows and outcome variables are jointly influenced by time-varying variables that are not readily measured. To address this, we adopt a supply- push instrument which is a weighted average of donor budgets, where the weights are fixed over time but vary across aid recipients. One advantage of the new instrument is especially worth a discussion. It is likely that much empirical work on aid conflates the effects of permanent and temporary variation, just as early work on consumption conflated the effects of permanent and transitory income (Carter, 2017). From a policy perspective, a researcher might be more interested in determining the effects of a permanent change in aid. One solution is to use an instrument that is correlated with the permanent component and not with the transitory component. Since our instrument is a weighted average of donor budgets, and individual donor budgets are persistent, it should come closer than some precursors to identifying the effects of permanent changes in aid. In using this instrument, we are assuming that the total aid budgets of most donors are not greatly influenced by the time-varying conditions of many individual aid recipients. As background, aid flows are increasingly fragmented. The number of significant donors has increased, and most donors provide aid to a large number of countries. This is documented in Annen and Moers (2017), Djankov et al. (2009), Easterly (2007) and Knack and Rahman (2007), among others. Even as early as the 1970s, the US accounted for less than a quarter of total aid flows. According to Annen and Moers, the average bilateral donor provided aid to about 20 recipients in 1960, rising to 87 recipients by 2011. For identification, we rely on within-donor and within-recipient fragmentation. We want to avoid correlations between the instrument and time-varying conditions in a recipient country. If aid is fragmented, endogeneity will arise only when the total aid budgets of multiple donors respond simultaneously to a recipient's conditions, and to a large degree. This does not seem especially plausible. Total aid budgets are likely to emerge from a medium-run political process, and the hypothesized strong responses would imply greater volatility in total aid budgets than we see in the data. If aid is not fragmented, the most serious problem for the instrument would arise when a donor gives to a small number of recipients, and those recipients receive most of their aid from that donor. There are few such cases in the data. The observed fragmentation suggests that the instrument will be uncorrelated with time-varying recipient conditions in most cases.4 To support this claim, we first look at within-donor fragmentation. We construct Herfindahl-Hirschman (HH) indices for the 29 donors that contribute to the supply-push instrument for at least one aid recipient in our main sample; low values correspond to aid that is distributed across many recipients, or high fragmentation.5 Fig. 1 shows box-plots for the distribution of the HH indices across donors, for each three-year period in our sample. If we take the box-plot for 1971–73, the upper border of the box indicates that 75% of the donors had HH indices below about 0.30, indicating high fragmentation even in the first period. Reading across the figure, the long-run trend is towards greater fragmentation (lower median HH indices). The dots indicate isolated exceptions, which relate to two minor donors and a small number of recipients. Even the exceptions are unlikely to be a problem, because these recipients will typically receive aid from multiple sources. Since the instrument is a weighted average of donor budgets, this should weaken the correlation between the instrument and recipient conditions even when a subset of the donors are highly specialized. But to check this, in the later analysis, we will consider versions of the instrument which exclude specialized donors. The results continue to indicate strong effects of aid on total consumption and net imports. The high degree of within-donor fragmentation supports our approach. It suggests that the total aid budgets of these donors are unlikely to be strongly driven by the domestic conditions of individual recipients. We next examine within-recipient fragmentation. As indicated previously, this will promote identification when the value of the supply-push instrument for a given recipient draws on aid from multiple donors. This is because, even if some donors to a given recipient have total aid budgets which are correlated with that country's domestic conditions, summing over a larger set of donors will tend to weaken the correlation between those conditions and the instrument. Note that, for identification, it will be fragmentation within the instrument at the recipient level that matters, rather than within a recipient's aid. We construct HH indices for each recipient and year, based on the contribution of each donor to that recipient's supply-push instrument in that year. Fig. 2 shows box-plots for the distributions of these within-instrument HH indices across recipients, drawing on the 1099 recipient-period observations in our main sample. The lines in the middle of the boxes show the degree of fragmentation within the instrument for the median recipient: for many recipients, the instrument is relatively fragmented, supporting identification. There are some exceptions where recipients have HH indices close to one, and hence where the instrument places a high weight on a single donor for that particular recipient. These are isolated cases which are unlikely to dominate the variation, and should not threaten identification given the within-donor fragmentation documented above.6 In summary, the analysis suggests that a supply-push approach can be applied to the study of foreign aid. We now turn to a different objection, which is that total aid budgets might be influenced by disaster and emergency relief. But even broadly defined, humanitarian assistance accounts for a small share of global aid flows: for 1995–2013, Qian (2015) finds that it ranged between 5% and 9% of official development aid. Some humanitarian assistance is long-term rather than emergency-related, and several of the major recipients are not in our data set.7 Finally, since emergency relief may be funded by reallocations within existing budgets, and disasters do not appear to have major effects on aid receipts (Qian, 2015), we do not see strong grounds to reject the supply-push approach on this basis. Importantly, allowing for heterogeneous effects of common factors will help to ensure that our supply-push instrument is exogenous. In contrast, conventional panel estimators with time fixed effects assume that common factors, such as global shocks, have exactly the same effects on all the countries in the sample.8 If the data generating process is more complicated, identification could fail, because the instrument might be correlated with the effects of the common factors. The combination of a supply-push instrument and a flexible approach to common factors is new to this paper, and should achieve identification under a wider range of circumstances than previous work. To address this, we filter out common factors using the approach of Pesaran (2006). His paper introduced common correlated effects (CCE) estimators for panel data. This class of estimators proxies for the combined effects of common factors using linear combinations of the cross- section means of the observable variables. To allow the effects of the factors to differ across countries, the combinations are estimated from the data and vary across countries. This is achieved by augmenting the regression with cross-section means of the dependent variable and the explanatory variables, all with country-specific coefficients. Under a condition on the number of linearly independent common factors, their combined effect will be captured by the cross-section means. The method thereby addresses an important class of omitted variables in a way that is straightforward to implement. It can accommodate various forms of cross-section dependence, and simulations suggest that it can perform well even in small samples and when the factors are non-stationary.9 The CCE method has been extended to the case of instrumental variables by Harding and Lamarche (2011), yielding a CCE IV estimator that we describe in Appendix A. This estimator is again easy to implement. The first and second stages of 2SLS are augmented with cross-section means of the observable variables, including the instrument, with country-specific coefficients. This approach has costs and benefits. It builds in robustness, helping to ensure that the instrument is exogenous. However, filtering out the common factors is parameter-intensive and could ask a lot of the data. The key point here is that, although the benefits of CCE IV may be offset by higher standard errors, our estimates are often precise enough to be informative.10 Although we report the results from several methods, we give most emphasis to CCE IV, as the estimator most likely to yield consistent estimates of the parameters of interest. We now discuss how the paper relates to previous work. The instrument we adopt was first used by Van de Sijpe (2010) to study aid and governance, but without allowing for a multifactor error structure.11 Work using other supply-push instruments includes Nunn and Qian (2014) and Werker et al. (2009). The instrument in the latter paper interacts the world oil price with a dummy for Muslim countries, since aid to Muslim aid recipients may be sensitive to the oil price. Werker et al. use this to study the effects of aid on a range of outcomes, including macroeconomic ratios. Their findings tally closely with ours. They find a significant effect of aid on consumption, where the IV estimate is much larger; no evidence that aid leads to higher government consumption; some evidence that aid promotes gross investment, but this is not robust; no evidence that aid affects exports; and strong evidence that aid leads to higher imports. Although the two papers differ in significant respects, the findings on absorption are remarkably similar. Our approach is related to other work on aid using instrumental variables, including Galiani et al. (2017), Jarotschkin and Kraay (2016) and Tavares (2003). The latter paper used the distance between recipients and donors, and whether they share a common border, language or religion, to instrument for aid. In our study, the initial shares in donor budgets can proxy for many possible connections between donors and recipients while remaining agnostic about their sources. Put differently, we infer connections from the data, rather than restricting them to take specific forms. In summary, the combination of a supply-push instrument and the CCE IV estimator is new to this paper. The approach has several benefits. First, by using the total aid budgets of many donors to construct the instrument, we are exploiting the fragmentation of global aid flows to lessen the risk that the instrument is correlated with the conditions of individual aid recipients. This will be the case even if those conditions are driven by country-level trends, for example. Second, we use the CCE IV approach to address a remaining concern, that total donor budgets could be influenced by common factors, such as world economic conditions, that are correlated with conditions in aid recipients. The combination of instrument and estimator allows us to go beyond the natural experiments studied in the literature to date, and study the effects of broadly-defined aid from a wide range of donors.","Our models will be estimated using three-year averages over 1971–2012.12 To construct the synthetic aid measure that we use as an instrument, we need the initial shares of aid recipients in donor budgets. These initial shares will be based on the period 1960–1970. Our aid variable is taken from Table 2a of the OECD Development Assistance Committee (DAC) data tables. We follow Arndt et al. (2010) in our treatment of some missing values: they argue that some apparently missing values in fact correspond to zeroes. In each year, we turn missing recipient-donor-year aid to zero for combinations of recipients that receive aid from at least one donor in that year and donors that disburse aid to at least one recipient in that year. Aid in recipient-year format is found by keeping the entries that list ‘All donors, total’ as a donor. Our focus is on net aid disbursements, and our final sample comprises 88 aid recipients. The dependent variables considered will include household consumption, government consumption, gross capital formation, imports and exports, again relative to GDP.14 Net imports are defined as imports minus exports. In the recipient-year data, before collapsing to three-year averages, observations for these variables are turned to missing whenever at least one of the other components of the GDP identity is missing. This keeps the sample consistent across the different outcomes we consider. In our final data set, the expenditure components sum to total GDP, or very close to GDP, for each country-period observation.15 We exclude countries with populations fewer than 500,000 people in the first period of the sample. These steps leave us with a panel with 1099 observations, based on 88 countries.16","For each dependent variable, we report eight regressions. For reference purposes, we report FE and pooled CCE results that do not instrument for aid. We report estimates for static models, and models that include a lagged dependent variable.17 In each case, the coefficient estimates indicate the effect of the aid-to-GDP ratio on the dependent variable, also measured as a ratio. For example, in the case of net imports, a point estimate of one implies that the ratio of net imports to GDP increases one-for-one with the ratio of aid to GDP, corresponding to full absorption. The standard errors that we report are heteroskedasticity-robust and clustered by country, and we make a small-sample adjustment to take into account the large number of parameters. In experiments, we compared adjusted standard errors to those obtained from a non-parametric block bootstrap, given that the asymptotic distribution of pooled CCE-type estimators is non-standard (Pesaran, 2006) and the asymptotic variance of the CCE IV estimator in Harding and Lamarche (2011) is not known. The bootstrapped standard errors are noticeably larger than the conventional standard errors in some cases, but our main findings obtain under either approach to inference.18 The long-run responses of macroeconomic ratios to aid may differ from those in the short run. The inclusion of a lagged dependent variable is one way to capture this. Whenever the estimated model is dynamic, the long-run effect of aid is estimated using the ratio of the short-run effect to one minus the coefficient of the lagged dependent variable, with a standard error approximated by the delta method. Since the long-run effect is a ratio, the estimates are likely to less precise than the estimates of short-run effects.19 Dynamic models also have the drawback that CCE-type estimators will be consistent under more restrictive assumptions than in the static case.20 With this in mind, our discussion will give more emphasis to static models; the use of three-year averages implies that, even in these models, absorption can extend over several years. Before we turn to the results, note that the effects of instrumenting for aid and allowing for common factors are likely to vary across the dependent variables. Macroeconomic ratios are likely to differ in their sensitivity to particular aid-relevant shocks, and in their sensitivity to common factors. We first study the effects of aid on trade-related variables, starting with net imports. Recall that net imports must increase if aid is absorbed. If the net import share rises one-for-one with the aid share, this should assuage concerns that aid is diverted abroad (capital flight), used to accumulate foreign exchange reserves, or spent on forms of technical assistance that do not have a direct effect on expenditure beyond the donor country. The results are shown in Table 1. In our IV estimates, we cannot reject the hypothesis that aid is fully absorbed domestically. The coefficient on aid is large, significantly different from zero, and not significantly different from unity, both in static models and in the long run derived from dynamic models. Note the contrast with the upper row of estimates: in the absence of an instrument, the evidence that aid is fully absorbed is weaker. We can also investigate whether there are symptoms of aid-driven Dutch Disease. An increase in domestic expenditure will often fall partly on non-traded goods, increasing both their relative price and the costs facing the export sector.21 Tables 2 and 3 show the effects of aid intensity on import and export shares respectively. The results for the import share are similar to those for net imports, with the exception of the static model estimated by FE IV. A strong positive effect of aid is found in the two CCE IV regressions in particular. In contrast, we do not find a clear-cut effect of aid on the export share. In the FE IV estimates of a static model, aid has a negative effect on the export share which is significant at the 10% level, but this finding is not robust to alternative models and estimators. In the dynamic FE IV estimates, and the two sets of CCE IV estimates, we cannot reject the hypothesis that aid has no effect on the export share. This does not rule out Dutch Disease – that finding would require zeroes estimated with greater precision – but nor is there robust evidence that aid adversely affects exports.22 Next, we study how aid is absorbed. Note that, since aid leads to an increase in net imports, the GDP identity implies that the sum of household consumption, government consumption and total investment must have also increased. The question is whether we can reliably identify the components of GDP which respond most strongly to aid. We first look at the effect of aid on total consumption. We define this as the sum of household and government consumption (C + G). While household and government consumption are distinct, there are sectors such as education and health where the distinction is somewhat artificial for welfare purposes, given a mix of public and private provision. The results are shown in Table 4 and suggest that aid has a large positive effect on total consumption. The difference made by instrumental variables can be seen clearly, by comparing the upper row of estimates with the lower row. Compared to the FE and CCE estimates, the point estimates from IV estimators suggest larger effects of aid on total consumption. We should avoid over-interpreting this, because the differences are not statistically significant. Nevertheless, it is easy to see how this pattern could arise. Donors may respond to country-specific adverse shocks by allocating countries more aid, and this form of endogeneity will weaken the correlation between aid and total consumption. By instrumenting aid we alleviate this source of bias, and find larger effects of aid. For the effects of aid on household consumption, the estimates are generally similar to those we find for total consumption, but less precise. These results are shown in Table 5. The use of an instrument again increases the estimated effect of aid. We find much less evidence that aid influences government consumption, as Table 6 shows. These results are similar to those found by Werker et al. (2009, Table 2) using a different instrument. Aid has often been characterized as primarily government-to-government transfers. Our finding that a substantial fraction of aid is reflected in higher household consumption, but not in higher government consumption, may be surprising. One mechanism could be lower taxes: increased aid to governments may not be used to increase government purchases, but to reduce taxes (Kimbrough, 1986). Alternatively, recipient governments may use aid to finance transfers for political ends, as in Adam and O’Connell (1999), Boone (1996) and Hodler and Raschky (2014), among others. Finally, some aid is given in ways which bypass domestic governments, such as off-budget aid projects and support for NGOs.23 In these cases, household consumption is where the effect of aid is most likely to be manifested in the national accounts. Finally, we look at the effect of aid on the investment rate, in Table 7. When aid is instrumented, the estimated long-run effect is significant at the 5% (FE IV) or 10% (CCE IV) level. But to anticipate our later discussion, the investment results are less robust than the consumption results. Overall, our findings are in line with Boone (1996), who finds that aid translates mainly into consumption rather than investment and growth. Clemens et al. (2012) present stronger evidence that aid has modest effects on investment and growth, but note that their results are ‘not incompatible’ with Boone's suggestion that aid is often consumed rather than invested. The 2SLS estimates of Werker et al. (2009, Table 2) suggest that aid has stronger effects on consumption than investment. They find no evidence that aid affects growth when using four-year averages, which tallies with the lack of robustness of an investment effect in this paper. Considering the magnitudes of the effects, the specification allows these to be interpreted easily. In Table 1, for example, a one percentage point increase in aid's share of GDP leads to a 0.772 percentage point increase in net imports as a share of GDP (in FE IV estimates) or a 1.085 increase (in CCE IV estimates). Although these two estimates are individually statistically significant, the difference between them may not be significant. To investigate this, we use a bootstrap-based test described in Section B.4 of the online appendix. For the results in Tables 1–7, we typically cannot reject the null that the CCE IV point estimate is not significantly different from the FE IV estimate, given their high standard errors. In choosing between them, there is a trade-off between robustness and efficiency. The CCE IV estimator should be consistent under more general conditions than FE IV, and hence more robust, but will usually lead to higher standard errors. Readers may legitimately differ in which set of estimates they prefer. We have not yet discussed the strength of our instrument. The tables report the first-stage F-statistic as a guide, indicating the significance of the single excluded instrument.24 This approach has been widely used, but the conventional Stock and Yogo (2005) benchmarks for first-stage F-statistics do not apply directly to panel data models. In keeping with other papers, our application of the F-statistic in panel 2SLS is best seen as heuristic. In most cases, and especially in the static CCE IV regressions, the first-stage robust F-statistic is reasonably high, and the Kleibergen-Paap LM test always rejects the null of under-identification at the 5% level.","We now consider several alternative models and estimators. These will tend to increase robustness – in the sense of reducing likely biases – at the expense of reduced efficiency. Our main conclusions continue to find support, even when we make adjustments to the instrument that weaken its explanatory power in the first stage. The estimates are summarized in Table 8, where row 1 shows the main results from Tables 1–7 for ease of comparison. Rows 2–9 then correspond to the robustness tests listed in the notes to the table, and that we discuss below. In each row, we report the estimated effects of aid in static and dynamic models, and the first-stage F statistics. In our main results, we followed much of the literature and excluded countries with small populations. If we include these countries, the instrument becomes weaker in the first stage of 2SLS: see row 2 of the table.25 A potential explanation is that, for aid recipients which account for small and volatile shares of donor budgets, the share of a budget at an initial date may be uninformative about that recipient's long-term exposure to changes in that budget. Hence, we would expect our supply-push instrument to have less explanatory power for aid to smaller countries. But despite the weakening of the instrument, the results are qualitatively unchanged. This is also true when we exclude aid observations for colonies prior to independence, as discussed in Section B.5 of the online appendix. Also in that appendix, we examine whether the estimated effects of aid are driven by distinct subgroups of countries, namely economies where natural resources play a major role, and countries with unusually strong or weak institutions. Estimates based on subsamples continue to suggest that aid increases consumption and net imports, while the evidence that aid affects investment varies across samples. As before, aid appears to increase imports, while there is no evidence of a negative effect on exports. Responses to aid may take time to emerge, as Clemens et al. (2012) argue. In Table 8, we estimate models with lagged aid rather than current aid (row 3) and using six-year averages of aid and the instrument (row 4).26 These point to stronger effects of aid on investment, but we interpret this cautiously. As we discuss later, our findings on investment are more sensitive to modelling choices than our findings on consumption and net imports. Section B.5 of the online appendix reports on some additional variations on these robustness tests. We now turn to potential criticisms of our instrument. Serial correlation in country-specific conditions might undermine exogeneity (Card, 2001). Another concern is that the relevance of the instrument could decline over time. Our IV strategy relies on the idea that shares in donor budgets in 1960–70 are informative about exposure to later changes in total donor budgets. If strategic or economic connections between countries evolve, the instrument may have less explanatory power for aid in later periods. We address these concerns as follows. Row 5 in Table 8 repeats the main analysis but excludes the first period. Row 6 uses an instrument based on initial shares calculated over the period 1960–73 and an estimation sample that starts with 1974–76. Row 7 excludes the first period from this sample. Section B.5 of the online appendix reports some further variations, including ones with the initial shares calculated over 1960–76. We might expect instrument strength to weaken when early time periods are excluded, and this is what we find. The estimated second-stage coefficients are fairly stable, however, when dropping early time periods. In rows 5–7, we continue to find that aid is absorbed via higher consumption and higher imports, without much effect on exports. In experiments that drop up to three periods, the point estimates are quite stable, but the results for total consumption become a little less precise once we drop three periods; see Section B.5 of the online appendix. We also note that, if serially-correlated shocks were a major problem in our static models, we would have expected a greater contrast between static and dynamic models in Tables 1–7. It could be argued that our instrument makes use of too many donors. By considering all DAC donors, we have included some whose budgets could be dominated by a few recipients, which risks endogeneity. To investigate this, row 8 in Table 8 shows results using an instrument based only on less specialized donors, whose HH index never exceeds 0.25.27 The instrument remains informative and the results are in line with our main findings, but yield larger point estimates for the effects of aid on consumption and on net imports. We cannot reject full absorption achieved through higher consumption.28 Next, we consider whether donors have been sorting themselves across recipient countries in a way that could undermine identification. We examine this using an approach developed by Greenstone et al. (2015). We implement two versions of their test and, in both cases, find little evidence that sorting has taken place. But there are some grounds for caution over the applicability of these tests in our setting, and Section B.3 of the online appendix examines these issues in greater detail. We also investigate the possibility of outliers. Given that we use 2SLS, outliers could arise in the first stage or the second stage. Some of our robustness checks give rise to large first-stage F statistics, which may be a warning sign of outliers. To address this, we use the robust instrumental variable estimator of Cohen Freue et al. (2013), after partialling out fixed effects and cross-section means. As they discuss, robust parameter estimates can then be used to identify multivariate outliers. Across our dependent variables, seven countries regularly give rise to one or more outlying country-period observations: Burundi, the Central African Republic, Chad, the Democratic Republic of Congo, Jordan, Madagascar, and Mauritania. The results when we exclude them are shown in Row 9 of Table 8. In line with our main findings, the instrument retains explanatory power, and we cannot reject strong effects of aid on net imports, household consumption and total consumption. The effects on net imports are sufficiently strong that we cannot reject the null hypothesis of full absorption. In the dynamic model, the increase in net imports appears to arise from higher imports rather than lower exports. The estimated effect of aid on investment is too imprecise to draw conclusions, and this remains the case in models (not reported) that allow for delayed effects. We now consider an alternative estimation method. Thus far, we have emphasized estimators that use a within transformation. For the static models, another way to eliminate country- specific effects is to first difference the model.29 A comparison of the two approaches should be informative about the validity of our assumptions, and help to address potential concerns. In particular, our dependent and independent variables both contain nominal GDP in the denominator, as does the instrument. This would be problematic if some function of nominal GDP cannot legitimately be excluded from the models we estimate: in that case, the instrument would be correlated with the error term. But, given the persistence of GDP, first differencing would weaken the correlation between the instrument and the error term in differences. Hence, we investigate the results from this approach, and compare them to our earlier findings. The first-differenced results suggest that our earlier findings are not spurious. Nevertheless, a sceptical observer might still be worried about a ‘false positive’, because the variables in our regressions take the form of ratios to nominal GDP. In principle, this could lead to an observed relationship between two ratios even when their numerators are unrelated.30 One reason to doubt this interpretation is the pattern of the results. We consistently find that the ratios of consumption and net imports to GDP are increasing in aid intensity, in line with theoretical predictions, while the effects of aid on the other ratios are weaker, again in line with theory. A convincing story based on ‘false positives' would have to explain why spurious correlations emerge for some ratios and not others, and why the magnitudes of the estimated effects remain plausible. But to investigate this further, Section B.5 of the online appendix reports on some additional analysis, including models with the dependent variable in logarithms and the logarithm of nominal GDP included as an explanatory variable. The results on absorption and its channels are in line with our main findings. In summary, we have carried out a range of robustness tests. These include demanding versions of CCE and CCE IV estimators with many additional parameters, and go beyond those typically used in the literature. It is not surprising that the effects sometimes become imprecise, but the point estimates are quite stable, and rarely change sign. Across a range of models and approaches, we continue to find that aid is absorbed in line with theoretical predictions, and with effect sizes that have plausible magnitudes. More precise estimates are likely to require longer spans of data or the use of additional instruments, such as those of Galiani et al. (2017) or Jarotschkin and Kraay (2016). In the meantime, we note that our findings are consistent with Werker et al. (2009), while adopting a new approach to identification and filtering out common factors. Moreover, in contrast to previous studies based on distinct natural experiments, the results we present relate to a broad concept of aid, for the full set of DAC donors.","Using cross-country data to study aid is fraught with difficulties, and yet some research questions are hard to answer any other way. This paper has aimed to make progress on two fronts. First, we have introduced a new instrument that can be used to identify persistent changes in aid, and combined it with a panel time series estimator that filters out unobserved common factors, even when their effects differ across countries. This approach is an advance on much of the existing literature. Second, we use the instrument to investigate the effects of aid on domestic absorption, consumption and investment, and whether aid is associated with Dutch Disease. The evidence suggests that aid is absorbed at least partially, and we cannot reject the hypothesis of full absorption, in which aid increases net imports one-for-one. Absorption seems to arise mainly via increased imports, and we find no evidence that aid lowers exports through Dutch Disease effects. We also investigate the relationship between aid and the separate components of domestic expenditure. Our estimates suggest that aid is absorbed primarily by increased consumption rather than investment. Although we do sometimes find a significant effect of aid on investment, it seems less robust than the effect on consumption. Overall, the supply-push instrument appears to work well. For many of the dependent variables studied, instrumenting for aid has a substantial effect on the results. The instrument is strong enough to generate some informative findings, and as future years of data become available, the prospects for robust and precise estimates should improve further. The instrument could have many possible applications in the future study of aid."],["In a novel experimental design, we study public good games with dynamic interdependencies, where each agent's wealth at the end of period t serves as her endowment in t + 1. In this setting, growth and inequality arise endogenously allowing us to address new questions regarding their interplay and effect on cooperation. We find that amounts contributed are increasing over time even in the absence of punishment possibilities. Variation in wealth is substantial with the richest groups earning more than ten times what the poorest groups earn. Introducing the possibility of punishment does not increase wealth and in some cases even decreases it. In the presence of a punishment option, inequality in early periods is strongly negatively correlated with group income in later periods, highlighting negative interaction effects between endogenous inequality and punishment. --------------------------------------------------------------------------------","Social dilemmas, where collective and private interests are in conflict, abound in economic and social life. Public good games have been used across disciplines as the standard tool to study a wide array of social dilemma situations. These include joint ventures (Grossman and Shapiro, 1986), R&D cooperation (Cozzi, 1999; Kamien et al., 1992), political action funds of special interest groups or parties (Dawes et al., 1986), multilateral foreign aid and effort provision in work teams (Ostrom, 1990; Hamilton et al., 2003; Tirole, 1986). But also pricing or market sharing agreements by firms (Green and Porter, 1984) as well as many economic activities in the family (Becker, 1981) can be thought of as instances of cooperation that can be modeled with public good games. One feature that is common to many of these examples is that there are dynamic interdependencies: not only will the same set of people interact again, but previous outcomes affect future endowments (both in terms of the stock of physical and social capital). In this paper, we present a novel experimental design that captures such dynamic interdependencies. Our design builds on what has become the workhorse model to study public good provision in experiments (see e.g. Isaac et al. (1984), Andreoni (1995) or Fischbacher and Gächter (2010) among many others): participants are matched in fixed groups of four people to play the public good game for 10 or 15 periods. As in most other public good experiments, we focus on the most challenging social dilemma situations, where the unique subgame perfect Nash equilibrium prescribes zero contributions by all group members, but where efficiency requires group members to contribute their entire endowment. We also conduct experiments where each group member can punish other group members by reducing their first stage earnings at a cost to themselves (Ostrom et al., 1992; Fehr and Gächter, 2000; Andreoni et al., 2003). The key difference to previous research using this “standard” design is that each participant's wealth at the end of a period constitutes their endowment for the next period, whereas in the “standard design” endowments are allocated exogenously and tend to be the same in each round.1 In our design, endowments are created endogenously, which leads to dynamic interdependencies. We focus on two important implications of introducing these interdependencies. First, if overall contributions today are high, then there will be higher wealth in the next period (growth). Second, heterogeneity in contributions today creates inequality in endowments in the following period. Growth and inequality can interact with the possibility of punishment in different ways. The threat of punishment can lead to higher growth if it induces higher contributions, but punishment executed on the outcome path can induce a multiplier effect which can hamper growth severely. Maybe more interestingly, punishment can interact with inequality in non-trivial ways. In particular, rich group members can be largely “immune” to punishment by poorer group members, if the punishment that poor group members can afford is too small relative to the richest group member's wealth. As all endowment can be used for punishment, rich group members, on the other hand, might be able to punish others harshly at a relatively low cost to themselves. This asymmetry of punishment possibilities translates wealth inequality into inequality in power to punish. In addition, in unequal groups richer group members will typically be free-riders implying that punishment power could be in the “wrong hands”. This raises the question of whether punishment will be as effective in increasing contributions and group income as it has been in settings without these dynamic interdependencies. Our main findings can be summarized as follows. Even in the absence of punishment contribution and wealth levels display a strictly increasing trend over time. In terms of the realized potential for growth and the level of inequality, there is a lot of variation across groups. Individual earnings range between 2 Euros and 241 Euros. The Gini coefficient assumes the full range between 0 (equal wealth of all group members) and 1 (one group member appropriates the entire wealth) in the experiment. Punishment (or the possibility thereof) does not increase wealth. This is true in both the 10 and 15-period variations despite the fact that people tend to contribute more in the treatment with punishment in the 10-period variations. We find evidence for two mechanisms behind this result: (i) in groups where inequality is high (above median) there is more anti-social than pro-social punishment, i.e. shirkers punish contributors more than vice versa and (ii) much of this punishment happens in early periods implying that resources are taken away exponentially.2 While the possibility of punishment does not increase wealth and in some cases even strictly decreases it, it also does not increase inequality on average. This is true despite the inequality-increasing presence of anti-social punishment. Analysis of data from Herrmann et al. (2008) shows that, in a comparable standard setting, punishment increases both wealth and inequality. In terms of the relationship between inequality and growth, we find that inequality in period 2 is strongly negatively correlated with wealth in period 10 in the treatment with punishment possibilities. In particular, a 1% increase in inequality in period 2 leads to a ≈ 0.5% decrease in wealth in period 10. Inequality and growth are positively related across groups with below median wealth and negatively related across groups with above median wealth. In interpreting these results, it is important to note how the setting we introduce differs from the standard public good game described above. There are two main differences: (i) there is no consumption until the last round in our setting, i.e. one's entire wealth can be reinvested at the end of the period and (ii) endowments are endogenous, i.e. determined by previous outcomes. R + D cooperation often displays these features, but also the evolution of societies could be viewed under this lens. In the standard setting, by contrast, there is full consumption, i.e. no wealth can be reinvested and endowments are exogenous (and stationary). Volunteering, e.g. at a food bank, seems a good example falling into this category. Other types of volunteering, such as in natural conservation and archiving, are examples involving stationary exogenous endowments, but no or little consumption. The case of full consumption with endogenous endowments describes the one-shot game. Finally, note that many applications, such as infrastructure investments, multilateral foreign aid or pricing agreements will fall somewhere in between these extremes with some, but not full consumption and with partially endogenous endowments. Literature review ~~~~~~~~~~~~~~~~~ To our knowledge, our experiment is the first to study public good provision with dynamic interdependencies and endogenously arising asymmetric punishment possibilities.3 As such, it contributes to studies of public goods with dynamic interdependencies more broadly. These include Battaglini et al. (2016) who study the Markov perfect equilibrium dynamics in the provision of a durable public good over time where there is consumption in each period. They find evidence of significant under-provision relative to the interior equilibrium. Duffy et al. (2007) studied threshold public good games with multiple contribution rounds, where, theoretically, “completion equilibria” (with positive contributions) do exist. They find that, as in the standard setting, contributions do decline over time (see also Croson and Marks, 1998 among others). Noussair and Soo (2008) and Cadigan et al. (2011) study dynamic public good settings where the current return from the public good depends on past contributions. Other studies link public good games over time via explicit reputation mechanisms (e.g. Milinski et al., 2002). Rockenbach and Wolff (2017) have recently studied a setting where games are linked via endowments. They use a non-linear exchange rate, though, which effectively eliminates the possibility of exponential growth and contains inequality. Consequently, their results are more similar to those obtained in the standard setting. Our results also contribute to research on the impact of inequality on public good provision. Most existing literature has studied the effects of exogenous income inequality. Chan et al. (1996) experimentally test a prediction by Bergstrom et al. (1986), where public good provision increases with inequality in the income distribution in an equilibrium with positive contributions. They find that group behaviour conforms with the theoretical prediction. Other authors have found that exogenous income inequality decreases contributions (van Dijk et al., 2002; Ostrom et al., 1994) or found no effect (Chan et al., 1999). Reuben and Riedl (2013) find that without punishment there is no effect of income inequality on contributions, while with punishment participants contribute proportionally to their endowments, leading to increased contributions. All these papers deal with one-shot games and exogenously imposed inequality. Sadrieh and Verbon (2006) study exogenous inequality in the Bergstrom et al. (1986) setting with the possibility of growth. They find that exogenous variations of inequality are mostly neutral to growth.4 The latter result contrasts with our finding that endogenous inequality is negatively related to growth in the presence of punishment. One possible reason for the difference between these results is that, as discussed above, with endogenous inequality the power to punish tends to lie “in the wrong hands” (those of free-riders). Our paper also contributes to the literature on the standard setting of repeated public good games surveyed in Chaudhury (2011). One insight emerging from this literature is that societies can maintain high levels of contributions and wealth if there is a threat of punishment to free-riders (Ostrom et al., 1992; Fehr and Gächter, 2000; Fehr and Gächter, 2002; Andreoni et al., 2003).5 While some studies have found that punishment can lead to lower payoffs if the horizon of the game is short (Egas and Riedl, 2008; Dreber et al., 2008), under a long horizon punishment in the standard setting has been found to be strictly beneficial (Gächter et al., 2008). Punishment does not increase wealth in our setting neither under a 10-period nor a 15-period horizon. We do not study longer horizons, but since many groups in the treatment with punishment destroy (almost) all wealth by period 10, these groups will be unable to recover even with a very long horizon. Some studies of the standard setting investigate how an increased marginal per capita rate of return affects contributions (Isaac and Walker, 1988; Goeree et al., 2002). There is some relation of these studies to our setting, since the possibilities of exponential growth provides increased incentives to contribute. Finally, the dynamic setting we study also relates to other dynamic games studied in economics. In the common pool resource game (CPR, Ostrom et al., 1994) players extract from a resource in each period with the non-extracted part growing at a fixed rate. There are various differences between this setting and our dynamic public good game (apart from the frame). One is that whatever is contributed in our public good game is mechanically shared among all participants. How the benefits of non-extraction are shared depend on players' strategies. On the other hand, what is withheld in our public good game in one period can be contributed in the next. What is extracted in the CPR game, however, cannot be reverted to the pool in future periods. As a consequence of these various differences, both games have different equilibria. The CPR game has an interior stationary Markov equilibrium where players extract too much relative to the efficient benchmark (Mailath and Samuelson, 2005). By contrast, in our dynamic public good game the only equilibrium has zero contributions in all rounds (Section 3). Herr et al. (1997) study two variations of the CPR game, one where extraction by one player increases the cost of extraction by others in the current period only, and one where extraction increases the costs for all future periods. While both versions have interior equilibria, the latter type of externality should worsen the CPR problem, i.e. lead to lower efficiency. Experimental behaviour is in line with equilibrium benchmarks. Botelho et al. (2014) study a game that lies in between a classic CPR and Centipede game, in that they study a CPR game which ends as soon as the resource stock falls below a certain level. They show that increasing uncertainty about that level increases extraction and hence the inefficiency. Battaglini et al. (2016) discussed above can be viewed as a public good study within the CPR paradigm. This paper is organized as follows. In Section 2, we present the experimental design. Section 3 discusses the theoretical predictions and summarizes our research questions. Section 4 contains our main results. In Section 5, we discuss the mechanisms underlying these results. Section 6 concludes. An Online Appendix contains screenshots, experimental instructions and the questionnaire, proofs of the theoretical predictions and several additional results, tables and figures.","In our experiment, participants play a sequence of dynamically interdependent public good games in two main treatments: (1) treatment NOPUNISH in which participants only play the public good game; (2) treatment PUNISH where, after each period, participants can also subtract tokens from other members of their group at a cost. We describe these treatments in turn. Treatment NOPUNISH ~~~~~~~~~~~~~~~~~~ Therefore, in period t + 1, the number of tokens that each participant can invest depends on the choices of all group members in previous periods (endogenous endowments). This is what makes our set-up different from the standard setting where the amount of tokens before each period is fixed and the earnings from previous periods cannot be used for investment into the public good (exogenous stationary endowments). At the end of each period participants observe information about endowments and contributions of all group members (Fig. A.1 in Online Appendix A). Additional treatments ~~~~~~~~~~~~~~~~~~~~~ As we outlined above, the dynamic interdependencies in this setting create two types of effects: (i) they create the possibility for endogenous growth and (ii) they create endogenous inequality and hence asymmetries in the power to punish others. To better understand the impact of these two forces we ran additional treatments, where we artificially eliminate growth (treatments NOPUNISH-NOGROWTH and PUNISH-NOGROWTH) but keep endogenous inequality or where we artificially eliminate inequality but keep growth (treatments NOPUNISH-NOINEQUALITY and PUNISH-NOINEQUALITY). We will discuss these treatments in more detail in Section 5. Length variations ~~~~~~~~~~~~~~~~~ In addition to the punishment variation, we added a second treatment variation which changes the number of repetitions of the public good game. We conducted sessions with 10 periods and sessions with 15 periods to understand how a longer horizon might affect behaviour and hence get some insights into how participants might perceive this environment strategically.7 Due to the (potentially) exponential increase of earnings over time in our setting, we limited the long horizon to 15 periods.8 Participants were informed about all these details of the design at the beginning of the experiment. Other details ~~~~~~~~~~~~~ Participants answered a sequence of questions at the end of the experiment (see Online Appendix B). We report summary statistics on these characteristics in Tables 14 and 15 in Online Appendix F, where analysis of the questionnaire data can be found. In all treatments, the amount that participants received at the end of the experiment was equal to a €2 show up fee plus the amount of tokens after the last period converted into Euros. 1 token was equal to 0.05 Euros. All experiments were run at the BEE-Lab at Maastricht University and used z-tree (Fischbacher, 2007). Table 8 in Appendix D shows the exact order of sessions for the main treatments. In total, 656 participants took part in our experiment: 152 in treatment NOPUNISH (38 groups), 144 in treatment PUNISH (36 groups) and the remainder in one of the additional treatments. Table 7 in Appendix D summarizes the treatment structure and the number of independent observations, participants and sessions for each treatment. Average earnings were 11.94 Euros with a minimum of 2 Euros and a maximum of 241.65 Euros. No other sessions apart from those reported here were conducted and there were no pilot studies.","In this section, we briefly discuss the theoretical predictions and then present our research questions. We focus on the standard solution concepts for dynamic games with observable actions, i.e., Nash equilibrium (NE) and subgame perfect equilibrium (SPE) under the standard textbook assumptions of self-interest and common knowledge of rationality. First, observe that the games that describe our treatments are not repeated games, as the set of available actions at each subgame depends on the moves that have been already realized. Therefore, unlike in the standard setting, we cannot directly apply the wide range of well-known results on finitely repeated games. Nevertheless, it turns out that SPE leads to very similar predictions as in the standard setting. On the equilibrium path all players contribute 0 in all periods, and moreover in the punishment treatment no player ever punishes. The purpose of stating these results is not to argue that they will be good predictions of behaviour, but rather to give the reader a sense of the incentive properties of the public good game with dynamic interdependencies. In this section, we will state the result informally. Formal statements and proofs can be found in Online Appendix C. Consider the public good game with growth as defined above. The unique SPE of both the game without punishment and with punishment is such that all players contribute 0 throughout the entire game (both on and off the equilibrium path). As is the case with the standard setting, positive contributions can be sustained under different assumptions. Particularly reputational models that invoke the existence of behavioural types such as e.g. Kreps et al. (1982) have the potential to induce positive contributions in either setting, but also to drive a wedge between the two settings in terms of the extent of cooperation found. More precisely, following Kreps et al. (1982), let us assume that there are two types of agents: (i) “standard” rational and self interested types and (ii) conditional cooperators who start out by contributing 10 tokens and then match at each time t the minimal contribution in their group at t − 1. We assume that all agents are of type (i) and that each agent has a minimal doubt that other agents may be of type (ii), i.e. conditional cooperators. More precisely we denote by μ the probability with which agents believe that all other agents are conditional cooperators. We also denote by T the length of the game, which in our experiment is either 10 or 15 periods. We can then state the following result. The proof of this result can be found in Online Appendix C. Because of the incentives provided by exponential growth, even a small doubt that some agents may be conditional cooperators suffices to induce positive contributions on the equilibrium path. In the standard setting, by contrast, such doubts have to be substantial (μ > 0.19) to even get one period of positive contributions. In fact, if 0.03 < μ < 0.19 and T = 10, then positive contributions can only be sustained in the setting with growth. For any given μ, more periods of contributions will be observed in the setting with growth. And, across the periods where contributions are positive, the amount contributed is constant in the standard setting and increasing in the setting with growth (for a related result in Centipede games, see McKelvey and Palfrey, 1992). We next summarize our research questions. First and foremost, we are interested in whether participants contribute positive amounts and in how much wealth is created in this game. Apart from wealth creation, we are also interested in the amount of inequality created both within and across groups. Of particular interest is the relation between inequality and growth. The sign of the relation between inequality and growth of societies as well as the causal link between the two has been at the center of a debate in macroeconomics and development (Barro, 2000; Forbes, 2000; Persson and Tabellini, 1991). Hence, while Q1 asks whether societies can see positive contributions and increasing wealth over time, our next question Q2 asks how much inequality is generated in this process and how inequality affects growth. Our setting allows us to address this question both within and across groups. In particular, we ask whether “rich” or “poor” groups will be more unequal and how initial inequality affects growth and hence final wealth. Last, we are interested in the effects of punishment in our setting. Punishment has been shown to be effective in the standard setting in securing high contributions and wealth (Gächter et al., 2008). In our setting, however, punishment does destroy resources which could otherwise be used productively in the following period. Punishment in the dynamic setting can have additional adverse effects that operate via endogenously created inequality. Because earnings are carried over across periods, free-riders in our setting will have more resources to punish others than contributors. Hence, the possibility of punishment could strengthen existing inequalities because it makes shirking individuals more powerful. This in turn could undermine the effectiveness of punishment.","This section comprises our main results. Section 4.1 focuses on Question 1, i.e. on how much participants contribute and how much wealth is generated in this setting. Section 4.2 focuses on Question 2, i.e. how much inequality is endogenously created. We study the effect of punishment (Q3) within each of these subsections. In Section 5, we discuss additional results and possible mechanisms behind our results. Provision of the public good and wealth creation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We start by discussing contributions. Panel (a) in Fig. 1 shows the average amount of tokens participants contributed over time. Contributions are clearly non-zero and are increasing over time in all treatments. Even without punishment participants contribute about 10 tokens in the first period and then steadily increase this amount over time. Contributions flatten out towards the end of the experiment and there is even a statistically significant drop in period 15 in the long horizon game. Note that increasing contributions over time imply that participants have increasing endowments over time. Hence, increasing contributions do not necessarily imply that participants contribute increasing shares of their endowments. Panel (b) in Fig. 1 shows the share of overall endowments contributed over time in all main treatments. In NOPUNISH, participants contribute around 55% of their endowment in period 1. This amount steadily decreases between periods 2 and 8 and is roughly constant afterwards at a level of about 35% of endowments. In the 15-period games, shares follow a similar pattern until period 10, then start to decrease again afterwards. Panel (a) in Fig. 1 presents a stark contrast to typical 10-period public good games in the standard setting, where contributions are decreasing over time (see e.g. Figure 3 in Fehr and Gächter, 2000).9 It should be kept in mind, though, that in the standard setting with stationary exogenous endowments, the dynamics of contributions over time is the same irrespective of whether it is measured in absolute terms or as a share of endowment. In Fehr and Gächter (2000), for example, contributions equal around 50–60% of endowments in Period 1 and decrease to around 10% of endowments by Period 10. Panel (b) in Fig. 1 shows that in our setting contributions stabilize at a higher level (≈ 35%) in the 10-period games even if measured as share of endowments contributed. One possible explanation is that the dynamic incentives with the possibility of exponential growth have a similar effect to that of increasing MPCR's in the standard setting (Isaac and Walker, 1988; Goeree et al., 2002).10 This higher level of shares contributed seems not to be sustainable in the longer run, though. By period 15, a similar share of endowments is contributed as in the standard setting after 10 periods (10–20 %). Fig. 2 shows the dynamics of wealth over time. Panel (a) focuses on all groups, panel (b) on those with above median wealth in Period 10 (“successful” groups) and panel (c) on those with below median wealth in Period 10 (“unsuccessful” groups). Average wealth is increasing across periods (see coefficients β1 and β4 in Tables 9 and 10 in Appendix D) and is substantially above 80 once period 10 is reached (Table 1). Not all groups achieve growth, however. There are several groups where wealth does not rise (substantially) above 80. Figs. E.1–E.2 in Online Appendix E illustrate wealth (and Gini coefficients) over time in different matching groups. Amounts contributed are positive and increasing over time even without punishment. Wealth is growing over time, but there is large variation across groups with the richest group earning more than 10 times more than the poorest after 10 periods and more than 55 times more after 15 periods. Effects of punishment on contributions and wealth We next consider the effect of punishment on contributions and wealth. Absolute contributions end up higher in PUNISH compared to NOPUNISH (Panel (a) in Fig. 1), but they start to differ only from period 7 onwards in the 10-period games and in the 15-period games they are higher only in the very last period. In terms of shares contributed (Panel (b) of Fig. 1), in PUNISH, contributions are roughly constant across all periods at a level of about 60%, though there is a slight decrease towards the end of the 15 period games. This is above the share at which groups stabilize in treatments without punishment. However, groups in PUNISH are also poorer. Median wealth is higher in NOPUNISH compared to PUNISH (Table 1) and the difference in mean ranks is statistically significant in the 10-period games according to a one-sided ranksum test (p < 0.0001 for all periods; p = 0.0928 for t = 10 data only). To assess the statistical significance of differences in means, we run OLS regressions where we regress wealth on a treatment dummy for PUNISH (Table 2). These regressions show that differences in means are only significant for below median groups in the 10-period games which earn on average 213 tokens less in PUNISH compared to NOPUNISH. Several of these groups, in fact, end up with zero income in PUNISH. In the 15-period games, differences are statistically significant across all groups with mean wealth in period 15 being about 790 tokens lower in PUNISH compared to NOPUNISH. The fact that punishment seems more harmful the longer the horizon stands in contrast to results obtained in the standard setting by e.g. Gächter et al. (2008). Columns (1) and (2) of Table 9 in Appendix D assess differences in time trends using a linear (column (1)) and square (column (2)) polynomial in period. While wealth is increasing in both treatments (β1,β1 + β3), the linear trend in column (1) is no different across treatments (β3). The square polynomial does reveal differences in time trends, though. While wealth is initially increasing at a lower rate in PUNISH (β3), it increases at a faster rate in later period (β5). Fig. 2 illustrates these results. It shows that wealth is lower in PUNISH compared to NOPUNISH across all periods except for the last three periods in the 10-period games. Overall, these results indicate that the threat of punishment is not needed to sustain high contributions in public good games with growth and does not increase wealth. Instead, the dynamic evolution of the public good drives contributions and wealth. This is in stark contrast to what we know from other social dilemma games, where punishment has been shown to be successful in both raising contributions (in Fehr and Gächter’s (2000) study on the standard setting by around 600%) as well as wealth in later periods (in Fehr and Gächter’s (2000) study from period 4 onwards (Result 8 in Fehr and Gächter, 2000). In Section 5, we discuss in more detail why punishment is less effective in our setting. There we also analyze data from an experiment by Herrmann et al. (2008) conducted in the standard setting and show that in this case punishment leads to unambiguously higher wealth. The possibility of punishment does not increase wealth. In the long horizon games (15 periods), average wealth is even lower in PUNISH compared to NOPUNISH. Inequality ~~~~~~~~~~ In this subsection, we focus on the amount of inequality created endogenously in our setting. As measure of inequality we use the Gini coefficient as defined in Deaton (1997). The smallest possible value the Gini coefficient takes is zero (if all four group members own one fourth of the wealth) and the largest possible value it takes is one (if one group member holds the entire wealth). Table 1 shows some summary statistics regarding the Gini coefficient. The period 10 Gini coefficient ranges between 0 and 0.43 in NOPUNISH with a median of 0.22 and assumes the full range between 0 and 1 in PUNISH with a median of 0.03. The period 15 Gini coefficient ranges between 0.03 and 0.49 in NOPUNISH and between 0 and 0.67 in PUNISH. Fig. 3 illustrates the dynamics of the Gini coefficient over time. The figure shows that in NOPUNISH inequality is sharply increasing from the initial value of zero and then increases more slowly between periods 2 and 10 to reach a level of ≈ 0.2 in period 10 (β2 and β4 in column (2) of Tables 11 and 12 in Appendix D). Interestingly, these patterns almost perfectly mimic the trends in inequality identified in a recent paper by Nishi et al. (2015) (see their Fig. 2) who study the effect of visibility of wealth (and inequality) in a cooperation game played on a network. A longer horizon seems to lead to slower growth in inequality, but by period 15 Gini coefficients are not statistically different from those reached in period 10 of the shorter horizon games. Panel (b) shows groups with above median wealth, which display lower levels of inequality than those with below median wealth illustrated in Panel (c) (ranksum test p < 0.001 all periods, p = 0.1096 period 10, p > 0.1 period 15). Mean Gini coefficients are increasing over time reaching ≈ 0.2 in the last period. There is substantial variation in Gini coefficients across groups, particularly in PUNISH where the Gini assumes the full range between 0 and 1 in the 10-period games. Effects of punishment on inequality We next study the effect of punishment on inequality. Mean Gini coefficients are similar across treatments and there are no statistically significant differences in mean Gini coefficients between NOPUNISH and PUNISH in period 10 or 15, respectively (see Table 4). By contrast, analysis conducted with data from a standard public good game in Section 5.4 shows that in this setting punishment leads to more inequality. In fact, in our case the median Gini coefficient tends to be even higher in NOPUNISH compared to PUNISH (Table 3). The difference in mean ranks is statistically significant (one-sided ranksum test p = 0.0395) only in the 10-period games. There is also more variation in PUNISH (one-sided variance ratio test on period 10 data only, p < 0.0001; period 15 data: p < 0.0001) where the Gini coefficient assumes the full range between 0 and 1 in the 10-period games and between 0 and 0.67 in the 15-period games. Fig. 3 illustrates these results as well as differences in time trends. Across all groups (Panel (a)) the dynamics of inequality seem similar across treatments with somewhat more volatility in PUNISH. Columns (1)–(2) of Table 11 in Appendix D show that there are few statistically significant differences in these time trends. The dynamics of inequality over time seem also quite similar in above median groups, where the difference is mostly one of means (Panel (b)). In below median groups, however, time trends are quite different across treatments (Panel (c)). In NOPUNISH, inequality is steadily increasing over time at a slow rate, while in PUNISH, the Gini coefficient seems to follow a different pattern. After increasing initially, it decreases sharply around periods 4–7 and then starts to increase again. One possible interpretation is that this pattern reflects cycles of reciprocal punishment. Depending on who punishes (shirkers or high contributors) inequality increases or decreases. We discuss differences in punishment behaviour between below and above median groups in more detail in Section 5.3. Mean inequality (Gini) is not statistically different in PUNISH compared to NOPUNISH in period 10 (15). There is more across- group variation in inequality in PUNISH.","In this section, we discuss some of the potential mechanisms underlying our main results. Section 5.1 discusses results on the relationship between growth and inequality. In Section 5.2, we will try to tease apart the effect of exponential growth opportunities from the effect of endogenously created inequality. In Section 5.3, we provide some additional results on punishment, which should help us get a deeper understanding of why punishment is less effective here than in the standard setting. Finally, in Section 5.4, we use data from experiments conducted by Herrmann et al. (2008) to study wealth and inequality in the standard setting. The relation between growth and inequality ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We study the relationship between growth and inequality across and within groups starting with across-group inequality. In NOPUNISH, the two variables are essentially uncorrelated. The Spearman correlation coefficient is −0.1840 in above median groups and −0.0630 in below median groups, neither of which is statistically different from zero. In PUNISH, by contrast, there is a substantial and statistically significant relationship between these two variables. In above median groups the correlation is negative (Spearman correlation coefficient ρ = −0.6091**): higher wealth comes with lower inequality. In fact, most groups that achieve very high levels of wealth in period 10, have a Gini coefficient of almost zero. For below median groups, the correlation is positive (ρ = 0.4683*): higher wealth comes with higher inequality. As a consequence inequality is highest in groups with intermediate levels of wealth in PUNISH. Fig. D.1 in Appendix D illustrates the correlation between period 10 wealth and Gini coefficients for both treatments. The same patterns hold in the 15-period games, but, possibly due to lower power, there is no statistical significance. These differences between above and below median groups appear already as early as in the second period of the game. For NOPUNISH, both successful (above median wealth in period 10) and unsuccessful (below median wealth in period 10) groups have a Gini coefficient of 0.13 on average in period 2. In the 15-period games, these numbers are 0.13 and 0.14, respectively. In PUNISH, on the other hand, there are substantial differences. Groups that are eventually successful have a Gini coefficient of about 0.11 in period 2 (0.10 in the 15-period games), while unsuccessful groups have a Gini coefficient of 0.27 (0.25 in 15-period games), more than twice as high. Hence, successful and unsuccessful groups already differ after the first round of the public good game. This observation motivates us to study the extent of path dependency within groups. Table 5 presents evidence in this respect. It shows the correlation between wealth in period 10 (period 15) and the wealth or the Gini coefficient in previous periods. Path dependence is evident. Early period wealth is strongly correlated with late period wealth. The Spearman correlation coefficient between periods 2 and 10 income is 0.48*** in NOPUNISH and even 0.82*** in PUNISH. Maybe more interestingly, also early period inequality is highly detrimental to final wealth, but only in PUNISH. The correlation between inequality in period 2 and wealth in period 10 is −0.47***, which is substantial. In NOPUNISH, by contrast, early period inequality (in periods 2–4) is not negatively correlated with wealth. In this treatment the negative correlation appears only from period 5 onwards. The fact that there is no negative correlation between periods 7, 8, and 9 inequalities and period 10 wealth in PUNISH is due to the fact that by then, several groups have zero wealth and inequality. Dropping these groups restores the negative correlation. Comparing 15 period wealth and early period inequality (bottom panel of Table 5) shows similar patterns. One question about these findings is whether they just reflect a stable distribution of “contribution types” or whether there is something more fundamental to it in the sense that the same people are more likely to end up with a much lower wealth if initial inequality is high. The fact that the negative correlation between early period inequality and wealth is only observed in PUNISH, suggests that this is not just a mechanical effect of having different distributions of “contribution types” across groups. Additional support for this view can be derived from our post-experimental questionnaire. Table 16 in Online Appendix F shows the average amount in Euros that participants decide to donate to Medics without Borders at the end of the experiment. We find that participants from above median groups do not contribute more on average than those from groups with below median wealth. This is despite the fact that participants from groups with above median wealth earn 178 tokens on average in period 10 (189 in treatment PUNISH), while those from groups with below median wealth earn only 56 tokens (23 tokens) on average in period 10. This suggests that participants in groups with above median wealth are not per se more altruistic than others. We also elicited 14 other personality characteristics as well as a measure of risk aversion in the questionnaire. Tables 17–18 in Online Appendix F show that none of them is able to explain the variation in wealth or inequality that we observe. 1. In PUNISH, wealth and Gini coefficient are positively correlated for poor groups (below median wealth) and negatively correlated for rich groups (above median wealth) in the 10-period games. These correlations are weak and not statistically different from zero in NOPUNISH. 2. (a) Early period wealth is positively correlated with eventual wealth in all treatments. (b) Early period inequality is negatively correlated with final wealth only in PUNISH. Taken together, both the findings on across- as well as within-group correlation point to a detrimental role of inequality if there are punishment possibilities. One of the reasons, hence, why punishment is not as effective in this setting seems to be that people react strongly to inequality. The findings are also indicative as to why we observe such substantial variation across groups both in terms of wealth and inequality. Groups in which initial behaviour leads to high inequality seem to get locked into a path of punishment and counter-punishment (interpreted as experimental equivalents of conflict) that eventually leads to a destruction of all wealth. We will see more evidence of such behaviour in Section 5.3. In fact, there is a strand of literature, where economic historians point out the importance of institutional lock in with institutions broadly understood as both formal constraints, such as rules and laws and informal constraints, such as norms of behaviour, conventions and codes of conduct (North, 1994). Our study provides an example of how a society (group) can get locked into dysfunctional behavioural norms. Before we conclude this section, let us point to potential links to other literatures. The sign of the relation between inequality and growth as well as the causal link between the two has been at the center of a debate in macroeconomics and development (see e.g. Barro, 2000; Forbes, 2000; Persson and Tabellini, 1991, among many others). Most of these authors find either a negative relation or no significant relation at all. In our context, the relation depends on the wealth of the group. For very poor groups in our data, inequality and wealth are positively related, while they are negatively related for richer groups. This is reminiscent of the famous Kuznets curve (Kuznets, 1955) which claims an inverse U-shaped relationship between growth and inequality. The connection between our setting and the fate of countries is too loose to draw any conclusions. Our results suggest, however, that there may be interesting links, between the level of social capital (cooperation, trust), the level of inequality and the level of growth of societies.11 Eliminating inequality and growth possibilities ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ This means that each participant's endowment ranges between 0 and 80 in each round and the sum of endowments in a group equals 80 in each period. This 10 period game hence, can be viewed as a public good game played in period 10, where the endowment each participant has is determined by the game played in periods 1–9. The unique SPNE is zero contributions in each period as in our main treatments. Note that this structure, while derived from our treatment NOPUNISH via a single change creates a situation where everyone contributing no longer Pareto dominates nobody contributing in periods 1–9. Consequently, participants may have different motives for contributing in these treatments compared to our main treatments and the results should be interpreted with this in mind. We do the analogous normalization for the punishment version (PUNISH-NOGROWTH). We had 116 participants (29 groups) in treatment NOPUNISH-NOGROWTH and 92 participants (23 groups) in treatment PUNISH-NOGROWTH. Table 13 (column (1)) in Appendix D shows the results of regression where we regress normalized contributions on treatment dummies. The baseline is NOPUNISH. To ensure a fair comparison, we also normalize contributions in NOPUNISH and PUNISH. Normalized contributions in these treatments are computed by multiplying the share of actual endowment contributed with normalized endowments. Maybe unsurprisingly, given the absence of even social incentives to contribute across periods 1–9 of the games without growth, contributions are substantially lower and close to zero. The second column in Table 13 (Appendix D) shows that inequality leads to lower contributions in both the treatments with and without punishment. Eliminating inequality increases mean contributions by 26 tokens without punishment and by 42 tokens with punishment. Furthermore, if inequality is eliminated then mean contributions are ≈ 24 tokens higher with punishment (PUNISH-NOINEQUALITY) than without (NOPUNISH-NOINEQUALITY), a difference that is statistically significant (χ2 < 0.001). It seems that endogenous inequality in endowments with the associated inequality in the power to punish undermines the effectiveness of punishment. Since we exogenously remove inequality in the treatments discussed here, the results from this section support a causal interpretation of the effect of inequality on wealth. Anatomy of punishment ~~~~~~~~~~~~~~~~~~~~~ Both Sections 5.1 and 5.2 have pointed to a negative role of inequality for contributions and wealth particularly in treatment PUNISH. To understand why this is the case it is helpful to study punishment patterns more closely. We first compare above and below median groups in treatment PUNISH to understand which patterns of punishment lead to low wealth in this treatment and are hence crucial for the ineffectiveness of punishment in this setting. Fig. 4 reveals an interesting pattern in this regard. It shows the amounts of tokens participants use to punish over time. In groups with above median wealth, the absolute amounts used to punish tend to remain stable or increase over time (OLS coefficient: 1.402* (10 periods); 0.098 (15 periods)),13 while they tend to decrease in below median groups (OLS coefficient: −0.830*** (10 periods); −0.332*** (15 periods)). Particularly striking is the fact that, in terms of amounts, the major difference between above and below median groups seems to lie in how much they punish in the first two periods of the game (two-sided ranksum test p < 0.0001). In successful groups, there is also an interesting peak in punishment one period before the game ends. This suggests that some participants may tolerate some degrees of free-riding while they wait to punish others harshly at the end of the game. This seems intuitive because of the detrimental effect that punishment can have on growth. While in below median groups, most punishment happens in the beginning of the game, above median groups punish at the end.14 We next ask under which conditions punishment is “pro-social” and when it is “anti-social”. Herrmann et al. (2008) have found that anti-social punishment strongly undermines successful cooperation in the standard setting, whereas pro-social punishment fosters cooperation. Punishment by player i is pro-social if i punishes a player who has contributed a lower share of his endowment to the public good than i herself. Punishment by player i is anti- social if i punishes a player who has contributed a higher share of his endowment to the public good than i herself. Table 6 shows that there is more pro-social punishment than anti-social punishment (one-sided ranksum test, p < 0.0001). Anti-social punishment is higher if inequality is high (pro-social: Spearman ρ = 0.0155 (ρ = 0.1231***, 15 period games); anti-social: ρ = 0.0777** (ρ = 0.1515***, 15 period games)). Anti-social punishment is not significantly related to wealth (ρ = −0.0109; ρ = 0.1184 in 15 period games), but there is more pro-social punishment in high-wealth compared to low-wealth groups (ρ = 0.1371***; ρ = 0.2302*** in 15 period games). Wealth and inequality in the standard setting ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Before we conclude, we have a brief look at data from the standard setting. We focus on wealth and inequality, as those are not commonly reported measures in the standard setting. In particular, we use data from experiments that Herrmann et al. (2008) conducted in Bonn. Bonn and Maastricht are at only 120 km driving distance and a large share of students at both Universities comes from the lower middle Rhine and upper lower Rhine area (between Koblenz and Duesseldorf) in Germany. This should make subject pools approximately comparable. Just as us, Herrmann et al. (2008) conducted 10 period games with an endowment of 20 tokens (in their case exogenously given each period). They also conducted NOPUNISH and PUNISH treatments with the 3:1 technology that we employ. The return on contributing in the one-shot game was 0.375 in our experiments and at 0.4 slightly higher in Herrmann et al. (2008). Fig. 5 shows wealth and inequality (Gini coefficient) over time in Herrmann et al. ’s ( 2008 ) experiments. We computed wealth by adding income of all group members in each period and subtracting 80 in each period 2,…,10. If all group members contribute zero in each period (and do not punish) wealth thus computed is 80 in each period, just as in our setting. If all group members contribute their full endowment in each period (and do not punish) maximal wealth in period 10 is 3075 in our setting and 1280 in Herrmann et al. (2008). Fig. 5 shows that both wealth and inequality are substantially higher under the punishment condition in Herrmann et al. ’s ( 2008 ) experiments (t-test period 10 sample means: p < 0.01). The effect of punishment, hence, differs from what we observe in the dynamic setting, where punishment, if at all, decreases wealth and inequality. Note also, that we do not observe the cyclical pattern in inequality identified in the PUNISH treatments in our setting. In summary, Section 5.1 demonstrates a negative correlation between early period inequality and wealth in period 10 in treatment PUNISH. The negative impact of inequality on contributions and wealth is confirmed with some caveats in Section 5.2 where we find that artificially removing inequality increases contributions and wealth both with and without punishment, but with a particularly strong effect in the punishment treatments. Section 5.3 provides some insights into why inequality is so detrimental in the punishment condition. There are two main effects: (i) in groups where inequality is high (above median) there is more anti-social than pro-social punishment, implying that shirkers punish contributors more than vice versa and (ii) much of this punishment happens in early periods in unsuccessful groups implying that resources are taken away exponentially. All these channels contribute to the negative impact of punishment possibilities on wealth creation in the dynamic setting. The absence of inequality in endowments in the standard setting seems to make wealth inequality less salient. This can potentially explain why punishment does not trigger the adverse affects identified in the dynamic setting in the standard setting (Section 5.4).","We studied public good games with dynamic interdependencies, where each agent's wealth at the end of a period serves as her endowment in the following period. We found that contributions are increasing over time even in the absence of punishment possibilities. The possibility of punishment does not increase wealth. These results suggest that in settings with a strong dynamic component societies achieve cooperation via the incentives provided by the dynamic evolution of the public good and not so much via the threat of punishment. Across groups, inequality in early periods is strongly negatively correlated with wealth in later periods. These results show that people are able to establish persistent cooperation in a setting that shares one key feature with many real-life interactions: past behaviour matters for future endowments. They also point to the limits of punishment in securing high contributions and wealth. The results highlight the importance of incorporating the dynamic aspects, present in many real-life interactions, explicitly in experimental designs. In this paper we have done so within the context of cooperation and public good provision. We should emphasize, though, that we view this design as complementary to the standard design where contributions occur from stationary exogenous budgets. Both settings have natural applications outside the lab. Public goods are often created by providing effort, as e.g. in the case of volunteering. In these cases, clearly, endowments cannot easily increase over time. In many other settings, such as those discussed in the Introduction 1, provision of public goods creates wealth from which future public goods can be provided. The evolution of societies might be viewed through this lens. We have seen that the two settings can yield quite different conclusions. In the standard setting with stationary budgets and full consumption, punishment has shown to be very effective at raising contributions and wealth (in the medium run). In the dynamic setting studied here, the effectiveness of punishment seems much more limited. It is not successful in raising wealth in the 10-period games and is even detrimental to wealth in the games with longer horizon. Future research should aim at getting a deeper understanding of the mechanisms behind these differences and adapt this setting to other contexts where dynamic interdependencies are likely to play an important role. Studying intermediate settings, with some but not full consumption and only partially endogenous endowments should also be of interest for future research."],["We study the implications of human capital hedging for international portfolio choice. First, we document that, at the household level, the degree of home country bias in equity holdings is increasing in the labor income to financial wealth ratio. Second, we show that a heterogeneous agent model in which households face short selling constraints and labor income risk, calibrated to match both micro and macro labor income and asset returns data, can both rationalize this finding and generate a large aggregate home country bias in portfolio holdings. Third, we find that the empirical evidence supporting the belief that the human capital hedging motive should skew domestic portfolios toward foreign assets, is driven by an econometric misspecification rejected by the data. --------------------------------------------------------------------------------","International finance theory emphasizes the effectiveness of global portfolio diversification strategies for cash-flow stabilization and consumption risk sharing.2 However, the empirical evidence on international portfolio holdings favors a widespread lack of diversification across countries and a systematic bias toward home country assets (see, e.g., Coeurdacier and Rey, 2013 for a recent survey). This discrepancy between theoretical predictions and observed portfolio constitutes the international diversification puzzle (see, e.g., French and Poterba, 1991). Moreover, albeit the degree of home bias has been reducing during the last few decades, it remains a first order characteristic of portfolio holdings.3 In the major industrialized countries, roughly two- thirds of gross domestic product goes to labor and only one-third to capital. Thus, human wealth likely constitutes about two-thirds of total wealth, suggesting that if investors attempt to hedge against adverse fluctuations in returns to human capital when making financial investment decisions, the mere size of human capital in total wealth makes its potential impact on portfolio holdings self-evident. Based on this observation, several contributions have argued that when the role of human capital is explicitly taken into account, the observed home country bias in portfolio holdings becomes harder to rationalize. The argument, originally formalized by Brainard and Tobin (1992) with a stylized example, works as follows: if returns to human capital are more correlated with the domestic stock market than with the foreign ones, labor income risk can be more effectively hedged with foreign assets than with domestic ones, and equilibrium portfolio holdings should be skewed toward foreign securities.4 As emphasized by Cole (1988), “this result is disturbing, given the apparent lack of international diversification that we observe.” However, as first suggested in Bottazzi et al. (1996), human capital hedging could also lead toward home country bias in portfolio holdings. For instance, the correlation between domestic return on physical and human capital can be lowered by idiosyncratic shocks that lead to a redistribution of total income between capital and labor. In the presence of these rent shifting shocks, foreign assets become a less attractive hedge for labor income risk—especially if total factor productivity shocks are highly correlated internationally. If the size of the rent shifting shocks is large enough, a situation in which domestic assets are the best hedge against human capital risk arises, therefore leading to home country bias in portfolio holdings. In this paper, we ask whether the human capital hedging motive is likely to have a sizeable effect on optimal portfolio choice, and what its implications are for the international diversification puzzle. Moreover, we propose a rationalization of the home country bias, based on a setting of endogenous portfolio formation and incomplete markets, that not only can rationalize aggregate portfolio holdings, but also the variable degree of home country bias in households' portfolios. In particular, Fig. 1 depicts a novel (to the best of our knowledge) finding about households' equity home bias.5 Panel (a) depicts the (locally weighted regression of the) share of foreign assets in U.S. household portfolios as a function of the household financial wealth to labor income ratio. Since (labor income) flows and stocks (of human capital) are cointegrated, Panel (a) shows that there is a systematic relationship between household specific home country bias and the household specific financial to human capital ratio: the degree of home country bias monotonically decreases as the human capital component of household total wealth becomes smaller relative to the household financial wealth. That is, in micro data, when the human capital hedging motive is more prominent relative to the financial wealth hedging motive, household portfolios show a higher degree of home country bias. Moreover, panel (b) of Fig. 1, that depicts the (locally weighted regression of the) number of stocks in U.S. household portfolios as a function of the household financial wealth to labor income ratio, shows that when the households' human capital wealth is relative larger than the financial wealth, the household portfolio will tend to be overall less diversified. Our paper provides a rationalization of both of these findings, as well as of the aggregate home country bias, and also shows that the canonical intuition that human capital should skew portfolio holdings toward foreign assets, and the related supporting empirical evidence, are both very fragile. In particular, we offer two main contributions. First, using novel estimates of the correlations of human capital and stock market return innovations, we calibrate an incomplete market model in which agents face both idiosyncratic and aggregate labor income risk, as well as borrowing constraints.6 The model is also calibrated to match the microeconomic (following Gourinchas and Parker, 2002) characteristics of the U.S. labor income and track the distribution of the asset wealth to labor income ratios observed in the Panel Study of Income Dynamics (PSID). The main findings of this calibration exercise are that a) investors that enter the stock market with a low level of liquid (i.e. financial) wealth to labor income ratio will initially specialize in domestic assets and, b) only as the level of asset wealth to labor income ratio increases do agents start diversifying their portfolios internationally by progressively adding different assets to their holdings, c) as a consequence, the aggregate portfolio of U.S. investors shows a large degree of home bias. What drives these results? Households face large human capital risk, but this is mostly of the idiosyncratic type—hence underestimated in a homogeneous agent setting. Moreover, in the presence of liquidity constraints, agents cannot borrow to construct an optimally diversified portfolio. Therefore, when their level of liquid wealth to labor income ratio is sufficiently high and they enter the stock market, agents try to minimize the overall wealth risk, investing first in the asset that has the lowest degree of correlation with labor income innovations—and, as discussed below, this assets is, in the data, the domestic stock. Only when the ratio of liquid wealth to labor income is sufficiently high, and the labor income risk hedging motive becomes less important relative to the financial risk hedging one, do agents start investing in foreign assets and diversifying their portfolios internationally. Since the distribution of liquid wealth to labor income is (in the data as in the model) concentrated in the region of low liquid wealth to labor income ratios, the resulting aggregate portfolio is heavily skewed toward domestic assets. Note that, in the absence of market frictions and idiosyncratic risk, the estimated and calibrated correlations of labor income and returns innovations, being very small, would have almost no effect on the optimal portfolios. Moreover, since in our model the aggregate home country bias depends on both the household optimal investment policy functions and the aggregate distribution of liquid wealth to labor income, a trend of increasing concentration of financial wealth (as documented by Piketty, 2014), and/or a negative trend in the labor share of income (as documented by Karabarbounis and Neiman, 2013), would both generate a negative trend in the degree of home country bias as found by Coeurdacier and Rey (2013).7 Since the pattern of correlation of innovations to labor income and returns plays an important role in our calibration, our second contribution hinges upon the identification of a common misspecification that has affected the previous empirical literature, and the provision of novel estimates that are not affected by this issue. In particular, we show that the seminal empirical result of Baxter and Jermann (1997) that, in the presence of a human capital hedging motive, investors should short sell the domestic capital stock—implying that “the international diversification puzzle is worse than you think”—is largely due to an econometric misspecification rejected by the data: the assumption that there are neither cross-country shocks to human and physical capital payoffs, nor common long run trends. We show that, once this restriction is relaxed, the effective degree of technological and economic integration becomes evident, therefore reducing the opportunities to hedge human capital risk by investing in foreign assets. Moreover, we also show that there is substantial uncertainty attached to the estimation of aggregate physical capital returns via the canonical Campbell and Shiller (1988) cum vector autoregression (VAR) approach. This feature of the data provides a rationalization for the apparently contradictory empirical evidence on the correlation between returns to human and physical capital found in the previous literature (that typically has not reported confidence bands for the estimated correlations): Lustig and Nieuwerburgh (2008) find a strong negative correlation between domestic returns to physical and human capital in U.S. data; Bottazzi et al. (1996) find such a correlation to be negative in all the countries they consider but the United States, where they find it to be strongly positive; Baxter and Jermann (1997) find this correlation to be positive and very close to one for all the countries they consider (United States, United Kingdom, Germany, and Japan). Nevertheless, we find that by restricting the set of assets available to hedge human capital to include only publicly traded stocks—what we consider as the relevant case for most households—much sharper estimates can be obtained, and these are the estimates used in our calibration exercise discussed above. We find that in this case human capital hedging can help explain the home country bias in portfolio holdings—since domestic returns to human capital tend to be systematically more correlated with foreign stock markets—but, we show, given the small magnitudes of the estimated correlation, the effect is quantitatively very small in a frictionless complete market setting.8 Nevertheless, as discussed above, this same small correlations have very large aggregate effects in an incomplete market settings in which agents face both idiosyncratic and aggregate labor income risk. Note that, overall, our result that domestic capital markets constitute a good hedge for domestic human capital risk are in line with the findings of a large empirical literature. Palacios-Huerta (2001) finds that if human capital is included in the definition of wealth, gains from international financial diversification for a mean-variance investor appear to be smaller than previously reported. Lustig and Nieuwerburgh (2008) find that innovations in current and future human wealth returns are negatively correlated with innovations in current and future domestic financial asset returns. Abowd (1989) finds a large and negative correlation between unexpected union wage changes and unexpected changes in the stock value of the firm. Davis and Willen (2000), using data from the PSID to construct synthetic cohorts, find that the correlation between domestic labor income shocks and returns on the S&P 500 is substantially negative for some categories.9 Moreover, they find that for six out of the eight sex-education groups considered in their study, a long position on the worker's own industry represents a good hedge for labor income risk. The empirical works of Gali (1999), Rotemberg (2003), and Francis and Ramey (2004) also document a negative correlation between labor hours and productivity conditioning on productivity shocks. Coeurdacier and Rey (2013), conditioning on exchange rate movements, find that wages and dividends growth rates comove negatively for all the countries they consider (also Coeurdacier et al., 2013 ; Heathcote and Perri, 2013 provide similar empirical evidence). Theoretically, we show that a situation in which labor income innovations are more correlated with the domestic payout to capital than the foreign ones—as we find in the data—is likely to arise once the degree of international economic integration observed in the data is properly taken into account. In particular, we show that very small redistributive shocks (shocks with a variance that is equal to as little as 6%–11% of the output variance), are enough to make the domestic equity market the best hedge for human capital risk. The analysis presented in this paper is part of the literature that has attempted to explain home bias as a hedge against non-tradable risks.10 Moreover, the potential rationalization of the international diversification puzzle we document in this paper should be interpreted as complementary, rather than alternative, to the ones based on transaction and information frictions (e.g. Van Nieuwerburgh and Veldkamp, 2009; Bhamra et al., 2014), nominal stickiness (e.g., Engel and Matsumoto, 2009), non-traded goods (e.g., Heathcote and Perri, 2013) and more broadly the role of real exchange rate fluctuations (e.g., Coeurdacier, 2009; Kollmann, 2006; Baxter et al., 1998). The remainder of the paper is organized as follows. Section 2 presents a calibrated model of human capital risk hedging in which households face both idiosyncratic and aggregate labor income risk, as well as liquidity constraints. Section 3 presents the empirical approach undertaken to measure factor returns and rationalizes the difference in results between our findings and the previous empirical literature. The final section outlines the conclusions of the paper, while a detailed data description, as well as additional results and robustness checks, are reported in the Appendix.","In this section we rely on numerical methods to compute the equilibrium outcome of a model that directly takes into account that i) most of the human capital risk faced by households is idiosyncratic in nature, and ii) households' optimal portfolio choice is influenced by liquidity constraints. The simple incomplete markets model presented below is a generalization of Heaton and Lucas (1997) to a multiple asset context and of Michaelides (2003), and builds upon the household income process estimated by Gourinchas and Parker (2002). Model setup and calibration ~~~~~~~~~~~~~~~~~~~~~~~~~~~ Given the assumptions on the primitives and the calibrated values, these conditions hold and there exists a unique set of optimal policies satisfying the Euler equations. To avoid the curse of dimensionality of numerical solutions and, most importantly, in order to have sufficiently long time series for the estimation of the variance covariance matrix of labor income innovations and asset returns using the unrestricted VAR approach presented in Section 3 below we focus on four countries: the United States (as domestic country), the United Kingdom, Japan, and Germany. As we show in Section 3 below not restricting the VAR representation to have country specific block exogeneity is both required by the data, and needed in order to uncover the true degree of hedging potential via international diversification.12 Since the U.S. domestic risky asset has enjoyed both the lowest variance and the highest Sharpe ratio compared to the other countries considered, and this pushes the optimal portfolio to be skewed toward the domestic stock, we calibrate all the countries as having the same mean return and Sharpe ratio as the United States. A summary of the calibrated preference and labor income process parameters are reported in Table 1. The crucial element in calibrating the model is the covariance structure of asset returns and innovations to the aggregate labor income process. We measure capital returns using broad stock market indexes and calibrate their covariance using the time series sample analogous. The calibration of the covariance structure of aggregate labor income shocks and stock market returns is summarized by the correlations reported in the first four columns of Table 2. These are based on the estimation approach discussed extensively in Section 3 below where we show that the (different) estimates obtained in the previous literature are due to misspecification. The crucial element in Table 2 is that the correlation between U.S. labor income innovations (fourth column) with the domestic stock market is marginally smaller than the ones with foreign stock markets returns (expressed in dollar terms). Note that these correlations are all small in magnitude and, as shown in Table A5, would have a very small effect on the optimal portfolio choice in a complete markets setting. As a benchmark, the last column of Table 2 reports the implied optimal portfolio shares of the domestic portfolio absent any human capital hedging motive and shows that, according to the estimated covariance structure of returns, the share of U.S. assets in the U.S. domestic portfolio would be about 25% in the absence of aggregate labor income risk. Investors' optimal policy rules and portfolio choice ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Having calibrated the model, we can estimate the optimal policy function by standard numerical dynamic programming techniques (see, e.g., Carroll, 1992 ; Haliassos and Michaelides, 2002) to compute the optimal consumption and asset holding rules. Since the time t optimal policy rules depend both on the normalized cash-on-hand and on the last labor income shock (εt −1), we numerically integrate out this last variable to have policy rules as a function of the cash-on-hand only,13 obtaining the investment rules {b(x),sd(x),si(x)}. Moreover, from Eq. (5) we can obtain the optimal consumption rule c (x). Optimal policy rules are plotted, as a function of normalized cash-on-hand, in Fig. 2. Not surprisingly, the optimal consumption policy rule has the same shape as in the buffer stock saving literature, with consumption being equal to cash-on-hand (no saving region) until a target level of cash-on-hand is reached and saving starts taking place. Once the saving region is reached, the consumers specialize in stocks, disregarding bonds. This result, well known in the literature, was originally obtained by Heaton and Lucas (1997) in a domestic portfolio choice settings, and it reflects the implication of the large equity premium for the optimal portfolio choice. More interestingly, when the consumer enters the saving region, she initially invests only in the domestic stocks and only gradually diversifies her portfolio internationally as the level of cash-on-hand increases. This happens for three reasons. First, only a small buffer stock saving is needed for the agent to protect herself from future labor income shocks. Second, when entering the saving region, the agent prefers to invest in the assets that have the smallest correlation with labor income shocks, in order not to increase her overall level of risk correlated with income. This is due to the fact that, when entering the investment region, almost the entirety of the agent's wealth is in the form of human capital. Hence, for relatively low levels of cash-on-hand, the human capital hedging motive dominates the portfolio diversification motive. As a consequence, the order in which the agents start investing in the different stock markets closely match the inverse rank of the correlations between labor income innovations and asset returns. Third, only for very high levels of liquid wealth to labor income ratio (high x) does the financial portfolio diversification motive become more important than the labor income hedging one, and the agent starts diversifying fully her portfolio. This is due to the fact that, as x increases, so does the non-human capital component of the household wealth, therefore reducing the human capital hedging motive. Comparing this result with the empirical distribution of cash-on-hand in the PSID data set, less than 1% percent of the population should be investing positive amounts in all four of the assets considered. Moreover, given the positive correlation between normalized cash-on-hand and asset wealth observed in the data, the results imply that only the richest households will be diversifying their portfolio internationally, coherently with the empirical evidence on households' portfolio holdings at the micro level (see, e.g., Jappelli et al., 2001). Using the estimated policy functions, we can compute the optimal portfolio shares as a function of cash-on-hand. These optimal shares are reported in Fig. 3. The figure shows a large bias toward domestic assets in all the relevant ranges of standardized cash on hand, implying that more than 99% of the households should have an asset portfolio strongly biased toward domestic assets. Compared with the optimal share of domestic assets in the market portfolios without aggregate labor income risk (25% in Table 2), this represents a home country bias of individual portfolios that ranges from 75% to 19%. Even investors in the top 1% of the distribution of cash-on-hand observed in the data would have, on average, more than 50% of their asset wealth invested in domestic stocks. Interestingly, this large home bias is generated by extremely small differences in the correlations between labor income shocks and market returns across countries, and a very small aggregate labor income risk component. Moreover, as shown by counterfactual calibration results,14 this effect is mostly driven by the ordering, rather than the magnitudes, of the correlations between labor income innovations and stock market returns. This implies that small shocks that lower the correlation between aggregate labor income innovations and market returns at the country level can generate, in the presence of short selling constraints and buffer-stock saving behavior, a very large degree of domestic bias in portfolio holdings. Implications for the aggregate portfolio ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ This subsection derives the implications of the optimal investment rules, obtained in the previous subsection, for the aggregate portfolio of U.S. investors. The standardized cash- on-hand in Eq. (6) follows a renewal process and can be shown to have an associated invariant distribution,15 and this can be used to compute the implied aggregate portfolio of U.S. investors. Moreover, given the estimated policy functions, the aggregate portfolio can also be computed using the observed empirical distribution of cash-on-hand. Second, we can alternatively draw random initial levels of x to reproduce the initial heterogeneity in wealth among agents, and then simulate dynamically the evolution of normalized cash-on- hand over time, generating what we refer to as the dynamic distribution of the model. We perform both procedures since the first one requires fixing ex-ante the relevant range of x while the second one instead determines the relevant range autonomously, therefore providing a robustness check of the construction of the model ergodic distribution. Fig. 4 reports the distribution of normalized cash-on-hand implied by the model and the observed distribution of normalized cash-on-hand in the PSID data.16 The model seems to reproduce fairly well the location of the mode and the shape of the right tail of the empirical distribution, but the model distribution is much less concentrated than the data around the boundary between the saving and the no saving zone, implying a higher participation rate in the market than what is observed in the PSID data, probably due to the absence of stock market entry costs in the setup of the model. With these distributions at hand, we can compute the implied aggregate portfolio shares of U.S. investors. The first column of Table 3 reports, as a benchmark comparison, the CAPM market portfolio implied by the calibrated covariance structure of returns in the absence of labor income risk. The implied aggregate portfolio shares of the model are reported in the second column. There is a dramatic effect of labor income risk on the aggregate portfolio with about 95% of the market portfolio invested in domestic assets. Moreover, the relative investments in foreign stocks are strongly affected, with a reduction of the portfolio shares in individual foreign stocks moving from the 18%–36% range to the 0%–4% range. Since agents with different levels of normalized cash-on-hand are likely to have different amounts of wealth invested in the stock market, the simple computation of the aggregate portfolio reported in column two of Table 3 could be a poor approximation of the aggregate portfolio. To address this issue, the third column of Table 3 weights the model distribution by the contribution to the aggregate portfolio of agents having different levels of cash-on-hand. This weighting of the distribution also corrects for the fact that the model implies a higher degree of market participation than what is observed in the data. The weights are constructed from the PSID data and are proportional to the total stock market holdings of households belonging to each category of normalized cash-on-hand. This weighting somehow reduces the degree of home bias relative to column two, but still delivers a portfolio share of domestic stocks of about 75%, implying that hedging human capital increases the portfolio share of domestic stocks by as much as 50% and decreases the portfolio shares of German, Japanese, and U.K. stocks by, respectively, 15%, 19% and 17%. The last two columns of Table 3 show that our main result also holds if we compute the aggregate portfolios using the empirical (again, for PSID data), rather than the model implied, distribution of cash-on-hand, with (column five) and without (column four) weighting. The aggregate portfolio shares implied by the empirical distribution are, in both cases, quite similar to the ones obtained by weighting the model distribution (column three) and carry the same message: the human capital hedging motive generates a very large home country bias, with an increase of the portfolio shares of domestic assets between 36% and 50%. But what is the key mechanism delivering a large home bias generated by the model in Table 3? The driving force of our results is that small differences in the correlation of aggregate labor income innovations and market returns, in the presence of short-selling constraints, lead to a gradual international diversification of investors' portfolio as their level of normalized cash-on-hand increases. In the presence of liquidity constraints, agents cannot borrow to construct an optimally diversified portfolio. Therefore, when their level of liquid wealth to labor income ratio is sufficiently high and they enter the stock market, agents try to minimize the overall wealth risk, investing first in the assets that have the lower degree of correlation with labor income. Only when the ratio of liquid wealth to labor income is sufficiently high, and the labor income risk hedging motive becomes less important relative to the financial risk hedging motive, do agents start diversifying their portfolios. Note that this is exactly the pattern found in the Survey of Consumer Finance data, and reported in Fig. 1. Since the distribution of liquid wealth to labor income is—in the data as in the model—concentrated in the region of low liquid wealth to labor income ratios, the resulting aggregate portfolio is heavily skewed toward the asset with the lowest correlation with aggregate labor income shocks. Therefore, the human capital hedging motive, once market frictions and idiosyncratic labor income risk are taken into account, is likely to explain a large fraction of the home country bias in several countries. The above results imply that domestic shocks that lead to a redistribution of total income between capital and labor, therefore lowering the correlation between return on physical and human capital, are likely to skew portfolio holdings toward domestic assets. In Appendix A.3 we show that very small redistributive shocks (shocks with a variance that is equal to as little as 6%–11% of the output variance), can indeed make labor income innovation more correlated with foreign, rather than domestic, market returns innovation. Many kinds of shocks are expected to have an effect on the income distribution that can rationalize the correlations observed in the data and used in our calibration. Common examples are political business cycles and changes in the bargaining power of unions relative to firms. Among others, the works of Bertola (1993) and Alesina and Rodrik (1994) suggest that changes in the time patterns of capital and labor returns may be the endogenous outcome of majority voting. Santa-Clara and Valkanov (2003) find that in the United States the average excess returns on the stock market are significantly higher under Democratic than Republican presidents. Moreover, if nominal wages and prices have different degrees of stickiness, demand and technological shocks will have redistributive effects on real payoffs to labor and capital. Supportive evidence for redistributive shocks can be found in the empirical literature: Abowd (1989), in a study on wage bargaining in the United States, finds a large and negative correlation between unexpected union wage changes and unexpected changes in the stock value of the firm; Bottazzi et al. (1996), using a VAR approach that imposes block exogeneity across countries (and hence, as discussed in the next sections, is likely to overestimate the benefits of international portfolio diversification), find that the correlations of returns to human capital with domestic market returns is smaller than the one with foreign market returns in 7 out of 10 countries in their study (with an average difference of 0.19); Lustig and Nieuwerburgh (2008)17 uncover a negative correlation between innovations to human and physical capital returns in the United States; Gali (1999), Rotemberg (2003), and Francis and Ramey (2009) document a negative correlation between labor hours and productivity conditioning on productivity shock. Note also that the above results have been obtained without considering the exchange rate risk connected with the investment in foreign assets. In the sample period considered, the lower bound on the estimated standard deviation of exchange rates in the three countries considered is about one-third of the standard deviation of market returns. Moreover, the exchange rates show a weakly positive correlation with the stock market of the foreign country and seem to be uncorrelated with the U.S. stock market and with labor income innovations.18 Therefore, adding exchange rate risk to the model would reduce the Sharpe ratio of foreign assets, making foreign investment less attractive and increasing the degree of home bias. Relaxing the borrowing constraints ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As a robustness check of the above results we now relax the short-selling constraint restriction. Relaxing this restriction reduces the buffer stock saving need, since short- selling increases the households' ability of smoothing wealth shocks over time via borrowing at the risk free rate (i.e. shorting the risk free asset). In particular, we relax the constraint by allowing the household to borrow (i.e. short-sell) up to a constant fraction of its annual (permanent component of) labor income. Table 4 computes the aggregate domestic portfolio as in Table 3 but considering a different level of households' borrowing capacity in each of its panels: in Panel A through D, respectively, short-selling is constrained to be no more than 20%, 50%, 100% and 200% of annual labor income. The table shows that the effect of relaxing the borrowing constraints is non- monotonic. Moderate and intermediate borrowing ability actually increases the degree of home country bias generated by the human capital hedging motive. On the other hand, allowing for extremely large borrowing reduces the degree of home country bias generated by the model. This non-monotonicity is quite intuitive. With moderate borrowing ability, when the household is sufficiently wealthy to invest in the financial market, it uses its borrowing capacity to leverage and hedge further the human capital risk by skewing holdings toward the domestic stock. As a consequence, the model in this case generates even more home country bias in portfolio holdings than in the baseline specification with no short-selling (reported in Table 3). When instead the household can borrow large amounts, given the large equity premia, borrowing at the risk free rate to invest in stocks is an attractive investment (for a power utility investor). Since the expected utility from the financial investment is maximised with a well diversified portfolio, a tension between human capital hedging and financial wealth diversification arises. As a consequence, when the household can borrow large amounts relative to the size of its human capital, the financial wealth diversification motive reduces the degree of home country bias in portfolio holdings. Nevertheless, even with the unrealistically high borrowing capacity considered in panels C and D, the model generates very large home country bias.19 This is due to the fact that the human capital of the household has a value that is a large multiple of its annual labor income (formally, the present discounted value of all future labor income), while the borrowing capacity, in realistic calibrations, is only a relatively small fraction (or a small multiple in panels C and D) of the current labor income.","To assess the role of human capital in determining optimal portfolio choice, and in particular to estimate the correlation between human and physical capital innovations, one needs to study the time series properties of the returns to human and physical capital. This task is complicated by the fact that neither the market value nor the returns to human and (total) physical capital are observable. Nevertheless, total payouts to both factors of production are directly observable from national accounting figures. Moreover, total payouts to the labor force and capital holders can be thought of as the aggregate dividends flows on the unobserved stocks of human and physical capitals. We can therefore use the Campbell and Shiller (1988) methodology to infer the time series properties of unobserved aggregate returns from the observed growth rates of dividends on human and physical capital. To make the above approach operational, we need to construct empirical proxies of the expected values in Eq. (9). We perform this task following Campbell and Shiller (1988) and use linear projections generated by a reduced form VAR in a similar fashion as in the seminal work of Baxter and Jermann (1997). There are two important implicit restrictions in Eq. (12). First, the first matrix on the right-hand side of the equation has all the off-diagonal matrices restricted to be zero, i.e., each country is assumed to be block exogenous with respect to the other countries: the first differences of log labor and log capital income of each country are not supposed to Granger-cause the first differences of log labor and log capital income in other countries. Second, the cointegration structure in the second term on the right-hand side of Eq. (12) rules out cross-country cointegrations between incomes of the factors of production—that is, it rules out international common trends (e.g., it rules out that capital income in different countries follows the same long-run stochastic trend). Our empirical analysis in the next subsection relaxes both of these restrictions and considers a more general class of VAR models for labor and capital incomes that allow for short- and long-run comovements across countries. Empirical evidence: a reappraisal ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We estimate Eqs. (9) and (12), as well as alternative VAR specifications, using annual data on labor income and capital income from OECD National Accounts for Germany, Japan, the United Kingdom, and the United States over the period 1960–2012. Our measure of labor income is total employee compensation. This is a less than ideal measure in that it is likely to contain components that are not purely compensation to labor (e.g. the wage bill received by an entrepreneur might contain capital compensation components), and consequently the literature has developed more accurate measures of compensation to labor (see, e.g., Gopinath et al., 2015). Nevertheless, using more accurate measures of wage compensation would require focusing on much shorter time series: depending on the approach, we would lose between 39 and 75 per cent of the time series dimension, and in such a reduced sample it would not be feasible to test for block exogeneity since the number of parameters to be estimated would be too large relative to the number of observations. Similarly, we are constrained to use a relatively small cross-section of countries (that, nevertheless, account more than half of the world equity market capitalization at the end of our sample, and more than 90% of the world capitalization at the beginning of the sample), since expanding to more countries would not only substantially increase the number of parameters to be estimated in each equation of the VAR system,22 but also shorten the available time series since the countries we consider are exactly the ones with the longest available history of wage bill data. This data limitation, nevertheless, has the advantage of making our results directly comparable to the previous literature and, in particular, to Baxter and Jermann (1997) since we use exactly the same set of countries and definition of the wage bill. Our baseline measure of capital income is GDP at factor cost minus employee compensation. A detailed description of the dataset is given in Appendix A.1. Model selection The restrictions imposed in Eq. (12) by Baxter and Jermann (1997) can be formally tested against more general specifications that allow for international comovements in the payoffs to the factor of production. We start by assessing whether the block exogeneity assumption is supported by the data. Table 5 reports frequentist tests of the null hypothesis of block exogeneity in Eq. (12). Both restricted and unrestricted specifications are estimated with only one lag (as in Baxter and Jermann, 1997), and we maintain the hypothesis of cointegration relationships only within the countries as in Eq. (12). As stressed by the p-values reported under the test statistics, the null hypothesis of block exogeneity is rejected at any standard confidence level. That is, the data suggest that there exist statistically significant cross-country links between the compensations of the factors of production. Table 6 reports the logs of the Bayes factor and the posterior probabilities defined by Eqs. (14) and (15) for a large set of models, under the assumption of flat priors and equal prior probability for each model. The models considered are as follows: i) vector error correction models—with (row 1) and without (row 2) block exogeneity restrictions—in which, as in Baxter and Jermann (1997), the only cointegrations allowed are within country and the fixed cointegration vector has the form [1, −1]; ii) VARs in levels with the block exogeneity restriction which relax the assumption of having a cointegration vector of the form [1, −1] (rows 3 and 4); iii) VARs in first differences (with, row 5, and without, row 6, block exogeneity) which rule out any form of cointegration among variables; and iv) unrestricted VARs in levels (rows 7 and 8) that allow for international comovements and arbitrary cross-countries—as well as within country—cointegration relationships. The maximum number of lags considered for each specification is naturally restricted by the sample size at hand, but nevertheless corresponds to the one chosen by Akaike and Bayesian information criteria. The econometric model considered in row 1 of Table 6 corresponds to the original Baxter and Jermann (1997) specification. The second row shows that relaxing the block exogeneity assumption leads to a dramatic increase in the log Bayes factor (log BFi). This increase is so large that if the models in the first two rows were the only ones considered, we would assign a posterior probability of about one to the specification that—unlike Baxter and Jermann—allows for international comovements among variables. That is, Bayesian testing confirms the strong rejection of the block exogeneity assumptions delivered by frequentist testing presented in Table 5. The models considered in the third and fourth rows maintain the block exogeneity assumption but, by considering VARs in levels, do not restrict the within country cointegration vector to take the form [ −1,1]. The Bayes factors of these specifications are of similar magnitude to the one in the first row, but much smaller than the one in the second row, providing additional evidence of a strong rejection of the block exogeneity assumption. The VARs in first differences with and without block exogeneity in rows 5 and 6 are relevant because they impose the restriction of no cointegration among variables. Since the Bayes factor in row 5 is smaller than the one in row 1, and the one in row 6 is smaller than the one in row 2, the data provide supporting evidence for within country cointegration in the payoffs to physical and human capital. Nevertheless, in Appendix A.4, we report a detailed frequentist analysis of within country cointegration and find mixed evidence in support of this hypothesis: the [ −1,1] cointegration vector is always rejected except for Japan, while relaxing this parameter restriction the results vary from country to country and with the lag length considered. Finally, the specification in rows 7 and 8 of Table 6 are unrestricted VARs in levels with one and two lags, respectively. These specifications allow for arbitrary cointegration within and, most importantly, across countries—that is, the variables are allowed to show both short- and long-run systematic comovements across countries. The specification with two lags (which can be mapped into a VECM) in row 8 delivers a Bayes factor that is substantially higher compared to all the other models considered. This large Bayes factor maps into a posterior probability (POj in the second column of Table 6) that is numerically indistinguishable from 1. That is, the data provide strong evidence of both short- and long-run cross-country comovements in the payoffs to human and physical capital, implying that the econometric model of Baxter and Jermann (1997)—that rules out both of these channels—is misspecified. To test the robustness of the above results, we use numerical integration of Eq. (13) (we used an importance sampling approach based on the asymptotic Normal- inverse-Wishart shape of the posterior to perform this task) to get alternative estimates of the Bayes factors and posterior probabilities in Table 6. We also experimented with non-flat priors over the parameters space. In both cases, the results are in line with the ones in Table 6. Overall, the results of this subsection imply that to accurately measure returns to human and physical capital, and their implications for international portfolio diversification, we should use an econometric specification that, differently from the ones used in the previous literature, allows for both short- and long-run international comovements in the payoffs to production factors. The correlation of human and physical capital returns In order to estimate factor returns using Eq. (9) and the selected VAR model, we calibrate the parameters ρ to 0.957 for both capital and labor income. This corresponds to assuming that the mean dividend–price ratio of labor income and capital income are identical and equal to 4.5% as in Baxter and Jermann (1997). Moreover, note that finiteness of the empirical estimate of the right-hand side of Eq. (9) is guaranteed if ρ times the largest eigenvalue of the companion matrix of the selected VAR model is within the unit circle. This condition is satisfied by our choice of ρ. Table 7 reports the correlations between returns on capital and labor computed using Eq. (9) and the estimations of expected Δd’s by the VAR in levels specification with two lags. The correlations are both qualitatively and quantitatively different from the ones derived by Baxter and Jermann (1997). The within countries correlations seem to be somewhat lower compared to Baxter and Jermann (1997): their estimates cover the range [0.78, 0.99], while our estimates have a maximum of 0.96 and a minimum of 0.67 in Japan.24 The between countries correlations appear to be higher compared to Baxter and Jermann (1997): their maximum correlation between returns on capital is 0.43 (U.S.–Germany), the maximum correlation between returns on labor is 0.35 (U.S.–Germany), the maximum correlation between domestic labor returns and foreign capital returns is 0.40 (Germany–U.S.). In our estimation, the between countries correlations for both rL and rK are much higher (with the exception of Japan, where the correlations between labor returns and foreign returns on both capital and labor are generally lower). The correlations between returns on capital, for example, cover the range [0.73, 0.98]. Moreover, for all countries but Japan, the correlations between domestic returns on labor and foreign returns on capital are similar to the correlation between domestic returns on labor and capital. These results suggest the presence of productivity shocks effective at the international levels. The factor returns that we obtain by applying Eq. (9) are generated regressors (Oxley and McAleer, 1993). To account for this, we compute the standard errors of the corresponding correlation matrix using a bootstrap approach to statistical inference (see, e.g., Efron and Tibshirani, 1993). More specifically, we apply a sampling-with-replacement raw residuals bootstrap scheme with 10,000 repetitions. Interestingly, as shown in Table 7, we find large empirical confidence intervals. For Japan, the confidence intervals indicate that the correlation between factor returns may be either positive or negative. In other words, we document the existence of substantial statistical uncertainty on measuring returns to the aggregate capital stock. The differences between our point estimates and the results of Baxter and Jermann (1997) are mostly driven by the relaxation of the block exogeneity assumption, which turns out to be strongly rejected by the data. Baxter and Jermann (1997) fit a model where they restrict the countries not to be economically and technologically integrated. As an outcome, the level of between countries correlation is underestimated and the within country correlation is overestimated (i.e., the countries appear not to be integrated). Once this restriction is removed, the effective degree of technological and economic integration becomes evident. This high degree of economic integration implies fewer opportunities to hedge the human capital risk investing in foreign marketable assets. For comparability, in Table A5 of Appendix A.6, we replicate the optimal portfolio implied by the complete spanning approach of Baxter and Jermann (1997). That is, we replicate the authors' main result, but after correcting their VAR misspecification (i.e. we use the VAR specification in Row (8) of Table 6). Raw point estimates indicate that the authors' original claim, i.e. that the home country bias is generally worsened by the human capital hedging motive, is not generally supported by the data: the table shows that, for some countries, we obtain exactly the opposite effect. Moreover, confidence intervals show that there is substantial uncertainty about optimal portfolios constructed with this approach, and that human capital hedging can potentially generate large home country bias.25 The correlation of human capital and stock market returns Estimating returns to physical capital using the approach in Eq. (9) might be inappropriate for evaluating international portfolio diversification, since i) only a subset of the claims to capital compensation in a country are tradable internationally with relatively little frictions—i.e., the ones of publicly traded companies, and ii) due to tax advantages, the compensation to capital elicited from national accounts tends to include de facto components of human capital compensation (e.g., for family owned and individual firms). As a consequence, a more appropriate measure of the correlation between human and physical capital compensations can be constructed replacing the VAR based estimates of returns to capital with the returns on broad stock market indices. In this subsection, we restrict the set of assets available to hedge human capital to include only publicly traded stocks, since we consider this as being the relevant case for most households. The innovations to human capital compensation are estimated as in the previous subsections. Table 8 reports the correlations between the returns to human capital and returns to the stock market. Since the correlation between U.S. human capital innovations and stock returns plays a key role in the model calibration presented in Section 2, for the United States we use three different stock indices: the Fama and French (1992) benchmark market return,26 the S&P 500 index, and the Dow Jones Industrial index. Overall, the table suggests that domestic returns to human capital are in general more correlated with foreign stock markets: for all countries considered, the asset with the highest correlation with human capital returns is always a foreign asset. Moreover, in the United States in particular, returns to human capital have the lowest correlation with the domestic, rather than the foreign, stock market index. However, the estimated correlations tend to be quantitatively small (this is consistent with the findings of Fama and Schwert (1977)), and the uncertainty attached to the estimation of human capital delivers large confidence intervals (in particular, only one of the estimated correlation coefficients is different from zero at the 95% confidence level). Nevertheless, as shown in Table A4 of Appendix A.5, despite the large confidence intervals, the pattern of lower correlation of human capital with domestic, rather than foreign, returns, is quite robust to alternative construction of the data. The small correlations in the tables have two important implications. First, when focusing on tradable claims to physical capital, the assumption of complete spanning for human capital return used in the previous literature seems to be unsupported by the data. Second, if one were to use a value weighted approach for the determination of optimal portfolios as in Baxter and Jermann (1997), the effect of human capital hedging would be very small, but it would tend to skew holdings in favor of domestic assets, as shown in Table A5 in Appendix A.6. For robustness, we have also estimated the correlations in Table 8 using two subsamples of equal length (pre and post 1987). The correlations in the two subsample are not statistically different from the full sample ones, albeit smaller in the second subsample. More importantly, the subsample estimates highlight the same pattern as in Table 8: labor income innovations tend to be more correlated with foreign stock markets than domestic ones. 27 Nevertheless, as we have shown in Section 2 above in the presence of borrowing constraints and both aggregate and idiosyncratic human capital risk, even the above small correlations can have a very large impact on optimal portfolio decisions.","This paper shows that human capital risk can help rationalize the home country bias in equity holdings at both the aggregate and household levels. First, we show that the theoretical intuition that short positions in domestic physical assets are a good hedge for human capital risk is a very fragile one, as very small redistributive shocks—e.g., with a variance of a mere 6% of GDP variance—are enough to reverse this intuition. Moreover, we find that the presence of this type of shocks is supported by the data. Second, we show that the commonly used approach of estimating country-specific VARs to compute returns to human and physical capital is rejected by the data and delivers mechanically biased estimates of the hedging benefits of shorting the domestic capital. Most importantly, we show that this misspecification largely drives the findings of Baxter and Jermann (1997)—i.e., the result that the home bias puzzle is unequivocally worse than we think once we consider human capital risk, is the outcome of an econometric misspecification strongly rejected by the data. Moreover, we show that when returns to physical capital are measured using broad stock market indexes, human capital return innovations tend to be more correlated with foreign rather than domestic stocks. Nevertheless, these correlations are small and, consequently, in a frictionless complete market setting, have very little impact on optimal portfolios. Fourth, calibrating a buffer stock saving model consistent with both micro and macro labor income dynamics—hence taking into account that individual labor income uncertainty is substantially larger than the aggregate one— as well as stock market data, we show that a large home country bias arises as an equilibrium result. This is due to the fact that household labor income risk is about one order of magnitude larger than aggregate labor income risk and, in the presence of liquidity constraints, optimal hedging becomes heavily skewed toward the asset whose innovations have the lowest correlation with the labor income innovations—the domestic asset. Moreover, our heterogenous agents model implies that, at the household level, the degree of home country bias should increase, and portfolio diversification decrease, in the labor income to financial wealth ratio—and these are exactly the novel empirical stylized facts that we uncover in the 1992–2013 U.S. Survey of Consumer Finances data."],["We study household income inequality in both Great Britain and the United States and the interplay between labour market earnings and the tax system. While both Britain and the US have witnessed secular increases in 90/10 male earnings inequality over the last three decades, this measure of inequality in net family income has declined in Britain while it has risen in the US. To better understand these comparisons, we examine the interaction between labour market earnings in the family, assortative mating, the tax and welfare-benefit system and household income inequality. We find that both countries have witnessed sizeable changes in employment which have primarily occurred on the extensive margin in the US and on the intensive margin in Britain. Increases in the generosity of the welfare system in Britain played a key role in equalizing net income growth across the wage distribution, whereas the relatively weak safety net available to non-workers in the US mean this growing group has seen particularly adverse developments in their net incomes. --------------------------------------------------------------------------------","Over recent decades, substantial changes in the distribution of incomes in both Great Britain (GB) and the United States (US) have placed increased pressure on government budgets.1 Declining employment and stagnant wages – each of which have affected both countries, to different extents and at different times – translate into reduced tax collections, while increased eligibility for and generosity of social insurance, means- tested transfer payments and work-based credits result in greater expenditures. The latter trend has been reinforced by the interplay between the labour market and the family, with increased inequality in family earnings and in assortative mating. The aim of this paper is to describe the relationship between inequality in labour earnings and the evolution of family income inequality. Tony Atkinson was the world leader in driving forward the study of economic inequality and its development over time, see Atkinson (1970, 1982, 1993, 1997, 2005). Many aspects of the work we present here take the lead from Tony's inspirational research in this field - in particular, the role of the tax and benefit system in mitigating earnings inequality and the interaction between the labour market and household income inequality, for example Atkinson (1992, 2000) and Atkinson (2006). Changes in wage inequality have been at the centre of much empirical research in labour economics. This includes large bodies of work aiming to identify causal channels (e.g. Bound and Johnson (1992); Katz and Murphy (1992); Card and DiNardo (2002); Bowlus and Robin (2004); Lemieux (2006); Autor et al. (2008); Blundell et al. (2016a, 2016b)) and to describe in some detail the key dimensions of change (e.g. Juhn et al. (1993); Katz and Autor (1999); Gosling et al. (2000); Piketty and Saez (2003); Burkhauser et al. (2012); Machin (2015); Guvenen et al. (2017)). However, there has been little systematic cross- country comparative work, and much less attention to the interaction between the tax and transfer system and family earnings in the evolution of household inequality. Family income inequality differs from wage inequality for a number of reasons. Family labour income depends also on hours of work and on how hours and wages covary between spouses, meaning the interplay between the intensive margin and jointness of the labour supply decisions, which may be heavily influenced by assortative mating in the marriage market (Blundell et al., 2016a, 2016b). In addition, the tax and transfer system can be a very important bridge between family labour income and living standards, through taxes, work- contingent credits and social assistance transfers. Tax and transfer systems are typically quite nonlinear, especially at low-incomes, and this can lead to very different inferences about levels of household income inequality; and major reforms to these systems can and do have large effects on the income distribution. We examine the labour market and tax and transfer system in its relationship with household income inequality in Britain and the US spanning the 36 years from 1979 to 2015. The approach we take is descriptive, but informed by structural changes in potentially-selective labour force participation, hours of work, assortative mating and income insurance provided by the tax and transfer system across the wage distribution. We develop an approach to study how the intensive margin of labour supply, family structure and the tax and transfer system have interacted over time to affect the link between wages and net family incomes right across the male and female wage distributions. To set the scene we begin by documenting and contrasting trends in male earnings and net (after-tax and transfer) income in each country. We then systematically trace out the path from individual labour market outcomes through to net family incomes, unpacking the underlying components of income inequality in the following sequence: Employment → Wages → Earnings → Family Structure → Family Market Income → Welfare → Gross Income → Taxes and Work-Based Tax Credits → Net Income. We explicitly consider the link between employment and wages with a median selection approach to bound wages in an effort to address selection into, and out of, the labour force, which has likely changed very differentially between the two countries over time (Johnson et al., 2000; Chandra, 2003; Blundell et al., 2007). In terms of the labour market, taking a relatively long-term view and considering trends since 1979, the basic background facts are that real wages have grown far less in the US than in Britain – and in fact have not grown at all at the median except for college graduates – while employment trends have looked relatively similar. However, over the past two decades, and especially since the Great Recession, employment has been more robust in Britain while wages have been more robust in the US. Britain has seen a large increase in male earnings inequality, not just during the much-documented 1980s inequality boom, but also since then. The increase over the past two decades was driven by a broadly secular decline in the hours of work of men at lower wage percentiles: inequality in male hourly wages between the 5th and 95th percentile changed little. The hours of work story has been the opposite among British women, among whom increases at the bottom of the wage distribution have reduced earnings inequality. This has not been enough, however, to stop family earnings inequality from rising. In the US, secular trends in hours worked (among workers) have been less pronounced, albeit with considerable cyclical variation around that, but male hourly wage inequality has increased. Meanwhile, employment among less-skilled men in the US fell over the sample period, and since 2000 has even fallen among higher-educated, and remarkably for women of all skill levels after a secular increase in the prior three decades. Using a bounding approach to account for the potential effect of selective entrances and exits from the labour market, we show that – especially since the Great Recession – wage trends among lower-educated groups may be more similar between the two countries than the raw data focused only on workers imply. Nevertheless, the basic qualitative comparisons between the countries prove robust to this bounding exercise. Even though there were sharp declines in hours of work among men in Britain, and some increase in assortative mating, the British welfare state has stabilized the economic inequality of tax units across the most of the net income distribution over the past two decades. For example, we show that 90/10 net income inequality fell slightly in Britain from 1994 to 2015 even though male earnings inequality increased. In comparison, we show that in the US 90/10 net income inequality rose sharply, suggesting that the US tax and welfare system is less successful at counteracting changes in the labour and marriage markets. The greater stabilization in Britain did come at a considerable fiscal cost, in particular due to large increases in the generosity of tax credits in the late 1990s and early 2000s which led to these credits trebling as a share of GDP from 0.5% in 1997 to 1.5% in 2004.2 The paper proceeds as follows. Section 2 gives a brief overview of the key policy context in both Britain and the US. Section 3 discusses the data we use in the paper, including how we harmonize the measurement of key variables across countries to the extent possible. Section 4 sets out the context of overall changes in net family income inequality in both countries, and how this relates to male earnings inequality. We then unpack the links between these. Section 5 begins with the labour market, including how it interacts with the marriage market, while Section 6 examines the impact of the tax and transfer system. Section 7 then brings these together by systematically tracing the links from wages right through to net family incomes. Section 8 concludes.","During the period considered in this paper there have been a number of key policy changes in both countries that are relevant for our analysis. In Britain there were significant cuts to income taxes during the 1980s, especially for higher earners. The top marginal income tax rate fell from 60% to 40% in 1988, and the basic rate of income tax fell in stages through the decade from 30% to 25%. Since 1994, which – for data reasons – we focus on for much of the analysis, the basic rate of income tax has fallen further in a number of incremental steps to 20%, and since 2011 the zero-rate band has been expanded rapidly. However, fiscal drag and some discretionary policy changes have pulled many more individuals into the higher tax bracket: the number paying the marginal rate of at least 40% has more than doubled since 1994.3 The net result is that the income tax system has become more progressive in recent years (with the opposite having happened in the 1980s). Since the late 1990s much of the key policy change in Britain has been on the transfer side. The Labour governments of 1997 to 2010 presided over large increases in the generosity of social assistance and tax credits, in large part as a means of pursuing ambitious quantitative child poverty targets for 2010 and 2020 (Joyce and Sibieta, 2013). The term ‘tax credits’ in Britain is in fact used to describe two very different forms of support: a genuinely work-contingent transfer4, currently named Working Tax Credit (WTC), and an additional means-tested element specifically for families with children (Child Tax Credit, CTC) which is available – since 2003 – to low-income families irrespective of work status. The out-of-work safety net was also made significantly more generous for families with children under Labour. Since 2011, however, a broad-based set of cuts to means-tested working-age transfers have been implemented as part of post-recession fiscal consolidation measures. These are clearly evident in the analysis we present later up to 2015, but they continued after that and are set to continue for several more years. Another important policy change in Britain was the introduction of the National Minimum Wage in 1999. It was subsequently increased in several stages, and by 2015 (the end of our period of analysis) it covered around 4% of employees. It is, however, now being extended much further and is set to cover around 12% of employees by 2020 (Cribb et al., 2017). Like Britain, the economic landscape of the United States over the past several decades has been characterized by massive changes to tax and welfare policy. The Economic Recovery Tax Act of 1981 and the Tax Reform Act of 1986 jointly broadened the tax base and reduced the number of federal income tax brackets from 16 to four, with the marginal tax rate on the highest income earners dropping from 70% to 28% by 1989 (Auerbach and Slemrod, 1997; Burman et al., 1998; Kniesner and Ziliak, 2002). The subsequent tax changes over the ensuing two decades eventually led to a return to seven marginal tax brackets and a top rate of 39.6% by 2009. Although the tax reforms expanded the standard deduction and personal exemption amounts, and thereby removed several million low-income households from the federal tax rolls, there were strong incentives for these families to file in order to claim refundable tax credits for workers; namely, the Earned Income Tax Credit (EITC) and the Additional Child Tax Credit (ACTC). The EITC was created in 1975 and targeted to low- wage workers (Nichols and Rothstein, 2016). The generosity was expanded several times in the 1980s and 1990s, and by 2014 the maximum credit was $5460 for a family with two qualifying children and annual earnings under $17,580. Over 28 million taxpayers claimed the credit that year at a current-year cost of over $68 billion, or 0.4% of GDP. The non- refundable Child Tax Credit and refundable portion ACTC were established in 1997 and (currently) provide a credit against tax liability of $1000 for each child under the age of 17. Initially eligibility was restricted to workers with annual earnings in excess of $10,000 in 2001 (and indexed to inflation thereafter), and most benefits went to the middle and upper-middle class. As part of the 2009 response to the Great Recession, the eligibility limit was lowered to $3000, thus better targeted the ACTC to part-time and part-year low-income workers. By the 2014 tax year, expenditure on the ACTC program exceeded $30 billion, or 0.2% of GDP. Concomitant with falling marginal income tax rates and expanding credits were substantial expansions in the payroll tax, which is used to finance Social Security retirement benefits, disability benefits, and Medicare health insurance for the elderly and disabled. While the rates have not changed since 1991 (15.3% combined employer/employee rate), the base applicable to Medicare tax (2.9 percentage points of the 15.3) was uncapped that year, and the retirement and disability benefit base subject to taxation was indexed to inflation and by 2014 was $117,000. Alongside the major changes to tax legislation were wholesale changes to means-tested transfers during the 1990s. The reforms altered significantly the economic rewards to work and to participation in transfer programs, and affected all segments of the low-income population. Some programs retrenched, while others witnessed dramatic growth (Ziliak, 2015). The 1996 Personal Responsibility and Work Opportunity Reconciliation Act abolished the cash welfare program Aid to Families with Dependent Children, which was an entitlement program for low- income and low-asset (single-mother) families with children under age 18, and replaced it with the time-limited, block-grant program Temporary Assistance to Needy Families (TANF). TANF limited eligibility to no more than five years, and less at state discretion, and imposed work requirements and numerous other restrictions on eligibility (Ziliak, 2016). While this program change effectively eliminated out-of-work cash welfare in the US, since 2000 there was huge growth in food assistance spending from the Supplemental Nutrition Assistance Program (aka food stamps), in health insurance coverage for children—first with state-directed Medicaid expansions, then federal creation of the Supplemental Children's Health Insurance Program, and finally the 2014 rollout of the Affordable Care Act—and steady growth of disability benefits both related to work (Disability Insurance) and childhood (Supplemental Security Income). Taken together, inflation-adjusted spending on the major US social insurance and means-tested transfers grew 60% to over $2 trillion by 2010, or over 13% of GDP (Ziliak, 2015).","We begin by providing a brief overview of our data sources, followed by a detailed description of how the various labour market and income sources were measured. We endeavoured to the extent possible to harmonize the datasets across countries over the past three and a half decades to provide a consistent and comprehensive portrait of the economic circumstances of individuals and their families in Britain and the United States. Great Britain ~~~~~~~~~~~~~ For the research on Britain, we draw on two distinct sources of data: the 1979–1993 survey years of the Family Expenditure Survey (FES), and the 1994–2015 survey years of the Family Resources Survey (FRS).5 Both datasets are annual household surveys and are commonly combined in this manner, including in the calculation of official statistics on poverty and inequality. The FES and FRS collect data on various sources of income received and taxes paid close to the time of interview, and all income and tax amounts are based on the self-reported values. A very small fraction of income components (typically less than 1%) suffer from non-response and any missing values are imputed. However, as neither survey identifies the observations and income components that have undergone imputation, we are unable restrict our sample to those without any imputed information. We restrict our sample to men and women aged 25–55 to focus on the prime working-age population, and thereby abstract from the part of the lifecycle where most human capital investments occur and that part associated with retirement. United States ~~~~~~~~~~~~~ For the US analysis, we use the Current Population Survey Annual Social and Economic Supplement (ASEC) for the 1980–2016 survey years. The ASEC is a stratified random sample of 60,000–90,000 household addresses from the noninstitutionalized population in the US. It serves as the official source of income and poverty statistics and has been the workhorse dataset for research on wage and income inequality. As with the British data, we restrict our focus on men and women aged 25–55. However, there are some important distinctions in the ASEC. First, all information refers to prior calendar year rather than the time immediately prior to the interview, as in the British data. Second, taxes and tax credits are self-reported in the British data, whereas the ASEC does not collect tax information. Instead we run the ASEC data through NBER's TAXSIM simulation program, which assumes 100% take-up among those eligible for tax credits. Third, nonresponse to earnings questions, and to the entire ASEC altogether, has been on the rise (Bollinger and Hirsch, 2006; Bollinger et al., 2017), and the US Census Bureau imputes values to nonrespondents. We drop those with imputed earnings and hours and reweight the ASEC data as described below. Measuring labour-market outcomes and incomes ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The primary economic outcomes in our analysis are employment, hours, real earnings and wages, and real before-tax gross income, and real after-tax and transfer (net) income. Employment rate In the British data, we measure the employment rate as the fraction of the population aged 25–55 employed during the survey week (sometimes referred as employment per capita). The measure is the same in the US, except employment is for any time in the prior year. Hours of work In both countries, hours of work refers to usual hours worked per week, where the reference period in Britain is “typical” hours in the current financial year, while in the US it is typical hours in the prior year. The data from Britain distinguishes between paid ‘basic’ and both paid and unpaid overtime hours. The hours measure we use is defined using paid basic and paid overtime hours only in order to more accurately reflect trends in formal labour market arrangements. No such distinction is made in the US. Overtime hours in the US primarily only apply to workers paid by the hour, and those workers are eligible to be paid 1.5 times the normal hourly wage. Real earnings and wages In the British data, information on earnings is obtained by asking respondents the amount they were paid on the pay date closest to interview. Raw responses are converted into nominal weekly amounts and we additionally convert these nominal values to real terms using a modified Consumer Price Index that includes an adjustment for mortgage interest. In the US, earnings are measured for the past year, and deflated by the Personal Consumption Expenditure Deflator. In both cases we use a 2010 base year. Real hourly wages are constructed as the ratio of weekly real earnings and usual hours per week in Britain, and the ratio of real annual earnings to annual hours of work (hours per week times number of weeks worked). We leave each country's earnings and wages in their respective currencies. For the analysis that relies on wage information, we exclude those with extreme gender- specific real average hourly wages (below 1st percentile; above 99.9th percentile) and adjust the survey weights using inverse probability weighting. Specifically, for each gender and year, we estimate a saturated probit model of the probability of not having an extreme wage using levels and interactions of age, race, education, marital status, and other demographics. We then divide the survey weight by the fitted probability of not having an extreme wage. For the US, we modify the procedure to also account for non-imputed employment and earnings. The reweighting approach results in consistent estimates under the assumption that the excluded observations are missing mean conditional at random. As we describe in the results section, this assumption is relaxed when we bound the wage series with worst-case bounds to account for possible nonrandom selection into employment. Gross and net income As we are ultimately interested in changes in family-level outcomes, in addition to individual-level employment and earnings we also construct gross and net income at the tax unit level. Tax units in the Britain are defined as an adult, their partner (married or unmarried), and any dependent children in their care. In the US data they are inferred from household relationship pointers and ages of occupants, where unlike Britain, cohabiting partners in the US do not file jointly.6 Our measure of gross income includes the earnings of the primary and secondary earner (if present), transfer income and nontransfer nonlabour income such as rent, interest, and dividend income. In the British data, transfers include all cash transfers and work-based tax credits, including the Child and Working Tax Credits, Child Benefit, Housing Benefit, Income Support and unemployment and disability benefits. For the US data, transfers include Social Security, Disability Insurance, Unemployment Insurance, Workers Compensation, Supplement Security Income, Temporary Assistance for Needy Families (cash only), Supplemental Nutrition Assistance Program (food stamps), Earned Income Tax Credit, and the Additional Child Tax Credit. Some of the benefits are recorded in the surveys at the individual level, and others at the family level. For the former we sum them up across all individuals in the tax unit. For both countries we rely on self-reported information when calculating transfer income (in the US, the EITC and ACTC are simulated with TAXSIM). In both the FES/FRS and the ASEC this approach is known to lead to systematically lower spending estimates than those observed in administrative data (Meyer et al., 2015; Brewer et al., 2017). While our main analysis does not account for such under-reporting, we provide additional results that adjust the self-reported benefit income amounts to match totals taken from administrative data and show headline trends are robust to this. Net income is constructed as gross income less tax payments, which in the British data includes income tax, employee National Insurance Contributions, and Council Tax.7 As noted previously, tax payments and credits are not reported in the US data and must be simulated. The NBER TAXSIM program receives as inputs the tax unit marital status, ages of members, number of (child) dependents for (refundable) tax credits, earnings, taxable and nontaxable transfers, and other items. It then returns a simulated estimate of federal, state, and payroll tax liability, inclusive of tax credits. For the payroll tax, we just assign the employee share. Finally, because household size and composition has changed substantially in both countries in recent decades, we equivalise gross and net income using a modified OECD scale.8 Education For many of our outcomes we split the sample into education groups, which is a standard proxy for skill and/or permanent income. Variables related to educational attainment in the British surveys have changed over time. In order to create a continuous time series we therefore focus on school-leaving age, which is consistently recorded over the entire 1979–2015 period, and use this indicator of education to define four groups: left education aged 16 years or younger; left aged 17 or 18; left aged 19 or 20; and left aged 21 or older. These age categories roughly approximate the four US education groups of less than high school, high school graduate (or General Equivalency Degree), some college (includes community college and associates degrees), and four-year college or more. Importantly, however, those leaving school at age 16 in Britain receive credentials, whereas they do not in the US, and thus the low-educated group in Britain likely has more qualifications than the typical US “dropout”. Appendix Fig. 1 demonstrates that there has been substantial education upgrading in both countries since 1979, with a reduction in half of the lowest education group. In Britain, 80% of men and women left school by age 16, and this plummeted to 40% by 2015. The comparable percentages in the US were roughly 20 and 10%, respectively. Notably, the most marked growth in both countries is the highest education level, especially among women when 35 (40)% of British (US) 25–55 year olds attained the equivalent of college or more in 2015, double the rate in 1979. Marital status The remaining key demographic outcome that factors prominently in our analysis is marital status. In the British data, couples who are married cannot be distinguished from those who are cohabiting, while in the US data cohabiting couples are treated as unrelated individuals and marriage only refers to those couples in a legally recognized union.9 Appendix Fig. 2 presents trends in the fraction of men and women married (or cohabiting in Britain) by the four education groups. The substantial retreat from marriage is most evident among the least skilled, especially men in the US. In 1979, the fraction of married US men with high school or less was just under 80%, and greater than the fraction married among those with a college degree. By 2015, the fraction of high school graduates or dropouts who were married was nearly 20 percentage points lower than that of college educated men. Similar patterns hold among US women, and both British men and women, though they are much more attenuated in Britain.10","Net income among ‘working age families’ in Britain (denoted as G.B. in all figures) and the US is presented in Fig. 1. It shows strong growth from 1979 to 2015 in household income across the distribution in Britain, and for the top half of the distribution in the US, though relatively flat net incomes in the bottom half, except for the brief window in the late 1990s. The experience in the two countries during the Great Recession, however, was markedly different. Real net incomes fell sharply in Britain, especially in the upper percentiles, while they continued to keep pace with inflation in the US. Although the top of the income distribution has grown considerably since the mid-1990s in both countries, Fig. 2 shows that the 90/10 ratio of net income inequality has been stable in Britain over this period, while increasing steadily in the US since 2000 (largely due to a rise in the 90/50 not shown in the figure). The British experience of stable 90/10 net income inequality stands in stark contrast to the sharp rise in male (individual) earnings inequality. This suggests the insurance against relatively weak earnings growth provided by family structure and the tax and benefit system may differ substantively from the US where earnings inequality has increased alongside net income inequality. Fig. 2 also highlights that male earnings inequality is much more volatile in the US than in Britain, which as will be seen below, reflects much greater cyclical sensitivity in hours of work, especially among low-income workers. To verify the trends in net income growth and in net income inequality documented here are robust to potential under-reporting of transfer income, Appendix Figs. 3 and 4 repeat the analysis shown in Figs. 1 and 2 using a measure of net income that rescales transfer income to match transfer spending totals taken from administrative data.11 Appendix Fig. 3 shows this adjustment leads to slightly stronger net income growth at the bottom of the distribution in both countries. Appendix Fig. 4 shows 90/10 net income inequality in both countries is slightly lower when one accounts for under-reporting of transfer income, although trends in inequality are broadly similar to those shown in Fig. 2, particularly since 1994 which is the period we focus on in later analysis.12","The dramatic differences in Britain and the US in terms of overall after-tax and transfer income inequality, in contradistinction to the rising male earnings inequality in both countries, forms the basis for the ensuing analysis, where we first examine differences in employment and wages in each country. Employment, hours and wage inequality by gender, education and race ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Fig. 3 sets out employment rates over time in both countries, by gender and education level. Comparing levels of employment, perhaps the most striking difference is how much larger the gap between the highest- and lowest- educated is in the US than in Britain – especially for women. Part of this difference is explained by the fact discussed in Section 3.3 that the lowest education group in the US are less likely to have obtained formal educational qualifications than the equivalent group in Britain. Looking at trends over time, male employment rates in both countries are lower than they were in 1979, especially for the lowest educated. However, in the US this is driven by a broadly secular decline since around 1990. In Britain, by contrast, male employment has been on an upward trend since the early 1990s (punctuated temporarily by the Great Recession), after falling sharply through much of the 1980s and during the early 1990s recession. The result has been a marked convergence of male employment rates in the two countries over the past 25 years, from a starting point at which male employment in the US had been considerably higher for all but the lowest educated. Among women, employment was stable or gently rising in both countries during the 1980s, but again it has since been in secular decline in the US – especially for the lowest-educated – while remaining stable or increasing slightly in Britain. Over approximately the past 25 years, trends in employment have been much more robust in Britain than in the US and this has been especially evident since the Great Recession. Appendix Fig. 5 documents further heterogeneity in employment trends by disaggregating by race and education groups.13 This shows the employment rate of less- skilled non-white men in both countries is substantially lower than the rates observed among other groups of men, especially in US. Higher-educated black men have employment rates comparable to high school dropout white men in the US, and the gap between both of those groups and higher-educated white men has expanded in the last decade. Remarkably, there is no race gap in employment for US women, only a gap based on education attainment. It is not just the extensive margin of employment that has been important in driving changes in incomes and inequality. Fig. 4 documents mean hours of work among workers in the two countries over time, split by gender and education. The figure shows a large difference between the US and Britain in the patterns of male employment at the intensive margin across skill groups with higher-educated men working far more hours than the low- educated in the US and vice versa in Britain. However, this contrast is in part due to differences in the treatment of unpaid overtime in the hours measure used in each country, as discussed in Section 3.3. Specifically, accounting for unpaid hours worked in Britain leads to the same ranking of education groups observed in the US, as unpaid work increases the average hours worked by the highest education group while leaving average hours of lower-educated workers largely unchanged. For women the relativities across skill groups are the same in both countries, with higher educated women working more hours; but US women work considerably more hours than their British counterparts, on average. Among women, average hours of work have been quite stable in both countries in recent decades, after rising during the 1980s. The one exception is the lowest-educated women in the US, whose hours of work have fallen since the mid 2000s. For men the key pattern has been a large convergence in hours of work across education groups in Britain. This has been driven by particularly large falls in hours among the lower-educated. Appendix Fig. 6 provides some detail behind this, showing percentage point changes in rates of “mini-jobs” (less than 16 h per week), part-time work (less than or equal to 30 h per week) and especially long hours of work (greater than 45 per week) between 1994 and 2015 across the hourly wage distribution. This highlights that reductions in hours of work among British men are in fact particularly concentrated in the bottom quintile of the hourly wage distribution and have been driven by both a reduction in the prevalence of long hours and an increase in the prevalence of part-time work. There has also been a sharp fall in the prevalence of “mini-jobs” among women in the bottom quintile of the wage distribution in Britain.14 By contrast, the hours changes among men and women in the US have been far more uniform across the wage distribution. Following from these changes in employment, Fig. 5 shows how median real hourly wages among those in paid work have developed for the different education groups. The significant contrasts in employment trends between the two countries suggest the observed wage trends may be in part driven by trends in the selectivity of the workforce. To account for this, we implement a modified version of the median selection model (see, e.g. Johnson et al., 2000; Chandra, 2003; Blundell et al., 2007) which bounds wage trends by assuming that all changes in employment rates are the result of entrances and exits at the bottom of the within-group wage distribution.15 The bounded series are indicated by dashed lines. The US has seen a remarkably long period of real wage stagnation, stretching back over most of the period since 1979, with the only clear exception being a short period during the boom of the late 1990s. In fact, for men it is only college graduates among whom median real wages are currently any higher than in 1979. The bounded series confirm that accounting for trends in selectivity would only make this conclusion stronger, due to large employment declines among lower-educated men over this period. In Britain, wage growth was considerably more robust until the early 2000s. The more recent comparison is different. Since the mid 2000s, and especially the Great Recession, Britain has seen marked declines in median hourly wages across most groups (but less so among the lowest educated). These wage trends tend to be worse than seen among similar groups in the US over the same period. It does, however, turn out to be quite important to assess employment and wage trends, and the link between them via selection, in a coherent framework. The potential for wage trends among less educated US men to have been flattened by selection (due to falling employment) in recent years is significant, and the bounded series show falls in wages more in line with their British counterparts.16 Nevertheless, overall Figs. 3–5 show a stark difference in the nature of the impact of the Great Recession on the US and British labour markets. Employment has proven more robust in Britain, on both the extensive and intensive margin, particularly through the pace with which employment rates recovered after the initial shock. By contrast much more of the adjustment in Britain has come through lower real wages, especially for the high educated. These developments resulted in the post-recession decline in top net incomes in Britain as shown in Fig. 1, while they reinforced pre-existing trends towards higher inequality in the US. In combination, these trends in wages and hours of work have led to increased male and reduced female earnings inequality in both countries. This is depicted in Fig. 6, which highlights just how influential intensive margin trends have been in Britain as the growth in male (female) earnings is far greater towards the top (bottom) of the distribution than growth in wages. In the US, however, the close alignment between growth in wages and earnings across the distribution suggests that changes in earnings inequality are primarily due to changes in wage inequality, rather than trends in employment on the intensive margin.17 To assess the relative importance of wage and hours trends more formally, we decompose the change in the log of individual weekly earnings into components that are attributed to changes in the variance of log hours and log wages and the covariance between log hours and log wages.18 Table 1 reports the results of this decomposition separately for men and women in each country over three periods: 1994–2015, 1994–2007 and 2007–2015. The first two panels of Table 1 confirm that male earnings inequality has risen in both Britain and the US over the 1994–2015 period, with the rise in Britain almost three times as large as that in the US (an increase of 0.103 compared to 0.035). One reason for this difference is that earnings inequality among British males rose consistently over the entire period, whereas in the US it was largely unchanged between 1994 and 2007 before increasing during the period after the financial crisis. This reflects Fig. 2, which showed a secular increase in the British male earnings 90/10 ratio compared to a far more cyclical trend in the US 90/10. The right-most three columns of the table show that increases in the variance of wages and the covariance between hours and wages are both important drivers of the rise in male earnings inequality in Britain, accounting respectively for 49.7% and 44.7% of the overall increase in the variance of earnings. Although the covariance between hours and wages has had a substantial impact on US male earnings inequality, the variance of wages is more important and accounts for over two-thirds of the increase in the variance of earnings over the 1994–2015 period.19 In contrast to the rise in male earnings inequality, the third and fourth panels of Table 1 show falls in female earnings inequality across countries. The magnitude of the change is again greater in Britain than the US (a reduction of 0.105 compared to 0.02), and is primarily due to reductions in inequality that occurred in the pre-recession 1994–2007 period. In both countries, these reductions have been driven by falls in the variance of weekly hours. In Britain this has been reinforced by a reduction in the covariance between hours and wages, whereas in the US hours and wages among women have become more positively correlated. In summary, something has happened in Britain in recent decades which goes against the conventional wisdom that male employment at the intensive margin is relatively fixed. The breakdown of this rule has had first order effects on earnings inequality in Britain. In a comparative context it tempers the conclusion that one would reach when focusing on the extensive margin alone, which is that male employment has been on a worse trajectory in the US with a particular problem among the lowest skilled. The British story becomes more reminiscent of the US story once the intensive margin is incorporated. Belfield et al. (2017) have shown that the increase in part-time work among low-wage British men has occurred among single men and those in couples, and those with and without children. Explaining the origins of this change, and in particular whether it represents a demand-side or supply-side shift, is a key challenge for future research given its implications for welfare and potential possible policy responses. A satisfying explanation would need to account for why Britain has not seen similar concurrent changes at the extensive margin, and why the adjustments in this respect have been the opposite of those in the US. Marriage and assortative mating ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We now make the important move from individual labour market outcomes to family-level outcomes. A key part of this link is the pattern of assortative mating, which is examined in Figs. 7 and 8. For each country and gender we rank by percentile of individual hourly wages and plot changes in spousal characteristics within each percentile group, comparing 1994 with 2015. In Britain, but not the US, we are able to observe non-married cohabiting partners – though for parsimony we use the term “spouse” to cover any cohabitation between partners. Although this introduces an inconsistency in measurement between countries, long-term non-marital cohabitation is comparatively less common in the US. Appendix Fig. 2 showed that living as part of a couple has generally become less common across education groups in both countries, while Fig. 7 shows that these changes are pervasive across the majority of the wage distribution. It also shows that this change has tended to be more pronounced for people in the bottom half of the gender-specific hourly wage distribution, and even more pronounced among non-working men (as indicated by the dots on the left-hand side of each panel). Changes in the probability of having a working spouse exhibit a similar gradient across the wage distribution. The gradient here is especially strong for women. In Britain the probability of having a working spouse has tended to decline in the bottom half of the female wage distribution and to increase in the top half. In the US, it has declined throughout essentially the entire distribution, but by more towards the bottom. Fig. 8 examines how the within-family correlation between wages has changed since 1994, plotting the average wage percentile rank of spouses by own-wage percentile (for those in each percentile group who have a working spouse). For both countries and genders, there is a clear positive correlation: people further up the individual wage distribution tend to have spouses who, if in work, are also further up the wage distribution. In the US, but not Britain, there is also clear evidence of an increase in this form of assortativeness as the gradient has become steeper over the past two decades.20 Table 2 shows inequality in total earnings among dual-earner couples has increased in both Britain and the US. In the British case, the increase is predominantly due to an increase in earnings inequality among the higher-paid member of the couple. Since main earners are overwhelmingly male, this result mirrors the rise in male earnings inequality documented above. Increases in the covariance between main and secondary earner pay have also acted to push up tax unit earnings inequality, particularly in the non-extreme section of the distribution, which indicates that two-earner couples in Britain have become increasingly assortative over the last two decades. The US results show tax unit earnings are more unequal than in Britain. Comparing the trimmed and untrimmed samples in the US reveals that increases in inequality among main earners has been entirely driven by the tails of the distribution: trends in main-earner inequality have actually acted to reduce tax unit earnings inequality among the middle 90%.22 Trends among secondary earners have acted to increase total earnings inequality in the US – due to both an increase in I2 inequality in their earnings and an increase in their share of total tax unit earnings – whereas they have slightly reduced inequality in Britain. Finally, Table 2 shows that increases in assortativeness have made a greater contribution to rising inequality in the US than in Britain, which mirrors the pattern shown in Fig. 8.","Another key bridge between individual labour market outcomes and family incomes is the government transfer and tax credit system. It makes sense to analyse this when moving to the family level: eligibility for such transfers is typically assessed at that level and so, at least where resources are pooled within families, transfer program participation measured at the individual level is not as meaningful. Fig. 9 documents trends in the generosity of the transfer systems of both countries by plotting the average share of family gross income that comes from the transfers and in-work tax credits by gender and education. The figure makes evident the greater generosity of the British welfare system across the education distribution in comparison to the US system, at least until most recently, with transfers and in-work tax credits accounting for a higher share of gross income among men and women of all education levels in Britain. This is in spite of the fact that, as discussed above, the lowest education group in the US is likely to be far less skilled than the lowest education group in Britain. The figure also shows the average generosity of the welfare system has tended to increase over the past 20 years in both countries, particularly among the lower educated in the US.23 The impact of the Great Recession on transfer income is also clearly shown. The increases in average welfare receipts in Britain that occurred in the years immediately following the financial crises have since been offset owing to the post-recession fiscal consolidation, which began in 2011 and included cuts to many transfer programs. The dramatic increase in average transfer generosity in the US emerged in response to the Great Recession – increasing average payments by 50% among the least skilled, but also more than doubling among those with some college – and unlike Britain, have remained elevated through the six years following the official end of the recession.","Bringing together the individual labour market outcomes, assortative mating and trends in welfare income, and adding in taxation, we can then trace the links from individual wages right through to net family incomes. To illuminate this, in Fig. 10 we rank people according to their position in the gender-specific hourly wage distribution and, keeping that ranking fixed, examine changes in different measures of income over the 1994–2015 period. The figure also shows growth in the different measures of income for non-workers which, as documented above, now account for a greater share of the working-age US population than in 1994. We start with family labour income, cumulatively add in work- based credits and then all other transfers (to make “gross income”), before subtracting direct taxes (to make “net income”). Family incomes are equivalised throughout this exercise in order to account for changes in family size and structure. The broad pattern in family labour incomes is one of increased inequality between higher- and lower-wage individuals, with the exception of the bottom male wage quintile in the US. These patterns are in line with the trends already documented in male earnings inequality (male earnings remain the dominant source of family labour income, on average) and the supporting role played by increases in assortative mating. However, important differences emerge between Britain and the US when looking beyond labour income. Transfers and taxes have had significant effects on trends in inequality between high- and low-wage people in Britain, but virtually no discernible impacts on those trends in the US. Work-contingent transfers actually have little to do with this, as they remain only a relatively small part of the overall transfer system in Britain (even for people in work). But increases in the generosity of the transfer system more generally, most importantly through CTC (most of which goes to families in work), have pushed the rate of growth in family gross income at the lower end of the wage distribution above the rate of growth in labour income alone. Direct tax cuts have had a further, similar impact towards the bottom, as the zero-rate income tax band has been increased sharply since 2010. Another striking point of contrast between Britain and the US is the experience of non-workers (represented by the dots on the left side of each panel). In Britain their net family incomes have grown robustly over the past 20 years, and more quickly than for the majority of the wage distribution. Unsurprisingly this is again due to increases in the generosity of the transfer system, particularly for families with children, both through CTC and through increases in the rates of out-of-work transfers. In the US, by contrast, non-workers have fallen further behind those in work over the past 20 years and in fact have seen barely any income growth at all, although the figure does suggest that growth in welfare income has mitigated to some extent the reductions in labour income among non-working US women. Appendix Figs. 9 and 10 repeat the analysis of Fig. 10 separately for singles and couples and those with and without children. These additional figures reveal that growth in total labour income has been more unevenly spread across the wage distribution among couples, re-emphasising that changes in the assortativeness of marriage have been an important driver of income inequality over the last two decades. The equalizing effect of changes in transfer income is most pronounced among those with dependent children, which is to be expected as the major welfare policy reforms in both countries have explicitly targeted this group. Overall, Table 3 shows the gap between net income inequality in Britain and the US has widened over the last two decades. The squared coefficient of variation in net income was around 23% higher in the US than in Britain in 1994, whereas in 2015 it was almost 40% higher. The most recent comparison shows that inequality in labour income is higher in the US and is offset to a far lesser extent by transfer and credit income.","Both Britain and the US have witnessed secular increases in 90/10 male earnings inequality over the last three decades. Up until the 1990s this was accompanied by similar increases in 90/10 inequality in net household incomes in both countries but since then trends have diverged with inequality in net family income declining in Britain while continuing to rise in the US. This paper has sought to shed light on the reasons for this divergence, taking inspiration from Tony Atkinson's extensive work on inequality, which emphasized the importance of accounting for the interplay between inequality in the labour market, the tax and benefit system and household income inequality. Since 1979, there have been sizeable changes in male and female employment in both countries. These employment changes have primarily occurred on the extensive margin in the US, with employment declining across gender and education groups from around 1990. In Britain, by contrast, the biggest changes have occurred on the intensive margin, with male workers experiencing declines in average hours of work that have been steepest for the lower-educated and most pronounced in the bottom quintile of the wage distribution. The impact of these trends in employment and hours on family-level income inequality has been mediated through several channels. First, changes in individual-level earnings inequality will also be influenced by changes in wage inequality. We find that wage growth has been relatively equal across the main part of the gender-specific wage distributions of both countries, although a novel worst- case bounding approach suggests that reductions in employment in the US may have attenuated growth at lower percentiles of the US wage distribution. As a result, the intensive margin changes observed in Britain led to a sharp reduction in female earnings inequality but a sharp increase in male earnings inequality. Second, the link between individual-level earnings and family-level labour income depend on changes in family composition and marital sorting. Focussing on the period since 1994, we find that both in Britain and the US, reductions in marriage have been greatest among low-wage workers and non-workers. In addition, the US has experienced an increase in assortative mating in terms of the correlation between wage percentiles of both members of a couple. The result of these trends has been an increase in inequality in family labour income among men and women in both countries. The most important final link from family labour income and net income is the tax and benefit system. Indeed, we find that the divergent trends in net income inequality in Britain and the US are largely due to the different policy regimes. Specifically, increases in the generosity of transfer payments in Britain under successive Labour governments between 1997 and 2010 boosted net income growth among low-wage workers and non-workers thereby equalizing growth rates in net income across the main part of the wage distribution. Policy changes on this scale have not occurred in the US with the result that the pattern of net income growth of US workers overall largely matches the pattern of family labour income growth. Differences in welfare policy are also key to understanding the differential fortunes of non-workers between countries. In Britain, many transfer payments are not contingent on work and therefore non-workers have witnessed relatively strong net income growth in comparison to workers. In the US, by contrast, a major part of the country's ‘safety net’ is the EITC and welfare that is targeted at non- working families has undergone successive reductions in generosity. As a result, non- workers in the US have seen the largest average falls in their net income, which is particularly worrying given this group now accounts for a greater share of the working-age population than in previous decades. In summary, changes in labour market outcomes in Britain and the US have undoubtedly influenced changes in net income inequality in both countries over recent decades. However, the impact of labour market trends has differed between countries both owing to differences in the nature of the trends themselves and the way they have been mediated by the tax and benefit systems of each country. A key difference between Britain and the US we have highlighted is the margin of employment that has been the source of greatest adjustment. In particular, the intensive margin of British male labour supply has become increasingly flexible over the past 20 years with low-wage male workers in particular experiencing large reductions in hours of work. This is in contrast to the US where the greatest change has been the reductions in extensive margin employment, which is somewhat puzzling given the very low level of transfer income available to non-workers in the US. Explaining the reasons for this difference is a key challenge for future research given its implications for welfare and potential possible policy responses."],["This paper uses population register data on inheritances and wealth in Sweden to estimate the causal impact of inheritances on wealth inequality. We find that inheritances reduce wealth inequality, as measured by the Gini coefficient or top wealth shares, but that they increase absolute dispersion. This duality in effects stems from the fact that even though richer heirs inherit larger amounts, the relative importance of the inheritance is larger for less wealthy heirs, who inherit more relative to their pre-inheritance wealth. This is in part driven by the fact that heirs do not inherit debts, which makes the distribution of inheritances more equal than the distribution of wealth among the heirs. Behavioral adjustments seem to mitigate the equalizing effect of inheritances, possibly through higher consumption among the poorer heirs. Inheritance taxation counteracts the equalizing inheritance effect, but redistribution of inheritance tax revenues can reverse this result and make the inheritance tax equalizing. Finally, we also find that inheritances increase intragenerational wealth mobility, but the effect is short-lived. --------------------------------------------------------------------------------","The evolution of wealth inequality and its determinants have received tremendous attention in recent years. After decades of decreasing or relatively low levels of wealth inequality throughout the Western world, wealth inequality may now be on the rise.1 A small but growing body of research has also shown that the importance of inherited wealth has increased recently (Piketty, 2011; Ohlsson et al., 2014). If wealthy children inherit from wealthy parents and inheritances therefore primarily benefit a small elite, there may be a link between increased inheritance flows and increased inequality in the wealth distribution. In this paper, we investigate the impact of inheritances on the distribution of wealth. Although we are not the first to address this issue, it is fair to say that a consensus has not been reached in the literature about whether inheritances increase or decrease wealth inequality. To the best of our knowledge, we are, however, the first to use population-wide individual-level data on both inheritances and wealth to estimate the causal effects of inheritances and characterize the underlying mechanisms. We also contribute by studying the impact of inheritances on wealth mobility and the ways in which inheritance taxation influences wealth inequality. At our disposal is a new population- wide database that contains detailed individual-level information about the estates and bequests of all Swedes who passed away during the 2002–2004 period. Our analysis is based on 168,000 decedents, and of all their family and non-family heirs, comprising 475,000 individuals. The panel dimension of the data allows us to follow heirs and their marketable net worth (which we will hereafter refer to as wealth) for several years—both before and after they inherit. Our identification strategy relies on observing inheritances and wealth distributions for yearly cohorts of heirs. Two different causal effects are identified. First, we estimate a direct mechanical effect (DME), which captures the immediate impact of inheritances, and occurs before any behavioral responses (i.e., before heirs can consume the inheritance). Although we ideally want to evaluate this effect by comparing inequalities just before and just after heirs receive their inheritances, we come close to identifying this effect by comparing wealth inequality at the end of the year preceding the inheritance year, with a measure of post-inheritance wealth inequality, obtained by adding the value of the inheritance to each heir's wealth in the year preceding the inheritance year. The second effect, denoted the behavior- adjusted effect (BAE), shows that heirs may change their behaviors in response to their inheritances, e.g., by consuming or investing part of their inheritances or by working less. We identify this effect by using a difference-in-differences estimator, which compares pre-inheritance inequality with post-inheritance inequality across the three sequentially inheriting cohorts. Heirs who inherit one or two years later serve as the control group for those who inherit in a given year. Note that our focus on heirs only is not very restrictive because everyone will inherit at some point (although a zero amount in some cases).2 This estimation strategy effectively removes biases stemming from macroeconomic events that might influence wealth inequality from one year to the next, as well as biases stemming from the aging of heirs. As pre-inheritance inequality trends are almost perfectly parallel across inheritance cohorts, we are confident in making a causal interpretation of the estimated effects. Our main finding is that inheritances reduce relative wealth inequality. The direct mechanical effect works to reduce the Gini coefficient by approximately 7%. As a point of reference, this decline is about as large as the equalization following the dotcom crash in 2000, when the stock prices of internet companies, presumably owned by the rich, plummeted. Examining different parts of the wealth distribution, we find that the top decile's wealth share decreases substantially, whereas the wealth share of the bottom half increases from a negative to a positive share. While inheritances reduce relative inequality, we find that they increase the absolute dispersion of wealth. This discrepancy between relative and absolute inheritance effects exists because, while wealthier heirs inherit larger amounts, less wealthy heirs receive much larger inheritances relative to their pre-inheritance wealth. Behavioral adjustments appear to dilute the equalizing impact of inheritances. The behavior-adjusted effects are generally smaller than the direct mechanical effects; for example, the Gini coefficient falls by 4% rather than 7%. This equality-diluting effect is consistent with previous research showing that less wealthy heirs spend a larger share of their inherited wealth than wealthier heirs (Druedahl and Martinello, 2017). We are also able to present the first register-based empirical estimates of how inheritance taxation affects wealth inequality, exploiting information about actual individual tax payments.3 The results indicate that the inheritance tax increases wealth inequality, reflecting that less wealthy heirs pay more in taxes relative to their wealth than wealthier heirs do. Still, wealthier heirs pay higher inheritance taxes, but their tax payments are almost always negligible relative to their wealth. However, we show that the redistribution of inheritance tax revenues can reverse this result and make the inheritance tax equalizing. Moreover, we estimate the effect of inheritances on wealth mobility. The welfare interpretation of our inequality results may partly depend on whether heirs switch places in the wealth distribution or retain their ranks after they inherit. We find that, overall, mobility rises substantially, with increased mobility across all parts of the wealth distribution. A series of sensitivity checks suggest that our main findings are robust across several dimensions. First, they do not change when the observed wealth levels are adjusted for potential measurement errors in our wealth and inheritance data. Second, they do not seem to be driven by unobserved inter vivos gifts from wealthy decedents; if anything, adding estimated gifts strengthens the equalizing impact of inheritances. Third, only analyzing inheritances from parents to their children (and neglecting one-third of heirs with more distant family or non-family ties) has a negligible impact on our conclusions. Fourth, we study the importance of young heirs (40 and younger), who could be driving the results because they tend to have relatively little wealth and thus should be affected relatively more by inheriting. While inheritance effects are indeed substantially larger in this younger group, inheritance effects are also important among older heirs. Finally, we exploit parent-child correlations in wealth accumulation and sudden deaths to examine whether heirs adjust their saving behaviors in response to expectations about future inheritances. If such responses were quantitatively important, we would miss a relevant aspect of how inheritances influence the wealth distribution. However, we find no indications of their importance or influence in the data. Our study contributes to the previous empirical literature on the distributional consequences of inherited wealth.4 One group of studies uses simulation methods to model people's savings and giving behavior to calibrate synthetic wealth and inheritance distributions. A sweeping generalization is that these studies tend to find that inheritances constitute a major source of wealth inequality.5 Another group uses individual-level data on people's self-reported wealth and their receipt of gifts and inheritances. The seminal contributions of Wolff (2002, 2003, 2015) and Wolff and Gittleman (2014) use data from the Survey of Consumer Finances to estimate how gifts and inheritances influence the distribution of wealth in the United States. A consistent finding in these studies is that the rich inherit more than the less affluent, but that the rich inherit less relative to their existing wealth, causing inheritances to have an equalizing effect on the distribution of wealth. Similar equalizing effects of inheritances are found in survey data from the United Kingdom (Karagiannaki, 2015; Crawford and Hood, 2016), Japan (Horioka, 2009), Sweden (Klevmarken, 2004) and eight EU countries (Bönke et al., 2017). In a study closely related to ours, Boserup et al. (2016) examine Danish individual-level tax register data on wealth to estimate the effect of inheritances on wealth inequality. The identification of the effect is based on following the wealth of children (45 to 50 years old) before and after the demise of their parents and then comparing this evolution to the wealth of similarly aged children whose parents did not pass away during the study period. The main findings are similar to ours, that is, inheritances cause an increase in the absolute dispersion of wealth and a decrease in the relative wealth inequality. They find larger equalizing effects than we do, although our studies cannot be directly compared with each other. While their approach has several similarities with our BAE analysis, our population is different from theirs in that it includes all adult heirs (not only children). The key difference, however, is that our data contain information about the value of their inheritances, which allows us to estimate the direct mechanical effect and dig deeper into how and why inheritances affect wealth inequality and mobility. It also allows us to study how inheritance taxation affects wealth inequality.6 The remainder of the paper is structured as follows. Section 2 presents the institutional context and the data. Section 3 presents our main findings. Section 4 explains how wealth mobility is influenced by inheritances, and Section 5 discusses the role of inheritance taxation. Section 6 discusses some implications of our findings.","In this section, we present the Swedish legislation regarding inheritances and inheritance taxation. Moreover, we provide descriptions of the data and the study population and discuss the various measures of wealth inequality that we use in the empirical analysis. Inheritance legislation and taxation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In Sweden, when a person passes away, an estate inventory report should be filed with the tax agency, reporting the values of the decedent's assets and debts. If the decedent has a positive net worth, his or her estate is distributed to the heirs according to a succession scheme that is based on genetic relationships. The decedent's relatives are classified into three groups of legal heirs: children and their offspring, parents and their offspring (the decedent's siblings, nephews and nieces), and grandparents and their children (i.e., aunts and uncles).7 Heirs in the second (third) group inherit only if there are no heirs in first (first or second) group. If the decedent has a spouse, the estate is transferred to him or her. If the spouses have common children, the surviving spouse receives what is referred to as free disposal of the estate, which means that the money could be spent but not bequeathed to others than the children. The common children receive the inheritance from the first deceased parent when the second parent passes away. The deceased's children who are not common with the surviving spouse will, on the other, hand inherit immediately when their parent passes away. The default succession scheme can be set aside by a will, but children are always entitled to half of what they would inherit in the case of intestacy, i.e., in the absence of a will. It should be noted also that heirs do not inherit any debts that the decedents may have at the time of death. Inheritance and gift taxes existed in Sweden until their abolishment by the end of 2004.8 In the early 2000s, inheritances exceeding SEK 70,000 (approximately USD 11,000)9 were taxed according to a progressive three-bracket schedule, with marginal tax rates ranging from 10% (paid by heirs who inherited amounts approximately between the 70th and the 90th percentiles in the inheritance distribution) to 30% in the highest bracket on inheritances over SEK 600,000 (USD 91,000, paid by, approximately, the top 2%).10 All inherited assets were taxable, but important concessions were made to keep the effective tax down on certain assets, especially firm equity (see also Ohlsson, 2011 and Henrekson and Waldenström, 2016). Data and study population ~~~~~~~~~~~~~~~~~~~~~~~~~ Our main data source is a population-wide register called Belinda. It originates from the Swedish Tax Agency and contains detailed accounts of the estates of, and inheritances, from all individuals who passed away in 2002–2004 and all of their biological and non- biological heirs. Data are available from this period because the tax agency was obliged to electronically codify all estate reports starting in July 2001, but this obligation was suspended in 2005 when the inheritance tax had been abolished. To these data, we have added information from other administrative registers, primarily those covering personal wealth but also other relevant economic and demographic characteristics for both the decedents and their heirs. In particular, the information about decedents in Belinda includes the net worth at death and its main components (total assets and total debts), the value of the estate, a list of heirs, special rules that apply to the estate and the bequests (e.g., will, prenuptial agreement, and life insurance policy) and personal details (e.g., identity number, marital status, and death date). The information about heirs in Belinda includes the value of their received inheritance, inheritance tax payments (if any), the taxable gifts received over the past ten years, the receipt of life insurance payments from the deceased and personal details (e.g., identity number and relationship to the decedent). Inheritances from a previous decedent (e.g., a late spouse), which the current decedent possessed with free disposal, are divided between the previous decedent's heirs, and the amounts are listed separately in the database. We define inheritance as the total net-of-tax value of inheritances and any insurance received from the decedent (unless it is explicitly stated to be the before-tax inheritance). For heirs who receive two inheritances when the decedent passes away (typically a child who receives one inheritance from the recently deceased parent and one from a previously deceased parent), we define the inheritance as the total sum of these transfers (plus any insurance payments from the two decedents, net of tax). It should be noted that the estates and inheritances observed in the data are reported at tax values, which are sometimes lower than the market values. For instance, real estate was valued at the tax-assessed value, intended to correspond to approximately 75% of the market value. In the main analyses, we use the amounts as given in the database but, in robustness tests, we investigate how the results change when we attempt to adjust the inheritance values to their market values. We define heirs as individuals who live in Sweden and receive an inheritance through the succession order, are beneficiaries of a will, or are beneficiaries of a life insurance policy. We focus on the final estate division of a household and, therefore, do not include heirs who were the spouse or partner of the decedent in our study population. We further restrict our attention to heirs who were at least 18 years old in the year when the decedent passed away because inheritances received by minors fall under the protection of a guardian and are, in practice, controlled by the parents.11 A key feature of our analysis is our classification of heirs into inheritance cohorts according to the year when the deceased passed away. We thus have three inheritance cohorts: 2002, 2003 and 2004, covering a total of 475,120 heirs connected to 168,055 decedents. Wealth data are collected from the wealth register of Statistics Sweden, which is available for the 1999–2007 period, i.e., several years before and after the 2002–2004 inheritance years. The wealth register contains detailed accounts of real and financial assets and debts, all recorded in market values at each year's end, for all individuals in the population. We focus on private net worth, which is the market value of real and financial assets less all debts. Specifically, on the asset side, the wealth portfolios comprise non-financial assets (owner-occupied housing, secondary homes, land, agricultural property, commercial real estate, etc.) and financial assets (bank deposits, listed stocks and bonds, mutual funds and other financial securities). Debts are mainly mortgage loans and state-subsidized loans for higher education. The wealth data are particularly advantageous because the bulk of the records come from third-party reports to the tax agency by financial institutions. The wealth register has limited information about some assets. The register does not cover funded pension assets. In addition, closely held corporations are incompletely covered, and compared with estimates of their aggregate value reported in the Financial Accounts, only about one tenth is accounted for in the wealth register. While these limitations are unfortunate, even when these assets are observed, e.g., in surveys, they are notoriously difficult to value and are, moreover, not always fully marketable. Moreover, consumer durables are not well covered by the wealth register. This may be problematic for an analysis of distributional consequences of inheritances since these goods can be important, not least in relative terms in less wealthy households.12 In the robustness analyses we, therefore, attempt to assess how sensitive our results are to the undervaluation of consumer durables by approximating the value of these goods using, e.g., estimated car values. Despite some shortcomings, it should be noted that our wealth data are the same ones as those used in the international wealth data project, the Luxemburg Wealth Study (see Sierminska et al., 2006). Descriptive statistics ~~~~~~~~~~~~~~~~~~~~~~ This study offers the first comprehensive view of the distribution of estates and inheritances in a population-wide register (see Fig. 1).13 First, we observe that the distribution of the decedents' estates is highly skewed, as most of the mass is located in the left tail and 17% of the estates have zero value. The median value is just over SEK 93,000 (approximately USD 14,000), the mean is approximately SEK 264,000 (USD 40,000) and the 99th percentile of estates is approximately SEK 2.2 million (USD 330,000). The top percentile share accounts for 19% of the total estate wealth, and the top decile accounts for 55%, which are levels that are consistent with those of previous wealth distribution studies (Roine and Waldenström, 2009). Second, the distribution of the inheritances that the heirs receive is similar to that of the estates—skewed, with 19% of the heirs inheriting nothing at all; the top tenth of inheritances represent 56% of the total inherited wealth. Third, the graph in the lower left-hand corner displays the wealth distributions in the year before inheritance (T − 1) for each inheritance cohort. These distributions are nearly identical across the cohorts, highly skewed (with Gini coefficients of approximately 0.8, as examined further in the next section) and show that a non-negligible fraction of the heirs have zero14 or close to zero wealth. Finally, the figure displays the heirs' age distribution. A slight majority of the heirs (56%) are between 50 and 70 years old. Table 1 presents additional descriptive statistics. The inheritance cohorts are nearly identical in all dimensions, which is also expected, as they comprise essentially the entire population of inheriting individuals for each year.15 The average wealth of the heirs one year prior to the inheritance year varies somewhat across cohorts. This variation likely reflects annual differences in macroeconomic conditions, particularly stock market and housing price changes. The bottom panel of the table shows statistics for the decedents. Similar to the statistics for the heirs, the differences are very small, and we thus conclude that the inheritance cohorts are also similar in terms of the characteristics of the donors. Measuring wealth inequality ~~~~~~~~~~~~~~~~~~~~~~~~~~~ The measurement of wealth inequality is somewhat more complex than the measurement of, for instance, income inequality because some individuals have negative wealth (i.e., when debts are larger than assets). Therefore, we conduct our analyses using various unidimensional inequality measures that are defined for variables containing positive as well as negative values.16 Our focus is on the Gini coefficient, which is the most widely used inequality measure. While the statistical properties of the Gini coefficient are fully intact when negative values exist, the normative interpretations from a certain level or trend may be less straightforward (e.g., How should the negative shares of a pie be distributed?). We complement the analysis with other unidimensional inequality measures that can handle negative values: top and bottom wealth shares, wealth percentile ratios, and a measure of absolute dispersion (the interquartile wealth range, and, in the Online Appendix, also the range between the 1st and 99th wealth percentiles, as well as the coefficient of variation).","We estimate two types of inheritance effects on wealth inequality among heirs: one direct mechanical and one behavior-adjusted. Conceptually, the direct mechanical effect (DME) represents the immediate distributional change that arises from adding the inherited amount to the heirs' pre-inheritance wealth. This is our main estimator of interest as it offers the clearest channel from inheritance to inequality change. The behavior-adjusted effect (BAE) accounts for behavioral responses among heirs, which reflect that receiving an inheritance may influence labor supply, consumption and investment decisions that, in turn, may affect wealth accumulation and inequality. The estimations of the two effects are performed both non-parametrically, showing how the distribution changes graphically, and for the different unidimensional measures of inequality. We focus on heirs of all the decedents who passed away between 2002 and 2004. Focusing on heirs is a natural starting point for our study of the distributional consequences of inheritances because almost everyone inherits sooner or later in life, whether the inheritance is a tiny amount (or even zero) or a larger sum. The direct mechanical inheritance effect ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The DME captures how the wealth distribution among the heirs will change if the heirs save their entire inheritances and nothing else happens. To evaluate this effect, we would like to compare the inequality in the wealth distribution in the period just before the heirs inherit to the inequality in the distribution in the period just after the inheritance. Denoting the measure of the wealth distribution of interest DW (e.g., the Gini coefficient), the time of the inheritance T and the length of time until the inheritance ε, the DME on DW would be given by DT+εW − DT−εW. To estimate the DME using this strategy, ε would need to be extremely small (e.g., one day) to avoid the influence of behavioral responses. However, we do not know the exact date when heirs received their inheritances (only the date when the decedents passed away), and we only observe their wealth on December 31 of each year. Comparing wealth distributions in the years before and after inheritance is clearly a too long time span to identify the DME because behavioral responses and changes in macroeconomic conditions may confound the estimates. To examine the statistical robustness of the effect, we compute standard errors by bootstrapping the estimates using 1000 repetitions. The standard errors are typically very small, reflecting both that the DMEs are mechanical in nature, without any stochastic element, and the large size of the dataset. Estimation results: direct mechanical effect We start by presenting a non-parametric estimation of the DME, which evaluates how the density distribution of wealth changes as a consequence of inheritances. Fig. 2 shows how the wealth distribution changes at different wealth levels when we add each heir's inheritance to his or her pre-inheritance wealth. Clearly, a pronounced drop in density occurs around zero wealth, and a sizable increase in density occurs at moderate wealth levels. Thus, heirs with zero (or almost zero) wealth move up in the distribution after having received inheritances. By contrast, no changes appear at very low (negative) and very high wealth levels. In these segments, the densities are similar both before and after inheritance, and the differences in the graph are accordingly quite close to the zero line. In other words, adding inheritances to the heirs' pre-inheritance wealth has the largest influence, quantitatively, on the middle parts of the wealth distribution, whereas the tails are nearly unaffected. We now shift focus to estimate the DME on unidimensional measures of wealth inequality. We seek to quantify the distributional effects of inheritance in terms of standard measures of inequality, which, in turn, facilitates comparisons with other factors and events that affect the wealth distribution. Panel A of Table 2 presents the DME on five unidimensional measures of inequality that were discussed in Section 2.4.17 The estimated effects with respect to these measures mirror the pattern displayed in Fig. 2. First, we see that the relative inequality decreases. The Gini coefficient falls from 0.804 to 0.748 (averaged over the three inheritance cohorts), corresponding to a reduction of 7% - a substantial equalization effect. For example, this drop in the Gini coefficient is larger than that following the dotcom bubble in 2000, when stock prices fell sharply, affecting primarily stockholders (who are typically in the top of the distribution). Similar results, showing an equalizing effect, are found for the other measures of relative inequality. The wealth percentile ratio P90/P50 falls quite substantially—by approximately 17%. The wealth share of the top decile also falls, from 55.9 to 52.3% (a 7% drop), while the wealth share of the bottom half of the distribution actually increases from minus 1.5% to plus 2%. Interestingly, the measure of absolute dispersion, the wealth percentile gap P75–P25, increases by 8%, which confirms the pattern of a high correlation of wealth across generations, in general, and between the heirs' wealth and inheritance amounts in particular. In summary, the results with respect to the DME indicate that inheritances equalize the wealth distribution of heirs, an effect that is consistent across both non- parametric estimation of the change in distribution and a range of well-known wealth inequality measures. In addition, the data confirm the conventionally held view that richer heirs inherit more than poorer heirs do, as indicated by the increased absolute dispersion of wealth. Our results so far show the effect of inheritances on the wealth distribution among all heirs, not only the children of the decedents. To facilitate comparisons with previous studies focusing on children heirs, and to rule out the possibility that our results are driven by including all family and non-family heirs, panel B of Table 2 presents the DME estimates for a sample consisting only of the decedents' children. The point estimates indicate a slightly larger equalizing effect of inheritances and a larger increase in the absolute dispersion among the children-heirs, but the differences between the two samples are small. We, thus, conclude that the choice of including or excluding heirs other than the children of the decedents is not important for our main conclusion that inheritances decrease relative inequality and increase absolute dispersion. In Table 3, we present results from several additional robustness tests. We show that the equalizing effects are more pronounced among the younger (and less wealthy) heirs, but still important across all ages and that our results are robust to the inclusion of heirs younger than 18 years, at the time of the inheritance (see panels A and B, and Online Appendix B.2.1 for details).18 Moreover, we perform several tests intended to assess how sensitive our results are to alternative measurements of wealth and inheritance measures. These tests show first, that the underreporting of consumer durables in our main wealth measures and, second, that using inheritances at market values (rather than at tax values as recorded in the data), have a minimal impact on our results (see panels C–E, and Online Appendices B.2.2 and B.2.3 for details). Finally, we dig deeper into how inheritance of financial and real assets affect inequality in the distributions of financial and real assets (see panels F and G, and Online Appendix B.2.4 for details). Since financial assets are more liquid than real assets, inheritances containing more real estate will affect consumption possibilities in the short run less than inheritances containing more financial assets. We find that inheritances reduce inequality in both financial and real estate assets, but that the effect is more pronounced for financial assets. This follows, as we show in Fig. B1, from the fact that inheritances typically contain equal amounts of real and financial assets and that a large fraction of heirs own real estate while only relatively few (but rich) had substantial amounts in financial assets. How can the equalizing effect be explained? The result that inheritances lead to lower relative wealth inequality is explained by the distribution of inheritances being more equal than the distribution of wealth among the heirs. If the distribution of wealth of decedents is more equal than the distribution of wealth among heirs and all decedents split their wealth equally between their children and have the same number of children, then the distribution of inheritances will also be more equal than the distribution of wealth among the heirs. These conditions appear to be met in our data. The Gini coefficient for the distribution of wealth among the decedents (in the year before the demise) is 0.76 and the Gini coefficient for the distribution of inheritances (net of taxes) is 0.73, while the Gini of the wealth distribution among the heirs (in the year before inheriting) is 0.80. Because wealth is positively correlated across generations (see, e.g., Charles and Hurst, 2003),19 wealthier heirs receive larger inheritances in absolute amounts, but because the distribution of inheritances is more equal than the distribution of wealth among heirs, less wealthy heirs will inherit more relative to the wealth they already hold prior to receiving the inheritance. Fig. 3 shows how the inherited amounts vary with the wealth of heirs. Looking first at absolute amounts (right axis), we see that wealthier heirs inherit more money.20 For example, heirs in the fourth wealth decile (ranked before inheriting) receive inheritances worth, on average, SEK 64,000, whereas heirs in the top decile inherited, on average, SEK 193,000. Thus, there is a positive association between the heirs' wealth and the amount inherited, which explains why absolute dispersion increases. When looking at the relative importance of inheritances instead, dividing inheritances by the heirs' wealth, the pattern is reversed. Heirs in the fourth wealth decile receive inheritances that are larger than their own wealth, whereas heirs in the top decile receive inheritances worth only one twentieth of their wealth. This pattern explains the decrease in relative wealth inequality among heirs. Relatively poor heirs often inherit amounts that are large relative to their own wealth, while this is typically not the case for richer heirs. The equalizing effect can be explained solely by the distribution of wealth among the decedents being more equal than the distribution of wealth among the heirs. Any factor affecting wealth accumulation processes among the donor generation and the heir generation may thus affect how inheritances affect the wealth distribution among heirs.21 Still, inheritances appears to be more equally distributed than the wealth of the decedents, which then contributes to the equalizing effect we find. Several mechanisms may be responsible for why the distribution of inheritances is more equal than the distribution of wealth among the decedents. Below, we address the importance of five such mechanisms. First, if wealthier decedents have more children, their estates will be distributed among more lots, causing each child to inherit less than he or she would have done had there been fewer children.22 However, we find no support for this mechanism. In particular, richer decedents do not have more children than less wealthy decedents and variation in the number of children has no important impact on the results. See panel A in Table 4, and Online Appendix B.3.1 for details. Second, if wealthier decedents testate a disproportionally larger share of their wealth to charities, the heirs would inherit less than they would have done in the absence of charitable bequests. In our data, we see that wealthier decedents indeed testate a larger fraction of their estate to charities. This in line with the literature on charitable contributions at death (see, e.g., Joulfaian, 2001).23 However, even among the wealthiest decedents, only 2.5% of the estate goes to charity. Moreover, a counterfactual analysis, in which we redistribute the charitable bequests to the heirs, produces DMEs that are essentially identical to the main ones. We are therefore confident that charitable bequests among the rich are not the driver behind the finding that inheritances lead to lower relative inequality. See panel B in Table 4, and Online Appendix B.3.2 for details. Third, if wealthier decedents circumvent the default succession rules by writing wills stating that part of their wealth should go to individuals outside the succession order, each heir would inherit less than he or she would have done in the case of intestacy. We address the relevance of this mechanism by calculating the hypothetical inheritance each child would receive in the absence of wills. The DMEs from this counterfactual exercise are largely similar to the main results, suggesting that the equalizing effect cannot be explained by richer decedents' preferences for distributing their wealth among more heirs. The results also suggest that the limited freedom to testate in the Swedish and Roman inheritance law tradition (Pestieau, 2003) is not the driver of the equalizing effect. See panel C in Table 4, and Online Appendix B.3.3 for details. Fourth, intergenerational transfers consist of both inheritances at death and gifts that the decedents give to their heirs during their lifetime, i.e., inter vivos. If substantial amounts were transferred during the years just prior to the inheritance, the interpretation of our results could be misleading. If richer parents were more likely to transfer wealth to their children in the years before the demise, the DMEs would show more equalization than if these transfers would instead take place as inheritances. We conduct several tests, using both data on reported taxable gifts and by assuming that the gift giving patterns in our sample are similar to those in other data sources (Ohlsson et al., 2014; Piketty and Zucman, 2015). We find that decedents with smaller estates make smaller gifts in absolute terms, but that they give away larger shares of their wealth. If we add the value of gifts to the inherited amounts, we find slightly larger equalizing effects. This suggests that the equalizing effects are not much affected by inter vivos gifts. See panels D–G in Table 4, and Online Appendix B.3.4 for details. Fifth, the inheritance law states that the heirs do not inherit the decedent's debts. If the decedent has negative net wealth, the heirs will inherit zero. This feature of the inheritance law makes the inheritance distribution more equal than the wealth distribution of the decedents. If we replace all negative wealth values with zeros and calculate the Gini coefficient for the wealth distribution of the decedents (in the year before the demise), we obtain a value of 0.73, which is akin to the Gini for the inheritance distribution but lower than the Gini for decedent wealth including negative values, which is 0.76. It is hence clear that this part of the inheritance law contributes to the equalizing effects that we find. Out of the five mechanisms discussed above, only the last one appears to be quantitatively important. Another possible explanation for the result is methodological and concerns the identification strategy. If heirs have adjusted their savings in the years prior to inheritance because they expect to receive inheritances, potentially important parts of the total wealth response to inheritances may be overlooked with the strategy. Heirs expecting large inheritances are likely to save less than heirs expecting a small inheritance. As such, the pre-inheritance wealth distribution will be more compressed than in a world in which heirs do not adjust savings decisions based on their inheritance-related expectations. Consequently, the total effect of inheritances—including both pre-inheritance and post-inheritance responses—might be more equalizing than what our estimates suggest. Quantifying expectation responses to inheritances is difficult, and only a few studies have attempted to do so (Wolff, 2015; Elinder et al., 2012).24 We conduct tests designed to assess how expectations about future inheritances may influence heirs' pre-inheritance wealth levels. A first test is based on the idea that if decedents (in the years before the demise) suddenly become richer (poorer) and heirs adjust their savings in response to changes in the expected size of inheritances, we expect that the heirs will respond by dissaving (saving) an offsetting amount of wealth. In a second test, we exploit the idea that heirs may respond more strongly to changes in the decedent's wealth in the years before inheritance if the decedent passes away as a result of a terminal illness rather than passing away suddenly. To investigate this idea more carefully, we use data from the Cause of Death register to identify heir-decedent pairs in which the decedent has passed away suddenly. The classification of sudden deaths (natural and unnatural) follows the classification in Andersen and Nielsen (2011). Neither of the tests provide evidence of responses in the heirs' wealth prior to inheritance. Altogether, the concern that heirs' saving behaviors depend on their inheritance expectations may be plausible, but we find little evidence in our data—or in the previous literature—that these behaviors will confound our main findings. While we clearly cannot rule out that such behavioral effects exist, they do not seem to matter much empirically. See Online Appendix B.3.5 for details. The behavior-adjusted inheritance effect ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ When empirically estimating the BAE on wealth inequality (and mobility), several challenges arise related to concerns about the ceteris paribus condition not being fulfilled. To illustrate the two most prominent challenges, consider first a strategy that compares the wealth distribution of heirs before and after the receipt of inheritance. The difference between the two distributions may be caused by inheriting only, though a singular source is unlikely. For example, macroeconomic events, such as housing market downturns, tend to slash middle-class wealth and thus increase wealth inequality, whereas financial market crashes primarily hit the wealthy and tend to make the wealth distribution more equal (Wolff, 2013; Lundberg and Waldenström, 2018). Second, age-wealth profiles generally imply that, within a birth cohort, wealth becomes more equally distributed with age (Paglin, 1975). Therefore, a simple before-after analysis may yield biased estimates of the effects of inheritances on the wealth distribution. In Eq. (3), Dc, yW denotes the wealth inequality that varies across cohort c and calendar year y. PostInheritance is a cohort-specific indicator variable, which equals one from the year of the inheritance and onwards. We also include year and cohort fixed effects, captured by λy and λc, respectively, and εc, y is a random error term. The estimation model is essentially a difference-in-differences estimator, where the identifying assumption is that the outcome would have evolved similarly for the inheriting cohort and the “to- inherit” cohort(s) in the absence of inheritance (i.e., a parallel trends assumption). While this assumption cannot be explicitly tested, it is possible to obtain indirect evidence of whether it holds by studying the outcome trends for the groups in the pre- treatment period, i.e., before inheritance. In the next section, we investigate the validity of the assumption and show that it appears to hold. Estimation results: behavior-adjusted effect Similar to the DME analysis, we start the BAE analysis with a non-parametrical, graphical illustration. The estimated effect of inheritance, by comparing pre- and post-distributions, may be biased by macroeconomic and demographic (aging) influences; to account for such potential confounders, we compare the wealth distribution changes of the 2002 and 2004 inheritance cohorts between 2001 and 2003, i.e., when the 2002 cohort inherits, and the 2004 cohort does not. Because both cohorts experience the same macro environment and aging process, the differences in wealth distributions effectively only reflect the inheritances of the 2002 cohort.26 Fig. 4 shows the results. An initial observation is that the pattern resembles the one seen for the DME in Fig. 2 (and reproduced here: dashed line in grey); inheritances positively affect a substantial mass of heirs from the bottom and the middle of the distribution. That said, the size of the BAE is apparently smaller than the DME. Next, we turn to analyzing the BAE on unidimensional wealth inequality and absolute dispersion. In contrast to the graphical analysis, we now use the 2002, 2003 and 2004 cohorts simultaneously and exploit the full wealth register data from 1999 to 2007. Therefore, the estimated inheritance effect can capture the average effect up to five years after inheritance, but in practice, most of the variation used in the estimations of the effect comes from the first two years after inheriting, which is why it is safer to say that we capture the effect up to two years after inheriting. As noted in the previous section, the identification strategy assumes that wealth inequality would evolve similarly for all cohorts had they not inherited. Fig. 5 depicts the evolution of the Gini coefficient for the three cohorts over the entire period. Until 2001, i.e., the year before the first cohort inherits, a near-identical development of the Gini coefficient occurs for all three cohorts, strongly suggesting that the parallel trends assumption is fulfilled. The Gini coefficients fall from approximately 0.85 in 1999 to 0.82 in 2000 and 0.81 in 2001. In 2002, the 2002 cohort inherits, and we see an immediate and sharp drop in the Gini coefficient to 0.78, falling further in 2003 when the heirs of the decedents who passed away in late 2002 received their inheritances. By contrast, the Gini coefficients of the two non-inheriting cohorts remain virtually unchanged in 2002. Starting in 2003, when the 2003 cohort inherits, the Gini coefficients of that cohort drop over the next two years. This pattern is repeated again for the 2004 cohort. Between 2005 and 2007, when all the cohorts have inherited, the Gini coefficients return to a common level and development. As clearly shown in Fig. 5, changes in wealth inequality differ across the cohorts only in the two years when they inherit. This strikingly consistent pattern offers strong evidence that the equalizing effect of inheritances on the wealth distribution persist for at least a few years. Table 5 reports the estimation results of the BAE on the five inequality and dispersion measures generated by the difference-in-differences estimator (Eq. (3)). The coefficient estimate in the first column shows that the inheritance effect causes a reduction in the Gini coefficient by 0.037 points, which is equivalent to a drop of 4.6% (when compared with the baseline of 0.804). In Online Appendix B.4.2, Table B5, we show that the BAEs are robust in the dimensions as the DMEs. The estimates of the inheritance effects on the other relative dispersion measures confirm what we have observed already in Figs. 4 and 5, and they are qualitatively similar to the DMEs. The P90/P50 decreases by 10.5%, and the share of total wealth held by the wealthiest decile falls by 5%. Notably, the poorest half increases their share of total wealth from minus 1.6% to just above zero. Finally, the estimated effect of inheritances on the distance between the 75th and the 25thh wealth percentiles indicates that wealth dispersion increases as a consequence of inheritances. Taken together, the results confirm the patterns we found when we estimated the DME, which was that relative inequality decreases while absolute dispersion increases, suggesting that behavioral responses to inheritances do not mute the equalization effect, at least not in the short-run. To be able to firmly assess whether the equalizing effect remains in a longer perspective, we would need more inheritance cohorts that inherit in later years and could serve as longer-run counterfactuals for the cohorts who inherit early. In Online Appendix B.4.3 Fig. B6, we use a slightly different dataset in an attempt to investigate whether the equalizing effects are still present five years after inheriting. The results suggest that the equalizing effect is present at least up to five years after inheriting. Unfortunately, data restrictions imply that we cannot say whether the effect decreases or increases over a period of more than five years. That said, the equalizing effects appear to be smaller when accounting for the behavioral adjustments of the wealth holdings. Therefore, we continue by investigating some possible explanations for this finding in the next section. Potential explanations for why is the behavior-adjusted effect is smaller than the direct mechanical effect? When comparing the two inheritance effects on wealth inequality, we see that the BAE is smaller than the DME. In the case of Gini coefficients, the BAE is 5%, and the DME is 7%. We discuss two possibilities for this discrepancy. The first relates to behavioral responses among the heirs, which may dilute the equalizing effect. The second relates to methodological differences between the DME and BAE. The task to separate the relative importance of the two explanations is, however, beyond the scope of this paper. Why might the equalizing effect of inheritances be less pronounced when behavioral responses are accounted for? Departing from a standard framework for wealth accumulation, several possible explanations are consistent with such a pattern.27 Compared with wealthier heirs, less wealthy heirs have a higher marginal propensity to consume their inheritances (see, e.g., Druedahl and Martinello, 2017). The second explanation is that wealthier heirs receive higher returns on their savings than poorer heirs do (Andersen and Nielsen, 2011). Both of these explanations would lead to increased wealth inequality and, in turn, mitigate the equalizing direct mechanical effect.28 However, without data on consumption, we cannot credibly assess the importance of these two explanations for our findings. The difference between the DMEs and the BAEs could potentially also result from differences in the estimation methods. A key difference between the two estimation methods is that the DME uses wealth data in the year before inheriting and data on the inheritance, while the estimation of BME is based only on wealth data, but for several years. A potential concern here is discrepancies between how wealth and inheritances are measured. While the wealth data are recorded at market values, the data essentially lack consumer durables. The inheritances, on the other hand, include durables but are recorded at tax values. As discussed in Section 2.4, for some asset types, most notably stocks and real estate, the tax value is lower than the market value. However, in Online Appendices B.2.2 and B.4.2, we show that the DMEs and the BAEs are not altered much when we attempt to correct for the underreporting of consumer durables. Moreover, in Online Appendix B.2.3, we showed that the DMEs are not altered when we attempt to adjust inheritances to market values. We, therefore, find no indications that the differences in the DMEs and BAEs stem from these types of measurement errors. However, we cannot rule out that the differences between the DMEs and BAEs are the result of any difference in how the effects are estimated.","Another possible consequence of the effects of intergenerational transfers is that they may influence the level of wealth mobility among the heirs. We are interested in mobility effects for two reasons. First, the normative interpretation of the inequality effects depends on how mobility is affected. Inheritances decrease relative inequality, but do inheritances also increase the chances of the poor to increase their wealth rank? Second, do inheritances set off mobility processes? If receiving inheritances makes people behave differently, perhaps taking more risk or changing their activity in the labor market, inheritances may affect wealth mobility long after the receipt. Even though our previous analysis indicated that behavioral effects exist, mainly working to mute the distributional consequences, perhaps the effect on mobility will still be more evident in this respect? This section presents an analysis—to the best of our knowledge, the first of its kind—on how inheritances influence wealth mobility. Our focus is on intragenerational wealth mobility, which means the rate at which heirs change wealth ranks in their distribution.29 There are several different ways to empirically measure intragenerational wealth mobility as discussed by Burkhauser and Couch (2009) and Jäntti and Jenkins (2015). We show results using one of the most common methods, which involves calculating transition probability matrices in the wealth distribution for heirs before and after inheritance. By comparing transition probabilities across quintiles, the matrix shows whether actual mobility patterns at the bottom, the middle and the top of the distribution differ. We then convert these matrices into a unidimensional metric, the Shorrocks-Prais mobility index (Prais, 1955; Shorrocks, 1978), which is an index centered on the diagonal elements in the matrices and ranges from 0 (perfect immobility) to 1 (perfect mobility).30 For robustness purposes, we also examine the change in Spearman rank correlation coefficients between the same distributions, i.e., heirs' wealth before and after receiving an inheritance. While the difference in Spearman correlations is not a direct measure of mobility, it imposes less structure and relies on fewer assumptions than the Shorrocks-Prais index. The DME for mobility is estimated by computing two transition matrices, one measuring the individual transitions from T − 2 to T − 1 (with the inheritance) and the mobility measure for the transition period from T − 2 to T − 1 (without the inheritance). Table 6 shows that the Shorrocks-Prais index increases by 43% when adding inheritances, which means that mobility increases substantially as a result of inheriting. When evaluating leaving probabilities across wealth quintiles (columns 2 through 6), we see that the mobility effect is larger among heirs in the lower part of the wealth distribution than among the heirs in the top. As a sign of robustness of this effect, the last column of the table shows that inheriting wealth also leads to a significant decrease in the rate of change in Spearman rank correlations across years. Such a decrease reflects a higher degree of rank movements along the wealth distribution when inheriting, which is in line with higher intragenerational mobility. We also estimate the DME on mobility for the Children sample. The results, reported in Online Appendix B.5.2, display a similar pattern to those for the full sample. To estimate the BAE for mobility, we implement the same approach as in the inequality analysis. In Fig. 6, we depict the evolution of Shorrocks-Prais mobility indices for wealth transitions around the year of inheritance (i.e., from Wpre to Wpost). The parallel trends assumption appears to hold, judging from the similar levels of wealth mobility in the pre-inheritance transition periods. Despite an overall rise in mobility from 2001 to 2002, the 2002 inheritance cohort exhibits an even higher mobility increase, from 0.23 to 0.32, than the two other cohorts (which increase from 0.23 to 0.28). One year later, the 2003 cohort experiences a similarly large mobility increase, and, another year later, the 2004 cohort experiences the same relatively large increase in mobility (the mobility increase is marginal in absolute terms, but, because the other two cohorts experience substantial decreases in the same period, the effect can be stated as a relative increase). To be more precise in determining the magnitude of the BAE, we estimate the effect using the difference-in- differences model of Eq. (3).31 The results in Table 7 show that overall mobility increases by 19% (a treatment effect of 0.048 compared with the average pre-inheritance Shorrocks-Prais index of 0.260). Although significant, this effect is less than half the DME reported in Table 6. When examining leaving probabilities across wealth quintiles (columns 2–6), almost no difference in mobility is found across the distribution, with mobility increases between 13 and 21% depending on the part of the distribution that is considered. The change in Spearman rank correlations exhibits once again a similar pattern as the mobility index, showing a decrease in correlation in the periods after inheritances are received, and the effect is also smaller than in the DME analysis. We also estimate the BAE on mobility for the Children sample. The results, reported in Online Appendix B.5.2, display a similar pattern to those for the full sample. In Online Appendix B.5.3, we attempt to see whether the mobility effect is persistent up to five years after inheriting. That analysis suggests, however, that the mobility effect is short-lived; it seems to last only two to three years. Two important messages emanate from the mobility analysis presented here. The first is that inheritances substantially increase wealth mobility. In particular, many of the heirs who were among the poorest before inheriting rise in wealth rank. The second message is that inheritances set off increased mobility in the first years after inheriting, but that effect does not seem to be persistent in the longer run.","In this section, we present the first empirical analysis of inheritance taxation on wealth inequality using individual-level register data.32 The distributional consequences of taxation on intergenerational transfers have received relatively little attention in the previous literature. Theoretical models that address this issue implement diverse analytical approaches, but most of them predict that inheritance (or estate) taxes increase wealth inequality (e.g., Stiglitz, 1978; Becker and Tomes, 1979; Atkinson, 1980; Davies, 1986).33 While these models typically focus on general equilibrium and long-term consequences of inheritance taxation, our analysis instead examines short-term consequences that are associated with the repeal of Sweden's inheritance tax. We begin by examining how the DMEs change due to the tax payments by the heirs. To estimate the effect of the tax, we calculate the difference between the DMEs using inheritances net-of- inheritance-tax payments (as in Table 2) and DMEs using inheritances before tax payments. The estimated effects are economically relevant as the repeal was effectively unexpected, announced just a few months before the repeal actually occurred. However, it should be noted that our analysis is limited by being mechanical and short-term. It does not account for any behavioral adjustments to inheritance taxation that donors or heirs may make over the longer run. The results are reported in Table 8 and should be interpreted as the effect of a repeal of the inheritance tax on wealth inequality. The tax repeal effect suggests that the inheritance tax increases relative inequality but decreases absolute dispersion. However, the magnitudes of these effects are very small. For example, the Gini coefficient falls by an additional 0.002 points due to the tax repeal. This relatively small effect is reasonable, given the rather small amounts of inheritance and gift tax payments in the early 2000s. To explain what causes the disequalizing effect of the inheritance tax, Fig. 7 displays the average level of inheritance and gift tax payments by the heirs' pre-inheritance wealth levels. Wealthier heirs pay more in taxes in absolute terms but less relative to their initial wealth. This finding implies that, for the wealthiest heirs, both their inheritances and inheritance taxes are relatively insignificant in relation to their pre-inheritance wealth, while both inheritances and tax payments are substantial relative to the pre-inheritance wealth of the less wealthy. We thus interpret the results of the test as evidence that the equalizing effect of inheritances would have been slightly stronger without the inheritance tax. The analysis has hitherto only considered the tax payments and not the possible uses of the tax revenues. To facilitate the interpretation of our results and put them into perspective, we show the possible redistributive role that the tax receipts can play. Second, we examine whether our results reflect the specific structure of the Swedish inheritance tax institutions of the early 2000s by assessing what would have happened if a confiscatory tax had been levied instead. Table 9 reports DME estimations under these extensions. In Panel A, we show DMEs under the hypothetical scenario, that the actual inheritance tax revenues are redistributed as lump sum transfers according to three alternative redistributive schemes: giving to all heirs, giving to heirs with below-median wealth and giving to heirs with wealth in the bottom quartile of the wealth distribution. The results indicate that redistribution can strongly counteract the disequalizing effect of the inheritance tax found in Table 8. When the revenues are redistributed among all heirs, the equalizing effect of the inheritances increases (instead of decreases, as it did when we only considered tax payments). Directing revenues to heirs in the bottom half or the bottom quartile of the wealth distribution leads to further equalization. Under all three redistribution schemes, the relative inequality falls more than in the baseline case. The absolute dispersion is also reduced as a consequence of redistributing tax revenues. In Panel B of Table 9, we simulate the redistribution effects under a fully confiscatory tax to determine how much the results are due to the specific institutional structure of the Swedish inheritance tax in the 2000s. Of course, the case of an imagined 100% inheritance tax would most likely have implications for wealth accumulation and the amount of inherited wealth. However, we prefer this scenario for two reasons. First, it represents an upper level for the redistributive impact of inheritance taxation, and milder variants will thus lead to outcomes within this case and the baseline cases. Second, it reflects an interesting counterfactual to our main analysis, namely, the case of literally “no inheritance”, whereas our baseline analysis compares the treatment of inheriting with an “inheriting later” counterfactual. Panel B of Table 9 reports the results from this exercise under the same three redistributive schemes as in Panel A. Redistributing the revenues from a confiscatory tax clearly has a sizable impact on the distribution: giving to everyone almost doubles the equalizing effect of inheritance, from the baseline of −7% to −13%. When applying the directed redistributive schemes, the equalizing effect grows even more, reducing the Gini coefficient by up to almost 22%. In summary, inheritance taxation alone does not seem to equalize wealth; instead, it slightly reverses the equalizing impact of inheritances. However, when inheritance tax revenues are also considered and used for redistributive purposes, the total effect may be increased equality.","Our findings of an equalizing impact of inherited wealth have implications for our understanding of the intergenerational transmission of resources and for economic inequality in general. First, if the poor tend to consume new wealth and the rich are more likely to save it, then the theory predicts the transmission impact to be one of disequalization, as noted by Scholz (2003). Our results are consistent with such behavioral responses, shown by the difference between the larger DMEs and smaller BAEs. However, these responses are not quantitatively large enough to balance the main equalizing impact, and the equalizing impact persists at least a few years after the inheritance treatment. Second, historical circumstances and the institutional context of Sweden may have specific bearing on the detected patterns. Inequality in marketable wealth in a country with an extensive welfare state is not necessarily directly comparable to inequality in countries in which people are more reliant on their own savings. However, since Sweden is no longer exceptional in terms of tax revenues as share of GDP (ranked seventh in 2014), our results could be well generalized to many other countries. It should be noted, though, that a major part of the inherited wealth analyzed was generated during the 1960s, 1970s and 1980s, a period in the Swedish history with peaking egalitarian welfare-state policies and relatively compressed income and wealth distributions. The years thereafter saw both liberalized policies and widening gaps. Could it be that these historical trends in inequality and redistribution are reflected in the equalizing impact of inheritance documented in the study, and thus that heirs not only inherited wealth but effectively also the previous, more equal wealth distribution? At this point, we can only speculate, but such an interpretation would be in line with Nybom and Stuhler's (2014) recent theoretical work on the mechanisms of intergenerational transmission, showing that past institutions and institutional change in a parental generation can have long-lasting effects and eventually affect the offspring generation through the transmission process. On the other hand, our findings are in line with the previous results from Europe, Japan, and the United States, suggesting that inheritances equalize the wealth distribution in many countries. Moreover, the fact that the inheritance law stipulates that heirs do not inherit any debts the decedent may have upon death also contributes to the equalizing effects of inheritances that we find. The Swedish inheritance law is, however, not unique in this feature. In most countries, debts are not passed on to the heirs and even in countries where they are (e.g., Spain and Japan), the heir can refuse the inheritance. Third, our focus on the inequality of personal wealth leaves out other relevant distributional dimensions. One closely related outcome is lifetime resources which is the sum of lifetime earnings and all gifts and bequests. Lifetime earnings are more evenly distributed and much larger than the sum of gifts and bequests, and this could make the distributional consequences of inheritances markedly different from what we observe (although early simulation studies by Blinder (1973) and Davies (1982) do not indicate such marked differences). From a more general perspective, investigating how inheritances affect other aspects of inequality, such as income, leisure, consumption and health, would be interesting."],["The paper reports the result of an experimental game on asset integration and risk taking. We find some evidence that winnings in earlier rounds affect risk taking in subsequent rounds, but no evidence that real life wealth outside the experiment affects risk taking. Controlling for past winnings, participants receiving a low endowment in a round engage in more risk taking. We test a 'keeping-up-with-the-Joneses' hypothesis and find that subjects seek to keep up with winners, though not necessarily with average earnings. Overall, the evidence suggests that risk taking tracks a reference point affected by social comparisons. --------------------------------------------------------------------------------","In spite of a voluminous literature in psychology and economics, risk taking decisions remain poorly understood. This is unfortunate given how critical risk taking is to important economic decisions. We use an original experiment to revisit a key issue that potentially affects risk taking: asset integration. This refers to the idea that individuals decide about risky prospects by considering the effect of decisions on their final wealth rather than on specific gains and losses (Kahneman and Tversky, 1979). We focus on two kinds of asset integration: (1) integration of winnings between successive tasks within an experiment; and (2) integration of winnings with real life wealth outside the experiment. We also examine whether risk taking is affected by social comparisons, while controlling for other factors that may affect risk taking, such as learning and imitation. The study population, partly made of Ethiopian farmers, faces considerable risk in their daily life and is thus well suited to investigate risk taking. Furthermore, because this population is poor, the winnings from the experiment are large relative to their normal income. We thus expect their behavior to be more representative of risk taking by experienced individuals. We also present evidence using university students in the United Kingdom and in Ethiopia. There is evidence that experimental subjects make risk taking decisions ‘as in a bubble׳, that is, ignoring their non-experimental assets.2 Perhaps the most convincing evidence of this is that experimental subjects reject profitable lotteries involving payoffs that are small relative to their wealth (Rabin, 2000; Rabin and Thaler, 2001; Johansson-Stenman, 2009). If participants integrated lottery stakes with their total wealth, only extremely risk averse individuals would avoid small profitable lotteries. One way to solve this paradox is to assume that individuals keep total wealth and experimental income mentally separate, which amounts to a lack of asset integration (Cox and Sadiraj, 2006). This was already implied by the risk attitude estimates of Binswanger (1981) and Gertner (1993), and has most recently been claimed by Schechter (2007). While asset integration has been tested using experimental data alone (e.g., Heinemann, 2008), we are aware of only one published study (Andersen et al., 2011) that directly tests asset integration by combining survey and experimental data.3 They find evidence of only partial asset integration. We revisit this issue using results from a multiple round experiment and test whether people integrate winnings between successive rounds of the same experiment. Participants are divided into groups of six. At the beginning of each round, three receive a high endowment while the other three receive a low endowment. Initial endowments are common knowledge within the group. From this endowment, participants are asked how much they wish to ‘invest’ in a lottery that yields, with equal probability, 0 or three times the amount invested. In the context of the experiment, risk taking is represented by the share of the endowment that players invest in the lottery. We take advantage of the fact that players are faced with the same decision three times to investigate dynamic individual and peer effects. We begin by showing that risk taking within the experiment is uncorrelated with the assets that participants hold outside the experiment, i.e., there is failure of integration with real life assets. For the Ethiopian subject population, we find that winnings from earlier rounds of the experiment increase risk taking. This result does not hold for UK subjects. Taken in combination, these results are consistent with narrow framing: what happens during the experiment is regarded by participants as being in a different frame from their daily lives4. But the size of the frame within the experiment may vary across populations. We also test whether risk taking depends on receiving a low or high endowment in a round. In an expected utility framework, the effect of a high endowment is predicted to be the same as that of past winnings, i.e., positive with an equal coefficient. This is not what we find: participants who receive a high endowment in a round invest a smaller share of it in the lottery than those who receive a low endowment, controlling for past winnings. This suggests that participants who receive a low endowment in a round try to make up for it by taking more risk. This effect can be understood as an application of prospect theory to our experimental setting, and suggests that reference points are affected by the endowment that participants receive at the beginning of each round. We investigate whether reference points are affected by the winnings of peers.5 At the end of each experimental round, participants observe the winnings and investment decisions of other players in their group. We examine whether risk taking is affected by how much others invested and won. There are several possible channels by which these may influence risk taking. We test – and reject – several of these channels. The results we emphasize here suggest that subjects’ risk taking is affected by social comparisons. We find that experimental subjects act as if they are in an implicit competition with each other: when others win big, they take more risk, presumably to try to keep up with them. We also investigate whether participants behave in a way similar to ‘keeping up with the Joneses’ in the consumption domain (Duesenbery 1949; Johansson-Stenman et al., 2002).6. One simple way of modelling this is by assuming that the reference point or aspiration level is a function of what others earn on average: if participants are falling behind the average of their peers, they then take more risk – up to the point where they are above the average.7 This idea is reminiscent of Bault et al. (2008) although their experimental setting is different. We test this prediction and find no evidence of such an effect. Taken together, the results indicate that experimental subjects take experimental earnings into account when deciding how much risk to take: some subject populations (i.e., in Ethiopia but not in the UK) take more risk if they won more in earlier rounds; all seek to compensate for a low endowment in a round by taking more risk; and all take more risk if some players in their group won more than them in earlier rounds. The latter behavior suggests that participants take high earners as reference point when deciding how much risk to take. Risk taking thus appears to have a competitive element, even when participants are quite poor and when the potential earnings from the experiment are large relative to their wealth or income.","We conducted an experiment in Ethiopia in four rural villages, mainly with farmers, and with university students in the capital city, Addis Ababa. As robustness check, the experiment was subsequently replicated in the UK. The details and findings of the UK replication are discussed in the robustness section. The rural fieldwork in Ethiopia was conducted between February and March 2009. The four villages are located in different agro-ecological regions of the country. Subjects were recruited among household heads and their spouses participating in the Ethiopian Rural Household Survey panel. All participants come from farming households; two-third of them are males. The games with university students took place in February 2010. Addis Ababa university did not have a permanent experimental laboratory at the time so students were recruited through ads posted around campus. Information about participants is given in the Data Section. The experimental design is based on earlier experiments organized by Zizzo (2003) and Zizzo and Oswald (2001) in a laboratory setting. Participants are asked to repeatedly make the same risky choice. This choice is framed as an investment decision. This enables us to examine whether choices evolve over time as a function of each participant׳s past winnings and information set. This aspect of the data is the focus of this paper.8 The design of the experiment is as follows. Thirty individuals participate in a session and these players are divided into five groups with six players in each, equally divided into high and low income players. Anonymity within each group is strictly maintained even though the thirty participants in a session can see each other. Each player plays three rounds. At the start of each round players are randomly given either a high (Ethiopian Birr 15) or a low (Birr 7) endowment to induce inequality.9 Each participant then decides how much of this endowment to invest in a more than actuarially fair lottery with a 50% chance of winning thrice the amount invested. Lottery outcomes are independent across participants. After lottery winnings are determined, players are informed of the winnings of the other five members of their group and how much they themselves have won from the lottery. The game is repeated three times.10 In each round new groups are formed with different participants. Players are informed about this. At the end of the game, participants leave with all the winnings accumulated over the three rounds plus a participation fee. This was implemented in four rural villages with a total number of 240 participants, and with 60 university students in the city of Addis Ababa. In addition, a slightly different version of the game was played with another 60 students. In this version, participants stay in the same group of six players over the three rounds. We call this treatment the fixed group treatment.11","We now introduce the econometric testing strategy. After presenting our notation, we explain how we test whether participants integrate their assets or past winnings when deciding how much risk to take. We then introduce social comparisons. At the end of the section we discuss how we address possible confounding effects induced by imitation and learning. Possible confounding effects are discussed in detail in Appendix. Risk taking ~~~~~~~~~~~ A positive a implies increasing relative risk aversion over the narrow range of values taken by Zit, something that is a priori unlikely among poor subjects. It is also difficult to reconcile b=0 with expected utility. We revisit these issues below when we introduce reference points and loss aversion.15 Social comparisons and relative utility ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ According to prospect theory, risk taking behavior depends on whether the decision maker is below or above his/her reference point (Kahneman and Tversky, 1979). Above the reference point, individuals are predicted to behave in the standard risk averse fashion. Below the reference point, individuals may behave in a risk neutral or risk loving manner. There is also a kink at the reference point, generating strong risk aversion when choosing between prospects just above and below the reference point. The largely unanswered question is what the reference point is (Koszegi and Rabin, 2006, 2007). If the reference point responds to what happens to peers, this opens the door to the possibility that risk taking is affected by social comparisons (e.g., Gill and Prowse, 2012). Keeping up with the winners This idea is related to the experimental evidence provided by Bault et al. (2008). But our experimental setting is different. Unlike Bault et al. (2008) and Linde and Sonnemans (2012), in our experiment social comparisons between subjects are not made salient by design. Furthermore, our subjects are not asked to choose between positively and negatively correlated outcomes. Rather they receive endowments that are, by design, negative and positively correlated with others in their group. Having received a different endowment, they are then given an opportunity to risk part of it, a dimension that is not present in these other papers.","In Table 1 we present descriptive statistics on participants from the four rural sites and for university students. Most participants are males but the proportion of males rises to 90% in the case of university students. Unsurprisingly, the average age of rural participants is higher than that of students. On average university participants take more risk: they invest a little over half of their initial endowment in the lottery, which is nearly twice as much as rural players; and the cumulative distribution of investment rates among students is everywhere above that of rural participants. University participants invest their entire endowment in 22% of the games compared to 3% of rural participants. Less than 1% of students invest nothing on lottery compared to 8% of rural participants. Hence, we clearly see higher risk taking among students compared to rural participants. If we assume, as is reasonable in the Ethiopian context, that university students have a higher permanent income, this is consistent with the idea that income affects risk taking. Other factors could also account for this difference – e.g., more risk taking by university students could be explained by higher cognitive abilities (e.g., Dohmen et al., 2010). Since taking risk is profitable in our experiment, it is not surprising to find that the lottery winnings of the students are on average higher than that of rural participants. Rural participants are covered by earlier household surveys from which we recover the value of their household assets. Household assets include agricultural tools like hoes and ploughs, household furniture and items like beds, tables, chairs and stoves and other valuables like jewelry and watches. Rural households hardly use financial assets. The value of household assets is a very good proxy for household income/wealth (e.g., Filmer and Pritchett, 2001). There is a lot of variation in wealth and expenditure within the participating rural population, as is clear from Table 1. Since there is no corresponding survey of university students, there is no information on their household assets. But even if we did have this information, it is unclear how informative it would be: education is probably a better predictor of students’ expected life earnings than whatever assets they may have. For farmers, average winnings from the experiment are equivalent to 1.5% of household assets. Average winnings for students are even larger in absolute terms. This ensures that the experiment provides sufficiently high powered incentives. There is no variation in the educational level of university participants as all of them are in higher education. Rural participants are more representative of the Ethiopian adult population, with much lower education levels. Half of rural participants have no formal education and more than 80% have at most incomplete primary education.23 Although vocational skills may increase agricultural productivity, only 2% of rural participants have any form of vocational training. The heterogeneity of the country in terms of religious beliefs is reflected in the subject population. In both sites, the traditional Ethiopian Orthodox faith is the most common, followed by Protestantism. Muslims are underrepresented compared to the Ethiopian population at large. Asset integration ~~~~~~~~~~~~~~~~~ We report three versions of each regression with different controls. The first version (columns 1, 4 and 7) only includes the above-mentioned regressors plus dummy variables for each of the experimental sites to control for differences in average attributes across sites. The second version (columns 2, 5 and 8) adds controls for whether the composition of the groups was the same across rounds or not, and for whether the experimental session took place in the afternoon – to control for possible mood effects correlated with time of day (e.g., Coates and Herbert, 2008). As further robustness check, the third version (columns 3, 6 and 9) adds individual controls such as age, gender, education and religion which may be correlated with risk taking. To investigate whether our results are an artifact of censoring, we reestimate Table 2 with a tobit estimator that allows for a lower limit of 0 and a variable upper limit Zit. Results, not reported here to save space, are very similar to those in Table 2 in terms of coefficient magnitude and significance. This is hardly surprising given that few observations are at the upper limit of Xit: 4.2% of high endowment observations take value 15 and 14.8% of low endowment observations take value 7. Social comparisons ~~~~~~~~~~~~~~~~~~ We also estimate model (KW) which is linear in Rit. If i׳s reference point is well above the average as in KW, we should only observe the declining portion of Fig. 2. As shown in Table 7, Rit has a significantly negative sign in 5 of the 6 regressions in linear form, as predicted by KW. Taken together, Tables 6 and 7 thus suggest that subjects want to keep up with above average players, that is, with the winners.27 Robustness checks ~~~~~~~~~~~~~~~~~ In addition, we investigate whether our results are an artefact of pooling student and rural subjects. To test the validity of pooling, we reestimate all Tables (except 3 and 4, which only apply to rural subjects) with interaction terms between regressors and a student dummy. Results are shown in online appendix Tables A2b–A10b. In the overwhelming majority of cases the interactive terms are not significant: of all the 135 interactive terms in all the regressions, only 12 are significant at the conventional level of 5%. Our conclusions stand even when we control for possible heterogeneity between students and rural subjects. Replication of the experiment ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As a final robustness check, we replicate the experiment with a different subject population, namely with College students at the University of East Anglia (UEA), United Kingdom. We ran 8 sessions of 12 participants each. Subjects are drawn from the pool of student subjects registered with the UEA experimental laboratory. To maximize comparability, the experimental protocol is as close as possible to that used in Ethiopia, and is virtually identical to that of Addis Ababa students. In particular, we use the same pen-and-paper design. The details of the experimental design are the same as those for the Ethiopia sessions. We took advantage of the replication experiment to take up on a suggestion made by one of the referees and echoed by the editor. In the original experiments, half of the subjects in a group receive an endowment of 7 and the other half received 15. Results show that subjects who receive 7 invest proportionally more than those who receive 15. This finding is hard to explain from expected utility, but can easily be accounted for using prospect theory if subjects have a reference point situated above 7. What is unclear is whether this high reference point comes from observing other subjects receiving 15 – which could constitute additional evidence of a ‘keeping up with the winners’ effect. To test whether the high reference point comes from observing other subjects ‘winning’ a larger endowment, we introduce two new treatments that only differ in terms of the endowment received at the beginning of each round. In the original experiment 3 subjects receive 7 and 3 others receive 15 (normal treatment). In the two new treatments, all 6 subjects in a group receive the same endowment, i.e., either be 7 (low treatment) or 15 (high treatment). By comparing subjects who receive 7 in a low treatment to those who receive 7 in the normal treatment, we can test whether their investment decisions are the same. If it is, this suggests that the high reference point does not come from observing other subjects receiving 15. For the rest, the experiment is the same as in Ethiopia with one caveat: in 6 of the 8 sessions we omit money burning. The purpose of this omission is to check the robustness of our findings to the absence of money burning. To summarize, treatments are as follows: T1, T2 and T3 are the normal, high, and low treatments, respectively – all without money burning. Treatment T4 is the normal treatment with money burning. We observe a number of behavioral similarities in the UEA and Ethiopia subject pools. Risk taking xit is higher for subjects who receive 7 than for those who receive 15 across all treatments, and this difference is strongly significant when we compare low and high endowment subjects within treatments T1 and T4: 56% versus 32% with p-value=0.0000 in T1; and 63% versus 46% with p-value of 0.0007 in T4. When we compare low endowment subjects in the normal (T1) and low (T3) treatments, we find virtually identical risk taking xit: 56% versus 58% with a p-value of 0.64. Similarly, when we compare low endowment subjects in treatments T3 and T4, we find that differences in average risk taking go in the anticipated direction but are not statistically significant: 58% versus 63% with a p-value of 0.22. From this we conclude that the endowment received by other players in the group is not what drives high reference points.","Using data on repeated risk taking in a sequential experiment, we have tested whether participants׳ behavior follows some commonly hypothesized patterns of behavior. Our key findings can be summarized as follows: Asset integration with total wealth: We find no evidence of asset integration between the experimental tasks and real world wealth. Participants apply a narrow framing by which they segregate the set of tasks at hand from their outside wealth . This finding provides support to the intuition of much of the literature and, if anything, is particularly strong evidence of narrow framing given that stakes are large relative to participants׳ normal income and that, unlike Andersen et al. (2011) who find at least some effect, we find no evidence that risk taking responds to wealth. Asset integration across experimental rounds: for the Ethiopia sample there is evidence that winnings from earlier rounds raise risk taking in later rounds; for the UK sample the evidence goes in the opposite direction. High reference point: Within each round, participants who receive a small endowment risk a higher share of it. This finding is difficult to account for under a reasonable expected utility model. But it can be explained if the aspiration level of low endowment recipients is higher than their endowment. For example, under loss aversion and a high enough reference point (due to this social comparison or otherwise), participants who receive a low endowment may seek to make up for it by risking relatively more than if they had received a high endowment. This finding holds in both the Ethiopia and UK samples. ‘Keep up with the winners’: We observe that subjects risk more when other participants they can observe have higher past winnings. We interpret this finding as suggesting that the reference point that subjects use increases in the winnings of others. Hence when others win more, they risk more in an attempt to catch up. This finding holds in both the Ethiopia and UK samples. In Appendix we test and reject the possibility that this may be due to imitation or to gambler׳s fallacy. ‘Keep up with the average’: We only find limited support for it in our experiment. Participants do take more risk when their past winnings are below that of the average of their peers, but not in a way that suggests they regard the average winnings of others as reference point. Combined with earlier results, this confirms that participants seek to keep up with winners, but not with the average. We believe that two of the above results are of particular interest. First, the evidence suggests that participants seek to keep up with the winners. This finding complements existing experimental evidence on social comparisons (e.g., Bault et al., 2008, Linde and Sonnemans, 2012). It highlights the need for further research to ascertain how sensitive social comparisons are to framing and to information about others׳ earnings. Second, for the Ethiopia sample – but not the UK sample – we cannot reject full integration of winnings within the experiment. We also find no evidence of asset integration beyond the narrow frame of the experiment. The Ethiopia findings provide some support for Cox and Sadiraj (2006) and Cox et al. (2008) distinction between total wealth and income and confirm the need to separate the two when estimating risk attitude in applied research. They also raise the question of how to interpret models that link risk attitude with overall wealth and income inequality in a population (e.g., Becker et al., 2005; Hopkins and Kornienko, 2010; Hopkins, 2011). One possible interpretation of these models is that economic agents are in a wealth tournament with everyone else in the population. The evidence presented here suggests that this interpretation may be unwarranted, in the sense that agents may not see themselves as part of a tournament involving their overall integrated wealth, but rather of one involving incomes earned in specific micro-decision environments (such as our experiment). More research is needed to ascertain whether our findings generalize outside our experimental setup."],["What do labor income dynamics look like over the life-cycle? What is the relative importance of persistent shocks, transitory shocks and heterogeneous profiles? To what extent do taxes, transfers and the family attenuate these various factors in the evolution of life-cycle inequality? In this paper, we use rich Norwegian population panel data to answer these important questions. We let individuals with different education levels have a separate income process; and within each skill group, we allow for non-stationarity in age and time, heterogeneous experience profiles, and shocks of varying persistence. We find that the income processes differ systematically by age, skill level and their interaction. To accurately describe labor income dynamics over the life-cycle, it is necessary to allow for heterogeneity by education levels and account for non-stationarity in age and time. Our findings suggest that the redistributive nature of the Norwegian tax-transfer system plays a key role in attenuating the magnitude and persistence of income shocks, especially among the low skilled. By comparison, spouse's income matters less for the dynamics of inequality over the life-cycle. --------------------------------------------------------------------------------","The aim of this paper is to examine the dynamics of labor income over the working life and to explore the impact of two mechanisms of attenuation or insurance to labor income shocks. The first is the tax and transfer system; the second is spouse's income. We focus on three dimensions of inequality: individual market income, individual disposable income, and family disposable income; and explore the relationship between them over the life- cycle. Our objective is to provide a detailed picture of the dynamics of inequality over the life-cycle, following individuals from many different birth cohorts across their working lifespan. By linking up individuals with other family members, we are able to examine the impact of spouse's income and the role of the tax–transfer system as mechanisms to smooth shocks to individual market income. There are a number of key questions addressed. What do labor income dynamics look like over the life-cycle? What is the relative importance of persistent shocks, transitory shocks and heterogeneous profiles? To what extent does the tax and transfer system attenuate these various factors in the evolution of life-cycle inequality? What happens when we add in income sources of spouses? Answering these questions has proved to be quite difficult. One problem that is often argued to hinder analysis is data availability. While the ideal data set is a long panel of individuals, this is somewhat a rare event and can be plagued by problems such as attrition and small sample sizes. An important exception is the case where countries have available administrative data sources. The advantages of such data sets are the accuracy of the income information provided, the large sample size, and the lack of attrition, other than what is due to migration and death. To investigate the above questions, we exploit a unique source of population panel data containing records for every Norwegian from 1967 to 2006. Norway provides an ideal context for this study. It satisfies the requirement for a large and detailed data set that follows individuals and their family members over long periods of their working career. It also has a well developed tax–transfer system, and our data provides us with a measure of income pre and post the payment of taxes and the receipt of transfers. To understand the role of taxes, transfers and the family in attenuating shocks to labor income requires a model that allows for key aspects in the evolution of labor income over the life-cycle. The extensive literature on the panel data modeling of labor income dynamics points to three ingredients of potential significance: shocks of varying persistence; age and time dependence in the variance of shocks; and heterogeneous age profiles. For example, our long panel data allows us to decompose income shocks at every age into two components: one is mean-reverting over short periods (we label ‘transitory’ shocks), while the other could either be permanent or persistent and mean-reverting over long periods (we label ‘permanent’ shocks). By way of comparison, the usual random walk specification would restrict the permanent shocks to be truly permanent rather than merely persistent. Additionally, by following many different birth cohorts across their working lifespan we are able to allow a flexible structure for time effects in deriving our life-cycle profiles.1 The size and detailed nature of the data we are using allow us to explore the importance of these three ingredients for labor income dynamics. Our key findings on the labor income dynamics of males are three-fold. First, the magnitude of permanent and transitory shocks varies systematically over the life-cycle. Indeed, we may strongly reject the hypothesis of age-independent variance of shocks. Second, there is essential heterogeneity in the variances of permanent shocks across skill groups. For low skilled, the magnitude of permanent shocks is monotically increasing in age. For example, a permanent shock of one standard deviation implies a 35% change in individual market income for a low skilled 30 year old; the corresponding number for a low skilled 55 year old is 50%. High skilled, on the other hand, experience large permanent shocks early in life; these shocks decrease in magnitude until age 35, after which they are relatively small and fairly stable. Third, the variance of transitory shocks exhibits a decreasing profile over the life-cycle. While this findings holds for all skill groups, high skilled tend to experience relatively large transitory shocks early in life. The evidence of heterogeneity in the dynamics of labor income by age, skill level, and their interaction motivates and guides our analysis of the insurance from taxes, transfers and the family. We find that the tax–transfer system reduces both the level and persistence of shocks to labor income. In particular, taxes and transfers lead to a remarkable flattening of the age profiles in the variances of permanent and transitory shocks for the low skilled. At age 55, for example, a permanent shock of one standard deviation implies a 50% change in annual market income for a low skilled; the corresponding number for annual disposable income is only 31%. After taking taxes and transfers into account, spouse's income matters little for the dynamics of inequality over the life-cycle. Taken together, our results suggest that a progressive tax–transfer system could be an important insurance mechanism to labor income shocks, especially for low skilled. These results may have implications for both policy and a large and growing literature on consumption inequality and the overall ability of families to insure labor income shocks (see e.g. Blundell et al., 2012). Economic theory predicts that consumption responds strongly to highly persistent or permanent shocks, and empirical evidence suggests little if any self-insurance in response to permanent shocks among individuals with no college education (see e.g. Blundell et al., 2008). Our study points to the importance of understanding the nature of risk that families face over the life-cycle, and the extent to which taxes and transfers crowd out or add to the insurance available in financial markets, the family or other informal mechanisms. Our paper also contributes to the literature on modeling of labor income dynamics. Identification of credible income processes is key for answering a number of important economic questions, including life- cycle consumption and portfolio behavior (see e.g. Gourinchas and Parker, 2002), the sources of inequality (see e.g. Huggett et al., 2011), and the welfare costs of business cycles (see e.g. Storesletten et al., 2001, 2004). The conclusions reached about these questions likely depend on the specification of the labor income process used to calibrate the models. The relatively small scale of the available U.S. panel surveys has forced researchers to rely on simple models that impose economically implausible restrictions (see the discussions in Baker and Solon, 2003; Meghir and Pistaferri, 2011). Using rich Canadian data from 1976 to 1992, Baker and Solon (2003) reject several of these restrictions, including no life-cycle variation in the variance of transitory shocks.2 DeBacker et al. (2013) use a large panel of tax returns to study income dynamics in the U.S. over the period 1987–2009.3 Their estimates point to the importance of allowing for time dependence in the variance components of income. Our study complements these studies by bringing new evidence on several issues pertinent to modeling of income processes. One key finding is that allowing for both age and time dependence in the variance components is essential to accurately describe labor income dynamics. In particular, when restricting the variances of the error components to be constant across the life-cycle, we miss the large permanent shocks that occur late (early) in life for the low (high) skilled. Another key finding is that allowing for heterogeneity by education levels is necessary to capture labor income dynamics of young and old workers. When we restrict the income processes at the variance level to be the same across skill groups, we find a U-shaped age profile in the variances of permanent shocks; however, this pattern is at odds with the age profiles of both high and low skilled.4 By way of comparison, the dynamics of income over the life- cycle change little when restricting the transitory shocks to be uncorrelated over time or allowing for heterogeneous experience profiles within each skill group. Indeed, only for the high skilled, there is evidence of significant unobserved heterogeneity in the income growth rates. Accounting for this heterogeneity lowers the overall persistence of income shocks somewhat, but barely moves the age profiles in the variances of permanent and transitory shocks. The remainder of the paper proceeds as follows. Section 2 presents our data and discusses institutional details. Section 3 describes our panel data specification for income dynamics and presents our findings on the labor income process of males. Section 4 explores the degree of insurance provided by taxes and transfers as well as the income of the spouse. Section 5 offers evidence on several issues pertinent to modeling of income processes. Section 6 concludes. Data and sample restrictions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our analysis employs several registry databases maintained by Statistics Norway that we can link through unique identifiers for each individual. This allows us to construct a rich longitudinal data set containing records for every Norwegian from 1967 to 2006. The variables captured in this data set include individual demographic information (including gender, date of birth, and marital status) and socioeconomic information (including years of education, market income, cash transfers). The data contains unique family identifiers that allow us to match spouses and parents to their children. The coverage and reliability of Norwegian registry data are considered to be exceptional (Atkinson et al., 1995). Educational attainment is reported by the educational establishment directly to Statistics Norway, thereby minimizing any measurement error due to misreporting. More importantly, the Norwegian income data have several advantages over those available in many other countries. First, there is no attrition from the original sample because of the need to ask for permission from individuals to access their tax records. In Norway, these records are in the public domain. Second, our income data pertain to all individuals and all jobs, and not only to jobs covered by social security. Third, we have nearly career-long income histories for certain cohorts, and do not need to extrapolate the income profiles to ages not observed in the data. And fourth, there are no reporting or recollection errors; the data come from individual tax records with detailed information about the different sources of income. We study income dynamics for the 1925–1964 annual birth cohorts during the period 1967–2006. The reason for this selection of cohorts is to ensure fairly long records on earnings for each individual. We restrict the sample to males, to minimize selection issues due to lower labor market participation rates for women in the early periods. In line with much of the previous literature, we exclude immigrants and self- employed. We further refine the sample to be appropriate for our analysis of labor income dynamics. In each year, we select males who are between the ages of 25 and 60. These individuals will likely have already completed most of their schooling and are too young to be eligible for early retirement schemes. In our baseline specification, we further restrict the sample to individuals with at least four subsequent observations with positive market income. This restriction gives us the largest possible sample, given that transitory shocks are assumed to follow a first-order moving average process. Applying these restrictions provides us with a panel data set with 40 time periods and 934,704 individuals. We will refer to this as the baseline sample. On average, this sample consists of 23,368 individuals per birth cohort. Our model estimates age-specific variance components from age 26 to 58. By following many birth cohorts, we are able to allow a flexible structure for calendar time effects in deriving our life-cycle profiles. For the 1942–1946 cohorts, we observe income at every age. For the cohorts born earlier (1925–1941), we miss one or more income observations between the ages of 25 and 41. For the cohorts born later (1947–1964), income is no longer observed at some point after age 42. As a result, our age-specific estimates are based on an unbalanced panel of income. Appendix Fig. C.1 shows the sample size by age. The number of observations declines late (early) in the working lifespan because we are not observing the labor income of younger (older) cohorts at these ages. It is therefore reassuring that the mean and variance of income display similar shapes over the life cycle across cohorts. The income variables that we consider are defined as follows. The first variable is individual market income, defined as the annual pretax earnings.5 The second variable is individual disposable income, incorporating annual earnings and cash transfers net of taxes.6 The third variable is family disposable income. Our measure of family disposable income pools the individual disposable income of the spouses (if the male has a spouse). Throughout the analysis, we partition the baseline sample into three mutually exclusive groups according to educational levels. The reason is that previous studies point to heterogeneity in the dynamics of labor income by educational levels. Low skilled is defined as not having completed high school (32% of the baseline sample), medium skilled includes individual with a high school degree (48% of the baseline sample), and the high skilled consists of individuals who have attended college (20% of the baseline sample). Institutional details and descriptive statistics ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Before turning to the estimation of the income processes, we describe a few important features of our data and the Norwegian setting. We first consider the pattern of labor market participation over the life cycle. Appendix Fig. C.2 shows the population share of males with positive market income by age. We see that the labor market participation rate starts out at around 90% when individuals are young. The participation rate remains at this high level until individuals reach their 50s, at which point they start exiting the labor market at an increasing rate. In particular, low skilled individuals are relatively likely to exit the labor market before they can receive (early) retirement benefits. Appendix Fig. C.3 shows the levels and growth rates in market income by the age at which individuals exit the labor market. We see that early exits from the labor market are associated with low and declining market income in the years prior to exit. Given the sample restriction of non-zero market income, our baseline sample will therefore be of higher quality toward the end of the life-cycle (especially among the low educated). This should put downward pressure on the magnitude of shocks late in life, and most likely give us a lower bound on the insurance from taxes and transfers at these ages. Next, we consider how individual market income varies over the life-cycle in our baseline sample. Fig. 1 shows the age profiles in the different measures of income by education levels. Each market income profile displays the familiar concave shape documented and analyzed by Mincer (1974), but the college-educated workers experience more rapid market income growth early in the working lifespan. Fig. 2 shows the variance of log market income over the life-cycle according to education levels. In line with the prediction of the Mincer model, the variances of medium and high skilled have a U-shaped profile.7 Among low skilled, the variance of log market income is weakly increasing until they are in their mid 40s, after which it rises rapidly. We then examine the extent to which the tax and transfer system affects the mean and the variance of log individual income over the life-cycle. Fig. 1 show how the progressive nature of the tax system dampens the income differentials between high skilled and low skilled after age 35.8 At the same time, low skilled are more likely to receive cash transfers while working (such as partial disability benefits), especially toward the end of the working lifespan. Fig. 2 shows how the tax–transfer system eliminates the increase over the life-cycle in the variance of log market income among the low educated. By comparison, taxes and transfers do less to the large income variance among the high skilled early in their careers. Lastly, we consider how family disposable income varies over the life-cycle in our baseline sample. Fig. 1 compares the age profiles in log family disposable income and log individual disposable income. In the beginning of the working lifespan, relatively few males are married and individual disposable income is therefore quite similar to family disposable income.9 When the males are in their mid 30s, the vast majorities are married and the income of the spouse plays a more important role. At this point, about 80% of the spouses are participating in the labor market, thus contributing significantly to family income.10 Fig. 2 shows the life-cycle variation in the variance of log family disposable income. The family income measure displays quite similar variance over the working lifespan as compared to the measure of individual disposable income. A panel data specification for labor income dynamics ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To understand the role of the tax and transfer system in attenuating shocks to income for individuals and families over the life-cycle requires a model that allows for key aspects in the evolution of labor market income over the working life. As we noted in the introduction, the extensive literature on the panel data modeling of income dynamics has pointed to three key ingredients of potential significance: shocks of varying persistence; age and time dependence in the variance of shocks; and heterogeneous age profiles. The size and detailed nature of the Norwegian population panel allow us to combine all three of these components and let the degree of persistence and the variance of the shocks vary in a quite unrestricted way by age and calendar time. In Appendix A: Estimation Details, we describe every step of the estimation procedure for the income process given in (1). There are, however, three important features to notice. First, we allow for the permanent component vi,ac to have a ρc coefficient less than unity. Since we have long enough panels for individuals in each of the cohorts, the parameters of this process together with those for the transitory process and the heterogeneous profiles can be separately identified. This illustrates that by allowing the variances of each component to differ by age, we are in effect, allowing the autocorrelation of income shocks to vary quite unrestrictedly over the life cycle (even though ρ does not depend on age). Lastly, the use of data that follows actual cohorts over the life cycle allows us to accurately measure their true earnings pattern and estimate the labor income dynamics experienced by individuals. The model given in Eq. (1) is estimated separately by education levels using an equally- weighted minimum distance approach applied to second order moments. At every age, we average the moments across cohorts before estimating the income process.12 Without further restrictions, the estimates can be interpreted as an average (or a typical) labor income dynamics experienced by these cohorts over their working lifespan. Our focus is on life- cycle effects and to determine the relative contribution of age and calendar time effects to the labor income dynamics we require further restrictions. If one were to assume no cohort effects, we would effectively control for calendar time effects by averaging the moments across cohorts. Heathcote et al. (2005) argue that time effects are required to account for the observed trends in inequality. One might, however, suspect that cohort effects should play some role in the distribution of fixed effects. For example, rising college enrollment rates may have changed the level of permanent wage dispersion of younger relative to older cohorts. We incorporate this source of heterogeneity across cohorts by estimating the income processes separately by education levels. Baseline estimates ~~~~~~~~~~~~~~~~~~ We begin by considering the labor income dynamics of males. The model given in Eq. (1) is estimated separately by education levels.13 For now, we impose homogenous experience profiles (βi = 0) in the estimation.14 Instead of presenting the labor income dynamics of each cohort, we average the moments across the cohorts before estimating the income process. As a result, the estimates should be interpreted as an average (or a typical) labor income dynamics experienced by these cohorts over their working lifespan. The first column of Table 1 reports the parameter estimates for individual market income of males. For each skill group, we find that the persistence parameter (ρ) is either one or close to one. This suggests that the shocks to (log) labor income can be described as the sum of a transitory shock and a highly persistent process. Because of the unit root, we do not identify the variance of initial conditions (var(αi)) for the low or medium skilled. The more persistent the shocks, the more important it is to know whether workers at different ages face the same variance of permanent shocks, or if the magnitude changes systematically over the life-cycle. Fig. 3 examines this by showing the age profile in the variance of permanent shocks (var(ui,a)) according to education levels. The magnitude of permanent shocks varies systematically over the life-cycle. Indeed, we may strongly reject the standard specification with age-independent variance components. Another key finding is the heterogeneity in the variances of permanent shocks by education levels. For low skilled, the magnitude of permanent shocks is monotically increasing in age. For example, a permanent shock of one standard deviation implies a 35% change in individual market income for a low skilled 30 year old; the corresponding number for a low skilled 55 year old is 50%. High skilled, on the other hand, experience large permanent shocks early in life; these shocks decrease in magnitude until age 35, after which they are relatively small and fairly stable. For example, a permanent shock of one standard deviation implies a 28% change in individual market income for a high skilled 55 years old; the corresponding number for a high skilled 28 (40) years old is 44 (22) percent. The variance of transitory shocks, shown in Fig. 4, exhibits a decreasing profile over the life-cycle. While this findings holds for all skill groups, high skilled tend to experience relatively large transitory shocks early in life. At the same time, the MA parameter differs by skill group. A larger proportion of the transient shocks persist for another period for high skilled workers than for low skilled workers. To see the importance of low incomes in determining the age profiles of labor market shocks, we present results in Appendix Figs. C.6 and C.7 where we exclude observations with low market incomes.15 The profiles are much flatter. The presence of low market incomes early and late in life is mirrored in hours of work over the life-cycle. When looking at decennial Norwegian Census data over the period of study, we find that mean hours across the life-cycle is inverse U-shaped. There is an increase until individuals are in their early 30s, then a flattening, and eventually a decrease toward retirement. The opposite is true for the variance of log hours, which is U-shaped. In particular, there is a sharp downward trend in the dispersion of hours worked before age 35.16 Model fit ~~~~~~~~~ We now examine the performance of the model given in Eq. (1) in fitting the variance of (residual) income growth rates as well as one-lag covariances. For each income measure and every skill group, we conclude that the baseline specification with homogenous profiles (βi = 0) and a MA(1) achieves a very good fit of these key moments over the life-cycle. Appendix Fig. C.10 shows the model fit for the variance of the growth rate, while Appendix Fig. C.11 displays the match for the one-lag covariance profile of the growth rate. We find that the model matches the variance of the growth rate observed in the data almost perfectly. When ρ = 1, we effectively target the variance of the growth rate in the estimation. As a result, the age dependence of the variances shocks allows us to match the age profile very well. When ρ < 1, the moments used in the estimation differ from those shown in the figure. It is therefore reassuring to find that the model also in this case fits the data very well.","The evidence of heterogeneity in the dynamics of labor income by age, skill level, and their interaction raises a number of important questions. To what extent does the tax and transfer system attenuate or insure the shocks to market income at different parts of the life-cycle? Does the addition of income sources from the spouse offset or enhance labor market shocks? In this section, we investigate these questions. Taxes and transfers ~~~~~~~~~~~~~~~~~~~ The second column of Table 1 reports the estimation results for individual disposable income of males. Importantly, the tax–transfer system reduces the level and persistence of both the permanent and the transitory shocks. The estimated persistence parameter falls the most for low skilled; when ρ = 0.87, the effect of an income shock is reduced to 25% of its initial value in ten years. At the same time, Figs. 3 and 4 show that taxes and transfers lead to a remarkable flattening of the age profiles in the variances of permanent and transitory shocks for the low skilled. At age 55, for example, a permanent shock of one standard deviation implies a 50% change in annual market income for a low skilled; the corresponding number for annual disposable income is only 31%. Shifting attention to the high skilled, we can see that taxes and transfers do little to the age profile in the variance of transitory shocks. As shown in Fig. 4, it exhibits a decreasing and convex profile also in individual disposable income; indeed, the magnitudes of the transitory shocks are only slightly lower for disposable income than for market income. The impact of taxes and transfers is somewhat larger for the variance of permanent shocks. Early in life, the permanent shocks to market income of high skilled are attenuated substantially, although they remain large. Toward the end of the life-cycle, the tax–transfer system reduces the magnitude of the permanent shocks somewhat. Taken together, our results suggest that the redistributive nature of the Norwegian tax–transfer system plays a key role in attenuating the magnitude and persistence of income shocks, especially among the low educated. This finding could have important implications for consumption inequality and the overall ability of families to insure labor income shocks. Economic theory predicts that consumption responds strongly to permanent shocks, and empirical evidence suggests little if any self-insurance in response to permanent shocks among individuals with no college education (see e.g. Blundell et al., 2008). Family income ~~~~~~~~~~~~~ We now shift attention to examining whether the addition of income from the spouse offsets or enhances labor market shocks. There are competing forces at play when going from individual to family income (see e.g. Blundell et al., 2012). The first is that the variance of market income is relatively large among females, reflecting considerable dispersion in hours worked. The second is that the stochastic components of labor income processes are likely to be correlated across spouses. If spouses were adopting perfect risk sharing mechanisms, they would select jobs where shocks are negatively correlated. Alternatively, assortative mating can imply that spouses work in similar jobs, similar industries, and even in the same firm; as a consequence, their shocks could be positively correlated. The third is that family labor supply is a possible insurance mechanism to market income shocks. For example, the wife's labor supply could increases in response to negative income shocks faced by the husband (see e.g. Lundberg, 1985). By comparing the dynamics of individual and family income over the life cycle, we are able to assess the overall impact of these three factors, but not identify the individual contribution of each factor.17 Figs. 3 and 4 display the age profiles in the variances of transitory and permanent shocks to family disposable income. By comparing these profiles to the ones for individual disposable income, we can see that the magnitude of the shocks change little when we add the income of the spouse. This suggests that risk sharing through (negative) assortative mating or labor supply offset the high variance in income of wives and any positively correlated shocks. Table 1 shows that permanent shocks remain highly persistent for low and medium skilled, while falling for the high skilled when we add spouse's income. We can further see that the persistence of transitory shocks change little when including the income of the spouse.18","This section takes advantage of the size and detailed nature of the Norwegian data and brings new evidence on several issues pertinent to the modeling of income processes. Nonstationarity in age and time ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our rich panel data allows us to let the variances components depend on age in an unrestricted way, while controlling flexibly for calendar time effects. This raises questions such as: What is missed by the standard specification in the literature with age-independent variance components? How important is it to account for calendar time effects such as the business cycle or tax reforms? Appendix Table C.1 investigates the implications of assuming age-independent shocks. We display the parameter estimates from a model in which the variances of the error components in Eq. (1) are restricted to be constant across the life cycle. For the high skilled, the estimated persistence parameter falls from almost one with age-dependent variance of shocks to .75 with age-independent variance of shocks. By comparison, the age-independent specification does not affect the estimates of the persistence parameter for the low and medium skilled. Figs. 5 and 6 show the misspecification bias from restricting the variances of transitory and permanent shocks to be constant over the life cycle. These figures highlight the importance of allowing for age nonstationarity to capture the labor income dynamics of low and high skilled workers. What features of the data give rise to the misspecification bias we observe? Recall that high skilled have a U-shaped age profile in the variance of individual market income. We argue that targeting these moments with an age stationary model puts a downward pressure on ρ. With a persistent parameter close to one, it becomes difficult to match the U-shaped profile with an age-independent specification because the permanent shocks would then accumulate over the life cycle, generating an increasing and convex profile in the variance of individual market income. In Fig. 7, we illustrate the importance of allowing for time nonstationarity to get a clear picture of the typical income dynamics over the life cycle. We estimate the model given in (1) separately by cohort, and graph the age profiles in the variance of transitory shocks to disposable income for different cohorts; for brevity, we do not split the sample by education. The 5 year interval between the cohorts allows us to clearly see the impact of a tax reform: For each cohort, we observe a spike in the variance of transitory shocks at that the time of the change in tax policy.19 By comparison, our baseline results control for such calendar time effects by averaging the moments across the cohorts before estimating the income process. After taking out calendar time effects, the variance of transitory shocks exhibits a smooth and decreasing profile over the life-cycle. Heterogeneous profiles ~~~~~~~~~~~~~~~~~~~~~~ Our findings suggest important heterogeneity in labor income dynamics by age, skill level, and their interaction. This raises questions such as: What happens if we do not allow for the possibility that individuals with different education levels face different income processes at the variance level? How important is it to allow for unobserved heterogeneity in the income growth rates within skill groups? Appendix Table C.2 displays parameter estimates from the baseline model of income dynamics when we do not split the sample by education. The persistence parameter in the pooled sample is one, suggesting that the shocks to (log) labor income can still be described as the sum of a transitory shock and a highly persistent process. Figs. C.8 and C.9 show the age profiles in the variances of shocks when we restrict the income processes at the variance level to be the same across skill groups. Because the results from the pooled sample mix the income processes of low and high skilled, we obtain an inverse U-shaped age profile in the variances of permanent shocks. However, this pattern is at odds with the age profiles of both high and low skilled: While the former group experience large permanent shocks early in life, the latter group faces the largest shocks at older ages. These findings point to the importance of allowing for heterogeneity by education levels to capture the labor income dynamics of young and old workers. So far, we have imposed homogenous experience profiles (i.e. βi = 0) within each skill group. We now relax this assumption and allow for a linear experience profile in the model given by Eq. (1). Appendix Table C.3 displays the parameter estimates for individual market income, while Figs. 5 and 6 show the misspecification bias from imposing homogenous experience profiles. The results suggest education levels do a good job in capturing heterogeneity in the dynamics of labor income over the life-cycle. Only for the high skilled, there is evidence of significant unobserved heterogeneity in the income growth rates; accounting for this heterogeneity lowers the persistent parameter from .98 to .90, but barely moves the age profiles in the variances of permanent and transitory shocks. Appendix Fig. C.12 illustrates the heterogeneity in market income profiles for the high skilled. There is a non-negligible fanning out of the income profiles. At the same time, there is a negative correlation between the initial conditions and the individual-specific income growth rate. This means that high skilled workers with relatively low market income at age 25 (the initial age) tend to have stronger income growth over the life cycle, offsetting some of the fanning out displayed in Fig. C.12. Serially correlated transitory shocks ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Because we have long panel of individuals, we can separately identify a transitory process with serially correlated shocks and a permanent process which allows for a persistence parameter less than unity. In many cases, however, this is difficult because the panel of individuals is too short (or plagued by problems such as attrition and small sample sizes). In Figs. 5 and 6, we examine the implications of restricting the transitory component in the model given by Eq. (1) to be uncorrelated over time. In the simple case of serially uncorrelated transitory shocks, all the persistence in the income data is attributed to the permanent income component. By comparison, the transitory component is assigned a larger share of the total variance in our baseline model, because the process captures short-duration persistence in the data. However, our estimates suggest the misspecification bias from assuming serially uncorrelated transitory shocks is relatively small compared to the biases from ignoring heterogeneity in labor income dynamics by age and skill level.","What do labor income dynamics look like over the life-cycle? What is the relative importance of persistent shocks, transitory shocks and idiosyncratic trends? To what extent do taxes, transfers and the family attenuate these various factors in the evolution of life-cycle inequality? In this paper, we used rich Norwegian panel data to answer these important questions. We estimated a process for income dynamics that allows for key aspects in the evolution of labor income over the life-cycle, including non-stationarity in age and time, shocks of varying persistence, and heterogeneous profiles. Our estimates of the labor income dynamics of males showed that the magnitude of permanent and transitory shocks varies systematically over the life-cycle, and that there is essential heterogeneity in the variances of these shocks across skill groups. We found that the redistributive nature of the Norwegian tax–transfer system plays a key role in attenuating the magnitude and persistence of income shocks, especially among the low skilled. Spouse's labor market income, on the other hand, matters less for the dynamics of inequality over the life-cycle. The size and detailed nature of the data we are using also allowed us to bring new evidence on several additional issues pertinent to modeling of income processes. One key finding was that restricting the age and time dependence of the variance of income shocks can lead to quite misleading conclusions about the income process. Another key finding was that allowing for heterogeneity by education levels is necessary to capture the labor income dynamics of young and old workers. By way of comparison, the dynamics of income over the life-cycle change little when restricting the transitory shocks to be uncorrelated over time or allowing for heterogeneous experience profiles within each skill group. Indeed, only for the high skilled, there is evidence of significant unobserved heterogeneity in the income growth rates. Accounting for this heterogeneity lowers the persistence of permanent shocks somewhat, but barely moves the age profiles in the variances of permanent and transitory shocks."],["We use firm-level data from Hungary to estimate knowledge spillovers in importing through fine spatial and managerial networks. By identifying from variation in peers' import experience across source countries, by comparing the spillover from neighboring buildings with a cross-street placebo, and by exploiting plausibly exogenous firm moves, we obtain credible estimates and establish three results. (1) There are significant knowledge spillovers in both spatial and managerial networks. Having a peer which has imported from a particular country more than doubles the probability of starting to import from that country, but the effect quickly decays with distance. (2) Spillovers are heterogeneous: they are stronger when firms or peers are larger or more productive, and exhibit complementarities in firm and peer productivity. (3) The model-implied social multiplier is highly skewed, implying that targeting an import-encouragement policy to firms with many and productive neighbors can make it 26% more effective. These results highlight the benefit of firm clusters in facilitating the diffusion of business practices. --------------------------------------------------------------------------------","Imports have large positive effects on firm productivity (Amiti and Konings, 2007; Halpern et al., 2015), yet there is much heterogeneity in similar firms' importing behavior. One explanation for this heterogeneity is the presence of informal trade barriers, when specific knowledge or a trusted partner is needed for a productive import relationship. When informal barriers are active, importing may diffuse from firm to firm through personal and business connections. Mion and Opromolla (2014), Mion et al. (2016), Fernandes and Tang (2014) and Kamal and Sundaram (2016) document such diffusion for exports, but at present we have limited evidence on the—equally important— import side of the market.1,2 Are there knowledge spillovers in importing? If there are, what factors facilitate or limit diffusion? The answers can shed light on the puzzling cross-firm heterogeneity in importing and its productivity benefits; and can guide trade policy to exploit indirect effects. In this paper we use firm-level data from Hungary to document and analyze knowledge diffusion in importing. In doing so, we make three main contributions. First, we develop a portfolio of empirical designs which rule out many alternative explanations and help advance the identification of trade spillovers in spatial and managerial networks. We address firm heterogeneity by identifying from source country variation, exclude spatial omitted variables by exploiting the precise neighborhood structure, and also use plausibly exogenous firm moves. We consistently find significant spillover effects. Second, we investigate the factors associated with stronger diffusion. We find that knowledge flows are stronger when firms or peers are larger or more productive. Knowledge flows also exhibit complementarities in firm and peer productivity, showing that positive sorting can increase the overall adoption of importing. Third, we demonstrate in a counterfactual analysis how network density and positive sorting combine to shape adoption patterns. We document that the model-implied social multiplier of importing is highly skewed in the number and type of peers, implying that import subsidies targeted at firms in buildings with many productive neighbors are much more effective. In Section 2 we present our data. We use a firm-level panel that contains rich information about Hungarian firms during 1993–2003. We combine three data sources: the Hungarian firm register, balance sheet data from the National Tax and Customs Administration, and trade data from the Hungarian Customs Statistics.3 The firm register contains, for the full universe of Hungarian firms, the precise address of the firm, all owners with their country of origin, and all firm officials with signing rights, as well as changes over time. As a result, we can trace changes in spatial and ownership links and the moves of people. The balance sheet data include additional information on the foreign ownership share and the industry of firms. And the customs data contain annual export and import flows at the HS6 product level for each firm, separately for each destination and source country. Section 3 presents our first main contribution: the empirical strategy and results on import spillovers. The key identification concern with estimating spillovers is one common to studies of peer effects (Manski, 1993): that a firm and its peer's import choices may be correlated for reasons unrelated to learning. For example, firms in a particular industry may make correlated location and import decisions. We address this endogeneity problem using two main research designs exploiting progressively narrower sources of variation, in combination with placebo tests and sample definition choices that rule out several omitted variables. Our first research design is a linear probability model measuring the effect of peer firms'country-specific experience on a firm's decision about starting to import from the same country. We implement this design by including firm-year and country-year fixed effects, effectively exploiting variation within a firm in a given year: we ask if having a peer which has past experience with a given country increases the probability of starting to import from that country, rather than from another country. To increase comparability we only look at four source countries similar in terms of imports: the Czech Republic, Slovakia, Romania, and Russia. And to ensure that all firms are the same distance from the border we only consider firms located in Budapest. We use this research design to estimate knowledge diffusion in two networks: close spatial neighborhoods and managerial networks. Within spatial neighborhoods we consider three types of peers: firms in the same building, firms in the two neighboring buildings, and, as a placebo, firms in the two closest cross-street buildings. In managerial networks, we define peers as firms from which an official with signing rights has moved to the firm of interest. To limit confounding effects we always exclude ownership-connected firms—defined as those which share an ultimate owner with the firm of interest—from the spatial and managerial peer groups. Our first design yields significant positive import diffusion estimates in both networks. For neighborhood networks we document highly spatially localized spillovers. Having a same-building peer with import experience from a specific country increases the probability of starting to import from the same country by 0.2 percentage points, which roughly doubles the baseline probability of starting to import from one of the four countries. The effect of a neighbor-building peer's import experience is only one-fifth as large, indicating fast decay by distance.4 The placebo effect of a cross-street peer's import experience is insignificant and small. Finally, in managerial networks the same design yields spillover estimates which are twice as large as the same-building effect. This design addresses several omitted variable problems which often plague estimates of knowledge diffusion. Most directly, by exploiting variation across source countries it addresses the basic concern that importers tend to be connected to other importers. Specifically, in the absence of source country variation the firm-year fixed effects would soak up all the variation in peers' import experience.5 In addition, our controls and placebo also address more subtle country-specific omitted variables. In particular, by controlling for ownership links we remove omitted variables based on joint ownership. Results below show evidence on diffusion across industries, addressing concerns with same-industry clustering. And, most important, the neighboring building versus cross-street building comparison rules out any remaining omitted variable as long as knowledge spillovers decay faster than the spatial correlation in that variable. One remaining concern with our first design is that it does not make explicit the source of variation in peer firms' experience, and therefore it may be subject to some unspecified—highly spatially concentrated—omitted variable. In our second design we address this problem by exploiting a concrete plausibly exogenous source of variation: firm moves. We conduct an event study of the impact of firms with country-specific import experience moving into an address where no such experience was present earlier. The move is a positive shock to local country-specific knowledge. We show that firms located in such an address start to import from the country known by the mover with a higher probability than from other countries, relative to firms in addresses where the mover had no such experience. Consistent with the logic of diffusion, the response of imports to the move is gradual. The magnitude of the estimate is comparable to that of our first research design. The consistency of the results identified in different networks and from increasingly narrow sources of variation further supports the knowledge spillovers interpretation. In Section 4 we present our second main contribution: the heterogeneity of the spillover effect. We explore heterogeneity both to internally validate our estimates and to obtain lessons about mechanisms. We measure heterogeneous effects both by the characteristics of the firm and those of the peer, as well as their interactions. Focusing on same-building peers, we find that larger, more productive and foreign-owned firms benefit more from peers' import experience. Firms also learn more from peers which are larger, more productive or foreign-owned. And spillovers are also stronger when more peers have import knowledge. These results are all consistent with the knowledge diffusion interpretation: better firms are likely to be both more receptive to information and more effective in passing it on, and multiple sources should further increase the rate of diffusion.6 We then document that the strength of the spillover also exhibits complementarities between the firm's and the peer's characteristics. We show that high- productivity firms tend to learn even more from higher-productivity peers than low- productivity firms do. Similarly, we show that the effect of peers operating in the same industry or importing the same product category is significantly larger than that of other peers. At the same time, spillovers from peers operating in different industries or importing different product types are still significant. The results on complementarities are potentially relevant because they suggest that positive sorting—even holding fixed the network structure—can generate aggregate gains in importing. In Section 5 we present our third main contribution: a counterfactual analysis to assess the policy implications of the estimated import spillover effect. Our results so far imply that spillovers should be stronger when (i) the number, and (ii) the productivity of experienced peers is higher. To quantitatively evaluate the combined impact of these forces, we compute the model-implied social multiplier effect on imports of a firm entering into an import market, which incorporates spillovers over the next five years. We calculate the multiplier using the same-building estimate which accounts for heterogeneity by the productivity of the firm and its peers, and also allows for an increase in spillovers with the number of experienced peers. Because the number and productivity of peers varies across the sample, we obtain a separate multiplier for each firm which has not imported yet from one of the four countries. The results show substantial skewness in the social multiplier. In particular, we find that the five-year social multiplier is 1.03 for the median firm and 1.13 for the firm in the 90th percentile. Thus, while accounting for spillovers is not important for the typical firm, it is potentially quite important for a substantial share of firms. An implication is that there may be significant gains from targeting trade policies. We confirm this by showing that a targeted import subsidy policy treating firms for which spillover effects are the largest can be 26% more effective than a non-targeted one. Because finding the firms with the highest expected indirect treatment effect only requires public information on firms' balance sheet and address, this targeting is in principle directly implementable. Our result quantifies the benefit of clusters—especially of firms with high productivity—in facilitating the diffusion of good business practices. We build on a literature on knowledge spillovers in trade, most of which studies the diffusion of exporting. An important part of the literature explores spatial spillovers. Early work focused on the diffusion of the decision to export, and obtained mixed results.7 More recent work studies the diffusion of specific knowledge, such as export experience with a particular country or product, and generally finds evidence for spillovers (Koenig, 2009; Koenig et al., 2010; Poncet and Mayneris, 2013; Castillo and Silvente, 2011; Ramos and Moral-Benito, 2013; Mayneris and Poncet, 2015). Using uniquely rich data on trade partners Kamal and Sundaram (2016) document the diffusion of concrete export partners. And Fernandes and Tang (2014) document export spillovers using for guidance a formal model that allows them to test specific predictions of the learning hypothesis. All these papers define spatial neighborhoods to be cities or similarly large agglomerations. Our spatial spillover results improve identification by using substantially more precise measures of neighborhoods. When networking benefits decay rapidly in space (Arzaghi and Henderson, 2008), spatial networks should be measured at a fine resolution to avoid confounding variation from omitted spatially correlated variables. Our results show that spillovers do decay fast, highlighting the relevance of our precise measures. More broadly, we also contribute to this literature by our focus on imports, our analysis of heterogeneous effects and the implications for targeted trade policies. Another part of the export spillovers literature studies spillovers through managerial moves. These papers show that having a manager with prior export experience join the firm increases the likelihood that the firm starts to export (Choquette and Meinen, 2015; Mion and Opromolla, 2014; Mion et al., 2016; Sala and Yalcin, 2015; Masso et al., 2015). We contribute to this work by focusing on import spillovers; by having a comprehensive study in which we compare spillovers in managerial networks to spillovers in spatial networks; and by our analysis of heterogeneous effects and the implications for targeted policies. Given this existing work on export spillovers, our main focus in this paper is the more novel and equally important topic of import spillovers. There is almost no work on this topic, the sole exceptions—to our knowledge—being Harasztosi (2011) and Harasztosi (2013), which estimate import spillovers in Hungarian NUTS4 agglomeration units. Our contribution to this work is the use of more precise neighborhood definitions, a variety of empirical designs that limit confounding factors, a more comprehensive analysis of multiple networks, the results on heterogeneous effects, and the policy counterfactual analysis. Finally, we build on a literature on firm networks and diffusion in networks. Chaney (2014) develops a model in which firms can acquire trading partners through existing contacts; Fafchamps and Quinn (2015) and Cai and Szeidl (2018) show that managerial meetings can facilitate the diffusion of business relevant information; and Banerjee et al. (2013) explore network-based targeting of microfinance in the presence of knowledge diffusion. Our study documents and analyzes these sort of network effects in the novel and important context of import spillovers. Data sources ~~~~~~~~~~~~ We create our panel of Hungarian importers by combining data from three sources. Firm registry 1993–2003 Data from the Hungarian Company Register contain basic information for the full universe of Hungarian firms, including the firm's name, tax identifier, and precise address: zip code, city, street, number, floor and door number. These variables have associated start and end dates, allowing us to track firm moves over time. The registry data also contain information about the firm's owners, and officials with signing rights which include directors, board members, the CEO, and some employees. As the employees with signing rights are usually at or near the top of the firm hierarchy, we sometimes—slightly imprecisely—refer to these people as managers. For firm owners the data contain the name and registry number; and for person owners and officials the name, mother's name and home address. These records also have start and end dates. We use the name, mother's name and address to create an anonymous unique identifier for each individual in the data. We use this identifier to track individuals across firms and over time. Our method allows for typos and slight variations in names, such as omitting the middle name. Balance sheets 1993–2003 We have balance sheet data for all double-bookkeeping Hungarian firms from the National Tax and Customs Administration of Hungary. These data also include the firm's industry at the 2-digitNACE level (Revision 1.1), and the shares of its capital owned by foreign entities, domestic private entities, and the Hungarian state. International trade 1993–2003 Detailed firm-level trade data come from the Hungarian Customs Statistics. These data contain yearly exports and imports by each firm to and from each foreign country at the Harmonized System (HS) 6-digit product category. The reason that our sample period ends in 2003 is that the firm-level trade data are not available for later years. We use unique firm identifiers to link these three datasets. Firm sample We focus on imports from four countries that are comparable in terms of their exports to Hungarian firms: the Czech Republic, Slovakia, Romania and Russia. To avoid variation in distance from the border, we use only firms with headquarters in Budapest, which account for over 20% of all the firms in Hungary. Accordingly, when a firm moves its headquarters out of Budapest, we let it exit from our sample. These exclusions result in our main firm sample which contains 211,598 firms and 1,189,402 firm-year observations. We conduct most of the analysis using our analysis sample, a (firm, source country, year) panel derived from our main firm sample. In this three-way panel we only include observations in which a firm in the main sample has not yet imported from the given source country up until the previous year. This sample construction allows us to estimate the probability that a firm starts to import from a particular country for the first time. We also make three additional exclusions. (1) We exclude firms for which the headquarters' address is missing, because for them we cannot define spatial networks. (2) We exclude firms which have >50 same-building peers to ensure that our results are not driven by large hubs. (3) We start the data in 1994 because separate trade data for the Czech Republic and Slovakia are only available starting 1993 and the analysis sample requires importer status of peers in the previous year.8,9 After these exclusions, the analysis sample contains 88% of the firms in the main firm sample and has 3,778,517 firm-year-country observations. About 5% of the firms in the main sample import from at least one of the four countries at least in one year during 1993–2003. Variable definitions We define the firm to have import experience with a country in a year if it has imported from that country in that year or in a previous year. This definition captures the idea that the firm has acquired import experience specific to that country by that year. We define experience with exports or with foreign owners in an analogous way. We classify a firm as foreign-owned in a year if it had majority foreign ownership that year. We classify imported products by their purpose using the Broad Economic Categories (BEC) classification. We create four product categories: Consumer goods (BEC 1, 6), Industrial supplies (BEC 2, 3), Capital goods (BEC 41, 51, 52) and Parts and accessories (BEC 42, 53). Using the Levinsohn and Petrin (2003) methodology we estimate from the balance sheet data total factor productivity (TFP) for each firm in each year, assuming a Cobb-Douglas revenue production function with capital and labor as factors and materials as an input, allowing coefficients to vary by two-digit industries. We normalize log productivity within each 2-digit industry to have mean zero in our main firm sample. We then assign firms to productivity quartiles in each year t, based on the average of their yearly 2-digit-industry-specific productivity percentile over the years t − 2, t − 1 and t. Taking the average over three years reduces noise, and results in a smooth but time-varying productivity index. Firm networks ~~~~~~~~~~~~~ A key ingredient in our analysis is data on peers in firm networks. We work with three classes of peers, defined based on spatial, personal and ownership connections. Spatial peers We use a highly localized definition of spatial connections. We create three different spatial peer groups. (i) Same-building peers, defined as firms with the same street address up to building number. (ii) Neighbor-building peers, defined for a firm with building number n as firms in buildings in the same street with numbers n − 2 and n + 2.10 (iii) Cross-street peers, defined as firms in buildings in the same street numbered n − 1 and n + 1. From all three peer groups we exclude firms which have an ownership link—as defined below—to the firm of interest in the given year. Because the address data has dates, all these peer groups are year specific. Person-connected peers We define a firm B to be a person-connected peer of firm A in year t if some person X is an official with signing rights of firm A in year t and was an official with signing rights of firm B at some earlier date. We will often focus on person connections that can transmit import experience with some country c, which happens when firm B had import experience with c before person X left that firm. In all person-connected definitions we exclude people with signing rights who are liquidators—officials assigned to handle liquidation of the company—as well as people who are officials or owners of >15 firms in the given year. We also exclude from the set of person-connected peers firms which are likely to have shared decision makers with the firm of interest: those ever connected to the firm through ownership links (as defined below), and those that have the exact same address including floor and door number. But we do include peer firms which are located outside Budapest. With slight imprecision, we sometimes refer to the person-connected network defined this way as the managerial network. Ownership-connected peers We classify firms A and B to be linked by ownership in year t if they have a common ultimate owner. This includes two types of connections: (1) when A and B have a direct or indirect common owner; (2) when one of the firms is a direct or indirect owner of the other. We also include peers located outside Budapest in the ownership-connected peer group of a firm. Summary statistics ~~~~~~~~~~~~~~~~~~ Table 1 presents descriptive statistics on the firms in our main sample. The first column refers to all firms in all years, the second column to firms in years in which they have already had import experience from one of our four source countries, and the remaining columns to firms with import experience from specific countries. Comparing between columns 1 and 2 shows that importers are on average older, larger, more likely to be foreign owned, more likely to export, and have higher productivity than the industry average. These patterns are familiar (Bernard et al., 2009). The remaining columns show that importers from the four countries of interest are fairly similar in terms of all the variables in the table, consistent with our intuition that these source countries are roughly similar in terms of their associated import barriers. Table 2 shows the number of firms and importers over time during our sample period. The rapid increase in the number of firms is likely due to the development of the capitalist economy in the 1990s. And the increase in the number of importers is probably a consequence of several factors: more firms, lower formal trade barriers, and a country more deeply embedded in the international economy. The considerable increase in importing shown in the table is a key source of variation for our analysis below. Table 3 reports the distribution of degree (number of peers) in the different firm networks. The average degree—shown in the bottom row—is the highest for the same-building network (8.4) and the lowest for the the person- connected network (0.3). The neighbor-building and cross-street networks are between these two extremes (average degrees of 5.2 and 3.3) and although the latter is more sparse, have a roughly similar degree distribution. In all networks a substantial share of firms are isolated, i.e. have zero neighbors. This heterogeneity in degree across firms is one key reason for our finding below that targeting import subsidy policies can substantially increase their effectiveness.11","This section presents our empirical strategy and results on the effect of peers' experience on a firm's import decision. Our main hypothesis is that importing requires source-country specific knowledge, which in turn diffuses in various firm networks. As a result, we predict that firms which—other things equal—have peers with experience importing from a particular country are more likely to start importing from that country. We divide this section into four parts. We begin by presenting motivating evidence which highlights a key component of the logic for identification: variation in peers' import experience across different source countries. We then present two empirical designs. The first design directly exploits this source country variation, and yields spillover estimates in both spatial and managerial networks as well as placebo estimates that confirm the logic of identification. The second design further improves identification for spillovers in spatial networks by exploiting plausibly exogenous firm moves. In the final part we assess the magnitude of our spillover estimates. Motivating evidence ~~~~~~~~~~~~~~~~~~~ Table 4 shows how we exploit source country variation in peers' import experience. The table reports the probability of a firm starting to import from a particular country in a year, conditional on it starting to import from one of the four countries that year, and conditional on different importing patterns of its peers. The four panels correspond to peers defined by the same-building, neighbor-building, person-connected and ownership- connected networks. Within each panel, the top row shows the share of firms which start to import from a country c, while the bottom row shows the share which start to import from a different country.12 The left column computes this share for firms with peers that have import experience with c but not the other countries; and the right column for firms with peers that have import experience with a different country but not c. We report the average share when c runs across the four countries, weighted by the number of observations per country. The table shows that in each network, the share of firms starting to import from country c is always higher when peers have c experience than when peers have non-c experience. This fact suggests that peers' experience influences firms' import decisions and forms the basis for our identification strategy. We now turn to more fully develop this empirical approach and derive statistical inference, explicitly address confounds, conduct placebo analysis and incorporate additional plausibly exogenous variation. Research design 1: Peers'country-specific import experience ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Here i indexes firms, c indexes source countries and t indexes years, thus each observation is a (firm, source country, year) triplet. We estimate the regression in our analysis sample, which contains observations where firm i has not yet imported from country c before year t. The left-hand-side variable Yict is an indicator for i importing from country c in year t. Given that the sample excludes prior importers from c, Yict measures entry into importing from c. On the right-hand side we include indicators for the presence of country-specific import experience in various peer groups n. Specifically, Xnic, t−1 is and indicator which equals one if there is at least one firm in firm i‘s peer group n in year t − 1 which has import experience from country c at time t − 1, that is, which imported from c in t − 1 or earlier.13 We use lagged peer experience because we expect information diffusion to take time. We consider the five different peer groups (n) defined in Section 2.3 above: (1) firms in the same building; (2) firms in the two neighboring buildings, (3) firms in the two cross-street buildings; (4) person-connected peers; and (5) firms in the same ownership network. Finally, αit denotes firm-year fixed effects, μct denotes country-year fixed effects, and εict represents other sources of variation in importing. Our main hypothesis is that, due to knowledge spillovers, βn > 0 for the spatial and managerial networks. We also expect βn > 0 for the ownership network, but in that network the mechanism need not be a spillover: it is also possible that the common owner's knowledge causes firms in the network to import from the same country. Because they play an important role in identifying our key coefficients, it is useful to discuss the fixed effects in Eq. (1). The firm-year fixed effects αit control for any omitted variable driving import behavior which is specific to the given firm in the given year. This is a rich set of fixed effects, and the only reason it can be included is because the data have an additional panel dimension: multiple source countries. In particular, estimating Eq. (1) in the absence of data on source countries, or with a single source country, would not be feasible because the firm-year fixed effects would soak up all the variation in the dependent variable. In this sense the key βn coefficients are identified from source country variation. An implication is that standard firm controls, such as sales, employment, ownership status, or other balance sheet variables need not be included in the regression, since they are already picked up by the firm-year effects. In turn, the second set of fixed effects μct pick up country-year specific variation, for example business cycle fluctuations in a source country that might affect the supply of imports. Due to their presence, we do not need to include country-specific controls such as the exchange rate or GDP of the source country. Beyond import spillovers, slightly modified versions of Eq. (1) can also be used to estimate other kinds of spillovers. We will look at cross-activity spillovers where on the right-hand side of the equation we measure peer firms'country-specific experience in a different domain, such as exporting to or having a foreign owner from the country; and (in Appendix A.2) we will also use a variant to present evidence on export spillovers. Identification Since Eq. (1) is essentially a peer effects regression, the main threats to identification are those highlighted by Manski (1993): endogenous peer groups and correlated omitted variables.14 Endogenous peer groups might arise because of clustering or because of peer choice. An example in the spatial network is when firms from one industry, or “high-type” firms, tend to both co-locate and make similar import decisions, creating spurious correlation between Xnic, t−1 and εict. An example in the managerial network is when a firm hires a manager because of her or his import knowledge. And an example of correlated omitted variables is when particular physical locations are better for importing from a country c, perhaps because they are close to c. Our first research design addresses these concerns in three main ways. (1) Source-country variation. By using this variation we address the basic concern that importers tend to be connected to other importers. As discussed above, if we were to estimate Eq. (1) ignoring the source of imports, the firm-year fixed effects αit would soak up all the variation. The implication is that remaining threats to identification must be based on country variation: for example, if certain types of firms tend to import from certain countries and co-locate with each other. (2) Sample definition. We use comparable source countries; firms based in Budapest; and we omit ownership-based links from the spatial and managerial networks. Our sample choices mitigate several concerns. Because the source countries are similar, it is less likely that “high-type” firms import from one, while “low-type” firms import from another. Because all firms are in Budapest, omitted variables based on distance from a country are muted. And by removing ownership links we address the concern that correlated decisions may be driven by a common owner. In addition, by focusing on imports we limit the concern of endogenous manager choice as knowledge of importing seems less likely to be a driver of hires than for example knowledge of exporting would be. (3) Placebo spatial peers. Perhaps the most convincing component of our design is that by exploiting the fine spatial structure we can compare same-building and neighbor- building spillovers with a cross-street “placebo spillover”. As long as spillovers are more spatially concentrated than the omitted variables—an assumption consistent with the results of Arzaghi and Henderson (2008)—estimating higher β coefficients for the closer spatial peers is evidence for knowledge diffusion. For the above reasons we feel that the most plausible confounds are accounted for by our current research design. Still, a possible concern is that, because the design does not make explicit the source of variation in peer firms' experience, it may be subject to some remaining—highly spatially concentrated—omitted variable. In the next subsection we address this concern by combining the current design with plausibly exogenous variation in peer firms' experience due to firm moves. Although that approach requires weaker identification assumptions, it can only be used to estimate spillovers in spatial networks. We therefore begin the analysis with the current design to demonstrate that knowledge spillovers about imports are present quite broadly across different types of networks. Results Table 5 reports estimates of regression (1). In this and all subsequent tables reporting regression results, coefficients are measured in percentage points. To account for spatial correlation in the error term, in all specifications we cluster standard errors by building. Column 1 focuses on spatial spillovers. The estimated effect of having a same-building peer with country-specific import experience is a significant 0.22. Intuitively, having a peer with experience importing from a particular country, e.g., Slovakia, increases the probability that the firm starts to import from that country by 0.22 percentage points. For comparison, the baseline probability that a firm starts to import from a specific country is 0.19%; thus having a peer with experience importing from a country more than doubles the probability of entering that import market. Column 1 also reports that the estimated effect of having a peer with country-specific import experience in a neighboring building is a significant 0.04. This is a fifth as large as the same-building effect, and shows that while spillovers to neighboring buildings are also present, their intensity declines rapidly with distance. The cross-street spillover effect is an even smaller and insignificant 0.03. This result lends support to our identification strategy: if a spatially correlated omitted variable was driving our estimates, we would expect that variable to also affect firms in buildings across the street. Taken together, these estimates strongly support the presence of spatial spillovers in importing. Column 2 reports the analogous estimate for the person-connected networks. Having a firm official who had prior experience importing from a country increases the probability of importing by a significant 0.43, or almost half a percentage point. This estimate is twice as large as the same-building spillover effect. The larger magnitude seems intuitive: same-building diffusion is likely to be more limited because interactions between members of different firms are probably less common and less intense. In contrast, for person-connected spillovers, interactions are almost guaranteed since the manager now works for the firm. Column 3 shows the analogous estimates in the ownership-connected network. Here we estimate an even larger coefficient of 0.53. Importantly, this coefficient cannot be interpreted as a knowledge spillover because it is likely partly driven by a common owner making sequential import decisions for her or his firms. Indeed, the reason we include this specification is to show that controlling for the common ownership channel—which we do by excluding ownership-connected firms from the other networks—is important to convincingly document knowledge spillovers in spatial and managerial networks. Column 4 shows that combining all three types of networks in the same specification leaves the estimates essentially unchanged, indicating that the different networks represent genuinely different spillovers. We conclude that there are significant import spillovers in both spatial and managerial networks. In Appendix A.1 we show that these results are also robust to a range of specification changes including various subsamples (Table A1, A3), additional controls for the firms' or its peers'country-specific experiences (Table A1) and different measures of connections (Table A2). Research design 2: peer moves ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In our second research design we exploit a specific, plausibly exogenous source of variation in peer knowledge, which is created by firm moves. Focusing on the same-building spillover, we explore the effect of a peer with particular import experience moving into the building on a firm's subsequent import decision. This design has power because moves are quite frequent, with >25% of the firms in our main sample moving at least once.15 As it is unlikely that the mover would internalize the effect of its import experience on other firms in the building when it chooses its location, we can plausibly assume that country-specific experience brought by the mover is an exogenous shock for the local firms. Similarly, although the owner of the building might want to attract good firms, it is less plausible that she would want to attract firms with specific import experience. We estimate the impact of moves using an event study, in which the event is when a firm moves from another address into a building. The sample consists of (i, c, t), that is, (firm, source country, year) observations where firm i is located in a building in some year t which is subsequent to some other firm j moving in the same building. The event is the earliest date at which another firm moves into the building of i. To limit the confounding effects of preexisting neighbors, we restrict the sample to observations in which no incumbent firm in the building had import experience with the country c prior to the event. We do not require that the mover firm j has import experience with the country c. Buildings with movers having no import experience serve as controls. Here Yict is an indicator for firm i having imported from country c in some year up to and including t. Ditτ is an event-year indicator which equals one if the mover firm came to the building of i exactly τ years before t; and the τ = 5 category also includes those observations in which the move occurred >5 years ago. Xic is an indicator for the mover firm having had import experience with country c by the time of the move. As before, αit and μct denote firm-year and country-year fixed effects and εict denotes the error term. In this specification the coefficients γτ measure the baseline dynamics of importing from a country c following a move by any firm. The coefficients of interest are the βτ which measure the additional gains in importing when the mover had prior experience with country c. Because of the firm-year fixed effects, similarly to the previous research design this regression is also identified from variation across source countries. Because Yict indicates if the firm has ever imported from c by t, and because the sample definition ensures that the i has not imported from c before the mover's arrival, Yict effectively measures if a firm with no prior import experience starts to import from c in the period between the arrival of the mover and t. Thus βτ captures the probability that the firm learns how to import by year τ, even if that firm does not import in every subsequent year. Fig. 1 presents visual evidence from the event study by plotting the estimates of βτ together with their 95% confidence intervals. Panel (a) shows the results from the specification without fixed effects, while Panel (b) from one that includes the full set of fixed effects. Although the point estimates in the second specification are somewhat lower and the standard errors wider because of the large increase in the number of controls, both specifications show the same basic pattern: a gradual and eventually significant increase in the probability of importing from a country subsequent to a new neighbor with country-specific import experience moving in. The fact that the increase is gradual is consistent with the idea of knowledge diffusion. In the more conservative fixed effect specification of Panel (b), the insignificant first-year effect of 0.12 percentage points increases to a significant 0.78 percentage points after four years.16 These estimates have the same order of magnitude as the estimated same-building effect of 0.22 percentage points in research design 1, but highlight the importance of explicitly considering the dynamic response to moves. The pattern revealed here serves as one motivation for examining the dynamic response of further import entries to a new entry in the counterfactual analysis of Section 5 below, where we will also be able to compare explicitly the dynamics implied by research designs 1 and 2. In summary, our research designs 1 and 2, exploiting different sources of variation, different networks, as well as a placebo design, consistently yield evidence in support of the presence and economic relevance of the knowledge diffusion hypothesis. We conclude that knowledge spillovers in spatial and managerial networks play an important role in shaping firms' import decisions. Benchmarking magnitudes ~~~~~~~~~~~~~~~~~~~~~~~ To get a better sense of the magnitude of the spillover effect here we compare it to three sets of benchmarks. As our first benchmark we use export spillovers, the existence of which was documented by Mion and Opromolla (2014), Fernandes and Tang (2014) and Kamal and Sundaram (2016) among others. To make this comparison meaningful, we use the same data and empirical approach for both types of spillovers: we employ our identification strategy 1 to also estimate export spillovers in Hungary. Table A6 in Appendix A.2 presents the results. Both the patterns and the magnitudes are similar to our import spillover results. For example, in the full model including other type of experience as well the same- building effect is 0.16 percentage points, the neighbor-building effect is 0.04 percentage points and the managerial peer effect is 0.37 percentage points. Relative to the baseline hazard of starting to export, 0.21, these estimates correspond to an increase in export probability of 76%, 19% and 176%, while the analogous numbers for the increase in import probability relative to its baseline of 0.19 are 116%, 21% and 216%. Export spillovers, like import spillovers, are also highly concentrated in space. We conclude that diffusion of knowledge about importing is about as strong as diffusion of knowledge about exporting. As a second benchmark we ask what increase in firm productivity would predict the increase in the probability of importing created by knowledge spillovers. In our sample the probability of starting to import from a country is 0.19% for not-yet-importer firms in the lowest productivity quartile,17 0.28% in the second quartile, 0.47% in the third quartile and 0.58% in the highest quartile. Consequently, the estimated same-building import spillover effect of 0.21 percentage points is comparable to the predicted increase in the probability of starting to import as a firm moves from the second to the third productivity quartile. This result further confirms the economic significance of the estimated import spillover effect. In our third benchmark we look not at the strength of the spillover but at its speed of decay in space. In particular, we infer a parameter of spatial decay that can be explicitly compared to similar decay parameters in the literature. Our approach is to convert the same-building and neighbor-building estimates of research design 1 to a distance-based metric. We work with the decay function βij = k ⋅ e−δ⋅distij, where βij is the estimated spillover from firm j to i, distij is the spatial distance between the two firms, and k and δ are parameters. In the 65% of the sample which we were able to geocode, we find that the average distance between two neighboring buildings is 28.1 m. Assuming that distance is zero if two firms are in the same building, calibrating δ and k to the specification of column (4) in Table 5, we obtain δ = 0.0579/m. This implies that spillovers decline by 5.6% every meter.18 This value is somewhat higher than other estimates of within-city spatial decay. Indeed, the estimates of Arzaghi and Henderson (2008) on networking benefits among advertising agencies in Manhattan imply a decay of 0.3% per meter; those by Rossi-Hansberg et al. (2010) on housing externalities in Richmond imply a decay of 0.2% per meter; and those by Ahlfeldt et al. (2015) on production and residential externalities in Berlin imply decays of 0.4% respectively 1% per meter.19 The main common feature of these results and ours is that they all represent fairly strong decay: knowledge spillovers appear to be highly spatially concentrated. And the fact that our estimate is the highest suggests that in our context building boundaries are important barriers to diffusion. Our decay parameter estimate may be useful for calibrating urban economics models that feature knowledge diffusion of business practices such as importing.","In this section we investigate the heterogeneity of import spillovers by firm and peer characteristics. We focus on same-building spillovers because these were the strongest and most cleanly identified. We first explore heterogeneous effects separately by firm and peer characteristics, and then investigate how the interaction between these characteristics influences the strength of diffusion. This analysis yields lessons about the mechanism of spillovers, highlighting the potential benefits of clusters and targeted policies, which we then quantitatively evaluate in the counterfactual analysis of Section 5.20 Firm heterogeneity Here h indexes firm categories by a characteristic, such as productivity quartiles; and Iith is an indicator which equals one if firm i in period t is in the particular category h, such as the highest productivity quartile. The variable Xsb is an indicator for peers' import experience in the same building. Accordingly, the coefficients βh measure the effect of experienced same-building peers for firms in category h. For completeness, the controls include the analogous interactions of the category indicators with import experience in the four other networks (neighbor building, cross-street building, managerial and owner network).21 As usual, αit and μct denote firm-year and country-year fixed effects and εict denotes the error term. Table 6 reports the results from estimating heterogeneous effects by firm size, productivity and ownership. Column 1 focuses on size measured as employment, and categorizes firms into four groups. Group 1 includes those firms with at most 5 employees, group 2 those with 6–20 employees, group 3 those with 21–100 employees and group 4 includes firms with >100 employees.22 The coefficient of 0.07 percentage points shows significant spillover effects for the smallest firms in group 1. The subsequent coefficients imply that the spillover effects for larger firms are larger than those for firms in group 1, and are increasing in the firm's size category. t-tests show that the difference between the estimated coefficients of subsequent groups is significant at 5% in each case (denoted by # in the table). Larger firms are more likely to respond to import knowledge in their building. Column 2 reports heterogeneous effects by firm productivity quartile, defined using our TFP estimates introduced in Section 2. Here we find no spillovers for the least productive firms in group 1, but significant and increasingly strong spillovers in the higher productivity quartiles. The coefficients of subsequent groups are significantly different in two of the three cases. Finally, in column 3 we look at ownership: group 1 represents domestically-owned firms and firms without information on ownership, while group 2 represents foreign-owned firms. Spillovers are significant in both groups and significantly larger for foreign firms. Taken together, these results suggest that absorptive capacity (Lychagin, 2016), which is more likely to be present in larger, more productive, and foreign firms, is important for the adoption of import knowledge. Peer heterogeneity Here too we create categories for a characteristic, such as size, and Xsbic, t−1(h) is an indicator for having a same-building peer in category h which has import experience. Thus βh measures the effect of having an experienced peer in category h. Similar to Eq. (3) the controls include the analogous variables for the other networks. Table 7 reports the results. Column 1 shows spillovers by peer size, using the same cutoffs of 5, 20 and 100 employees already used above.23 Spillovers are significant even from peers in the smallest group. Although the differences are not significant at 5%, the point estimates show that spillovers are larger when peers are larger, except for peers in the highest quartile where the coefficient is imprecisely estimated. Column 2 shows the analogous specification using peers' productivity quartiles. Here too, spillovers are always positive, and point estimates are larger for higher productivity peers. The difference between the third and fourth quartile is significant. Finally, column 3 shows significant spillovers from domestic peers (group 1) and significantly larger spillovers from foreign peers (group 2). Although the coefficients in this table are slightly less precisely estimated, their general pattern strongly suggests that the import knowledge of larger, more productive and foreign firms—perhaps because they are more successful importers or more trusted peers—is more likely to diffuse. To further confirm this logic, in Table A7 of Appendix A.3 we show that spillovers are stronger from “more successful” importer peers, where import success is measured with the persistence of the peer's import experience. Number of peers We next explore whether having more peers with country-specific import experience increases the probability of importing. Simple models of diffusion would predict such an effect, as with more informed peers there are more opportunities for learning. We consider a specification in which the effect is linear and use the number of peers with country-specific experience as a right-hand side variable. Column 1 of Table 8 shows that increasing the number of experienced peers in the same building by one increases the average probability of import entry by 0.2 percentage points. Column 2 presents similar results from a more flexible specification in which we separately estimate the effect of having exactly k experienced peers in a specific peer group. These coefficients are comparable in magnitude to the 0.2 effect of the linear specification, and given the standard errors we cannot reject that in this range the number of experienced peers linearly increases the probability of importing. Taken together, the above results reveal plausible heterogeneity in knowledge spillovers: diffusion is stronger when firms are better, when peers are better, when the quality of knowledge is higher, and when there are more learning opportunities. Interaction between firm and peer characteristics ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We turn to explore how firm and peer characteristics interact in shaping diffusion. Interaction effects are potentially important because their presence indicates that sorting firms can further increase the adoption of good business practices. Productivity complementarities Table 9 shows the results from estimating this regression. Column 1 reports a specification in which high productivity is defined as the top quartile in the productivity distribution. The fact that the coefficients of the non-interacted indicators of high-productivity firm (βhl) and high-productivity peer (βlh) are positive and significant is familiar from the previous subsection. The key novelty in the specification is that the coefficient of the interaction between high- productivity firm and high-productivity peer is a significant 0.5 percentage points. In column 2 we change the definition of the indicator for high- productivity firm to be above the median of the productivity distribution. The patterns obtained here are similar, and in particular the coefficient of the interaction continues to be significant and positive. From these results we conclude that there are statistically and economically significant complementarities between firm and peer productivity for the adoption of good business practices. One implication of these results concerns the benefits of sorting. Because of positive complementarities, sorting firms by productivity can generate aggregate gains in the overall adoption of good business practices. This force is distinct from the basic idea that having more informed peers increases adoption: it suggests that even holding fixed the average number of informed peers—that is, the neighborhood structure—changing the pattern of sorting can further increase adoption. Table 10 shows the results, using the top productivity quartile for the definition of high-productivity firms. The positive and significant coefficients show that spillovers are positive for any firm type and peer type, so that having more knowledgeable peers increases the probability of importing. And the fact that βhh is much larger than the other coefficients shows the complementarity effect: diffusion is stronger when both the firm and the peers are more productive. Same-industry and same-product effects To further explore the nature of complementarities, we investigate whether spillovers are larger between same-industry firms, and within a given imported product category. For same-industry effects our strategy is to include separate indicators for experienced peers operating in the same 2-digit industry as the observed firm and operating in different industries. We do this for all networks, but only report the results here for the same-building network. Column 1 of Table 11 shows that same-building peers have a larger effect if they operate in the same industry as the firm. Relative to the significant different-industry spillover of 0.17 percentage points, the same-industry spillover is larger by 0.42 percentage points. Column 2 shows a similar pattern for the restricted sample of manufacturing firms, but perhaps due to the reduction in power the difference between the effect of the two peer types is not significant any more. The positive and significant cross-industry spillovers mitigate identification concerns related to clustering by industry. And the larger same-industry spillovers highlight the societal benefit of sorting firms based on industry for increasing the overall adoption rate of good business practices. Finally, to measure import diffusion within a product category, we modify our specification in two ways. First, we estimate separate regressions for each product category, using a sample of observations in which the firm has not yet imported the given product category from a specific country, and including as controls indicators for whether the firm has imported other product categories from that country before. Second, our right- hand side variables are indicators for “same-product importer peers”—that is, peers which have imported in the past the given product category from the specific country—and “different-product importer peers”—that is, peers which have only imported in the past different product categories from the specific country. The last four columns of Table 11 show the results for each of four product categories defined based on the BEC categories. The effect of different-product importer peers is significant in all four categories; and same-product spillovers are always higher, significantly so in three of the four cases. We conclude that spillovers are larger within a product category, which is intuitive if part of importing knowledge is product-specific and further strengthens the argument about sorting firms based on industry to maximize spillovers.25","In the presence of spillovers, policies that encourage firm trade can have additional indirect effects through a social multiplier (Glaeser et al., 2003). And when spillover effects are context-dependent, so is the size of the multiplier, opening the possibility that targeted trade policies generate larger social gains. In this section we use our estimates of the import spillover effect in a counterfactual analysis to explore how the size and composition of a firm's peer group shape the social multiplier. Our goal is to compute the model-implied effect on the number of importers of a non-importer firm's exogenously induced entry into importing. To do this we assume that import spillovers follow a simple diffusion model whose parameters are determined by our estimates. For simplicity, in the model we only allow import spillovers between peers in the same building. We assume that the probability that a non-importer gets “infected” is linear in the number of importing peers.26 We allow the diffusion probability to depend on both the sender and the receiver firm's productivity type, measured with an indicator which equals one if the firm is in the highest productivity quartile. We also allow firms to become importers independently of spillovers, with a baseline probability which is constant over time and across source countries, but can depend on the firm's productivity type. We assume that all spillover and baseline adoption realizations are independent from each other and over time. Given these assumptions, the model generates a Markov process, and we can track its dynamics, for each building, with four state variables: the number of high/low productivity importer/non-importer firms in the building.27 To parametrize this model we use specification (6) in Table 10, which estimates different spillover parameters by firm and peer productivity category, and also reports the change in spillovers by the number of experienced peers. We calculate baseline probabilities by firm productivity category using the subgroup of firms which have no experienced peers in the same building. Starting from an initial year s which we set to 2003, we then study dynamics in the diffusion model in each building over a five-year horizon. In doing this, we assume that firms do not move in or out of the building and do not enter or exit production. We also investigate the benchmark case of the model with no spillovers, in which the diffusion parameters are set to zero. The social multiplier ~~~~~~~~~~~~~~~~~~~~~ Here Mca(i), s+5 is the number of importers from country c on address a of firm i in year s + 5 and Tsc(i) refers to the “treatment status” of firm i in year s, taking the value 1 if this firm is induced to start importing from country c. The numerator shows the expected change in the number of importers after 5 years of firm i being treated. This term incorporates import spillovers. The denominator is the corresponding treatment effect in the benchmark model in which import spillovers are set at zero. Thus the multiplier measures how much larger is the treatment effect in the presence, relative to the absence, of import spillovers.28 Fig. 2 plots, in increasing order, the implied 5-year social multiplier for all non-importer firms that have non-importer peers in our data in s = 2003. The figure reveals substantial heterogeneity. Interestingly, for about half a percent of firms the multiplier is smaller than one: treating these firms results in a smaller number of total importers in the presence of spillovers than in the absence of spillovers. This is because spillovers have two effects: they increase the impact of treating firm i, but they also increase spillovers from other importers in the building. Because of this second force, the net effect of treating firm i can be reduced when spillovers are introduced, essentially because spillovers from peers of i crowd out spillovers from i. However this subtle crowding-out effect only overcomes the more intuitive positive effect for a small share of observations. The median multiplier in the figure is 1.03: inducing the median firm to import is 3% more effective in the presence than in the absence of import spillovers. The 90th percentile of the multiplier is 1.13. Thus inducing a firm to import which is located at this point of the multiplier distribution is 13% more effective once import spillovers are taken into account. While spillovers may not be very important for the typical firm, they seem quite important for a significant share of firms, suggesting that targeting policies to such firms can generate substantial benefits. Targeted trade policies ~~~~~~~~~~~~~~~~~~~~~~~ We next use our counterfactual to evaluate a hypothetical import-encouraging trade policy, which demonstrates how targeting can improve policy effectiveness. For policy evaluation the object of interest is not the multiplier, but rather the numerator of Eq. (7), which measures the five-year treatment effect of inducing firm i to import from country c. In our analysis we compare two policies: one in which we target firms for which this treatment effect is large, and another with no targeting. For simplicity we consider an import-encouragement treatment which is completely effective in teaching the firm how to import from a particular country. Thus we assume that treating a firm results in it starting to import from the country under consideration with certainty. Our targeted policy is to treat the 1000 firms for whom the estimated treatment effect is largest, while our non-targeted policy is to treat 1000 randomly chosen firms. To avoid complications arising from treating multiple firms in the same building, we restrict both policies to treat, for any given source country, at most one firm per building. And to induce some amount of diffusion we only treat firms which have not yet imported from the country and which have at least one other non-importer peer in the building. Evaluating the targeted policy is straightforward, as it requires computing the numerator of (7) for the selected firms. For the non-targeted policy the impact also depends on the specific set of firms treated. To measure its average effect, we draw the 1000 random firms 1000 times, compute the treatment effect for each draw, and average over draws. The differences between the impacts of the two policies are remarkable. The targeted policy yields after five years 285 additional importers for a total of 1285 importers. In contrast, the non- targeted policy yields, on average, 16 additional importers. In this example the targeted policy is 26% more effective than the non-targeted policy (1,285/1,016 − 1 = 0.26). Since the targeting is based entirely on observable firm characteristics such as the productivity of the treated firm and its peers in the building, in principle it can be implemented using public data. Overall, the result suggests that there can be large potential gains from targeting interventions to firms which are likely to be good seeds for diffusion. Internal consistency ~~~~~~~~~~~~~~~~~~~~ We now connect the simulation results of the diffusion model and the estimates of the mover design in Section 3.3. Both of these designs evaluate the dynamic impact of having an additional importer peer. Because they exploit different sources of variation and use a different combination of reduced-form and structural approaches, their comparison provides a useful test of internal consistency. As we have just seen, the counterfactual implies that turning 1000 random firms in different buildings with non-importer firms into importers would result in an expected 16 additional importers after 5 years. The point estimate of the mover design implies (Table A5) that five years after an importer moves into the building, the probability of an incumbent starting to import increases by 0.73 percentage points. Because the average number of incumbent firms in a building is 4.6, the latter estimate implies that turning 1000 firms in different “non-importer” buildings into importers would result in 0.0073 ⋅ (4.6 − 1) ⋅ 1000 = 26.28 new importers. This has the same order of magnitude as the counterfactual, and given our standard errors we cannot reject that the two are equal. We can also check intervening years. Table O9 in the Online Appendix reports the expected number of firms starting to import 1–4 years after the above treatment in both designs. Here too, the numbers have the same order of magnitude and given the confidence intervals we cannot reject that they are equal. These patterns are especially remarkable because the mover and the counterfactual design use somewhat different samples: for example, in the mover design the 5-year effect is identified from moves in the subperiod 1994–1998. We conclude that exploiting different designs and sources of variation lead to similar estimates of the dynamics of knowledge spillovers, providing internal consistency to our results.","In this paper we documented evidence for import spillovers. Exploiting source-country variation, precise spatial neighborhoods and plausibly exogenous firm moves in two complementary research designs, we obtained credible estimates of diffusion in spatial and managerial networks. We also documented that spillovers are stronger when firms or peers are better, and exhibit complementarities in firm and peer productivity. Taken together, these two results show that both high network density, and positive sorting in a given network, can increase diffusion. We then conducted a counterfactual analysis showing that due to the combination of these two forces the social multiplier of importing is heterogeneous, so that targeted import subsidy policies can have substantially larger effects. In combination, our results highlighted one concrete benefit of firm clusters: that of facilitating the diffusion of good business practices. More broadly, our analysis contributes to a growing literature highlighting the importance of business networks in shaping economic outcomes."],["Natural disasters impact economies not only through physical damages, but also by affecting survivors emotionally and psychologically. This can alter their economic behavior, in ways that remain poorly understood. We present a model of post-disaster savings that reveals two opposing tendencies: the need to self-insure through increased savings, and the drive to “enjoy life while it lasts” through increased spending. We use panel datasets from China's Sichuan province, and isolate psychological impacts by focusing on those who lived in quake areas but did not themselves suffer damages or injuries. Although they did not bear economic losses, they saved less, spent more on alcohol, and played majiang (a Chinese game)more often, suggesting that the “no tomorrow” tendency dominated over the precautionary tendency. The magnitude of the savings rate impact, a drop of 0.17 percentage points for each percent of distance closer to the epicenter, is economically significant, and persists in the medium term. --------------------------------------------------------------------------------","Economists have studied how households cope with negative shocks using mitigation strategies such as consumption smoothing (Townsend, 1994), income smoothing (Takasaki et al., 2010), or saving (Paxson, 1992). Udry (1995) showed that households increase their saving when they anticipate a weather shock, then tap into their savings after the shock strikes. These studies have in common that they describe behaviors that can largely fit a “precautionary savings” model: households build up buffers to protect themselves against potentially negative outcomes when they anticipate a loss. Such behavior may become more salient after experiencing a shock, such as an earthquake. However, exposure to a shock, in particular a deadly one, may also lead people to reflect on their own mortality. Realizing that life is fleeting leads people to discount future events, prompting them to save less. We call this the “live like there’s no tomorrow” or “carpe diem” model.1 For instance, after negative shocks, people are more likely to lower investment in education (Fortson, 2011) and increase unsafe sexual behavior (Oster, 2012). This is also sometimes referred to as a “nothing to lose” attitude (Harris et al., 2002; Hill et al., 1997). These two channels pull an agent’s savings behavior in opposite directions. The “precautionary savings” channel leads someone to spend less and save more, while the “no tomorrow” channel leads to higher spending and lower savings. Different people may be more affected by one channel or the other. Thereby the overall impact of a disaster on savings behavior will depend on which of these two effects dominates on average. The first contribution of this article is to provide a theoretical and analytical framework which explains why the empirical results in the literature may go in opposite directions. We show how the existence of two types of risk, the risk of losses and the risk of death, can produce such antagonistic responses. The second contribution of our work is empirical. We study how these phenomena were reflected in people’s economic behavior following the 2008 earthquake in Sichuan (China), benefiting from before–after panel data. To isolate the psychological impacts of the quake, in most of our analysis we exclude from the sample disaster victims who suffered damages or injuries. This is because their spending and savings are affected by the necessities of reconstruction, so by excluding them we can rule out the “consumption smoothing” impact channels. Instead, we are looking at households who did not directly suffer damages nor injuries, but did witness destruction around them, and thus would still be affected psychologically. We examine whether this group of households start saving more (the “precautionary savings” channel), or rather start saving less (the “live like there’s no tomorrow” or “carpe diem” channel). Our empirical analyses suggest the latter channel dominates. The data we have access to are uniquely appropriate for a study of disaster impacts for three reasons. First, earthquakes are among the least predictable and most sudden type of natural disaster: while some areas around the globe are known to be more earthquake-prone than others, the exact location of the epicenter within those zones is as good as random. This makes earthquakes uniquely suitable for use as a natural experiment. Second, the intensity of earthquake damage can be precisely measured. For most types of disasters, intensity can be hard to gauge, and is often approximated by the extent of damages. But using damages as an intensity variable poses problems of endogeneity, because areas more damaged are not necessarily those where nature struck hardest. The extent of damages may be more related to preparedness (an endogenous variable) than to the intensity of the natural shock (an exogenous variable). Those who tend to save less for the future may also be those who incur the most damages in a disaster, so the cause-consequence relationship between damages and savings behavior is blurred.2 Modern seismology provides measurements that can pinpoint the origin and magnitude of tremors with great precision. The distance to the epicenter is a fully exogenous proxy for intensity and is also remarkably easy to compute.3 Third, we have access to both pre-earthquake and post-earthquake data. Using three separate sources, we were able to compile a unique panel dataset spanning a vast area around the epicenter of the earthquake over three years: 2007 (pre-earthquake), 2009 (one year post-earthquake), and 2011 (three years post-earthquake) for our analysis. In addition to household savings, we also look at other behaviors: the time people spent playing majiang (a Chinese 4-player game sometimes called mah-jong), as well as their expenditures on alcohol, as supporting evidence. Our study contributes to the literature on natural disasters. Economic studies of natural disasters tend to focus either on physical damages, such as the costs of losses and reconstruction (Anderson, 1990; Mechler, 2004; Toya and Skidmore, 2007), or on macroeconomic consequences, such as the impacts on growth (Cavallo and Noy, 2009; Hallegatte and Przyluski, 2010; Noy, 2009). The psychological literature, on the other hand, places greater weight on emotional damages, such as the prevalence and persistence of depression or post-traumatic stress disorder (PTSD) (Madakasira and O’Brien, 1987; Yule et al., 2000). At the nexus of these literatures lies the fact that disasters influence economic systems through changes in the psychology and behavior of economic agents. Yet while this phenomenon is sometimes alluded to in both literatures, it is seldom explicitly treated. Our paper fills this lacuna. Our work also relates to the literature on the impacts of exposure to shocks on risk preferences. That literature tends to focus on risk aversion as elicited through experimental designs (choice lotteries etc.) rather than on economic variables (savings) as we do in this paper. However, our focus on the psychological nature of these impacts makes the connection relevant. As shown in Chuang and Schechter’s (2015) compelling review, the findings of this literature are mixed: while some find that agents became more risk-averse after a disaster (Cameron and Shah, 2015; Cassar et al., 2011; Chantarat et al., 2015), others find the opposite (Bchir and Willinger, 2013; Cameron and Shah, 2013; Eckel et al., 2009). Another comprehensive review suggests that measures of risk preferences are unstable over time (Schildberg-Hörisch, 2018). While our paper does not treat of risk aversion directly, the theoretical framework we develop provides a lens that accommodates such conflicting findings. One paper standing out in that literature is Hanaoka et al. (2018): it shows that males who experienced a higher intensity of the Great East Japan Earthquake exhibit decreased risk aversion, and also increased risky behaviors such as drinking and gambling. Their finding is consistent with a dominance of what we call the “no tomorrow” channel. However, they do not examine the impact on savings behavior, which is a focus of our paper and a key economic variable. The next section illustrates the “precautionary savings” and “no tomorrow” savings behaviors in a theoretical model, reviews the literature, and outlines our empirical strategy. Section 3 describes the background and data. Section 4 presents results as well as discussions on the robustness and the economic significance of those results. Conclusions follow. A simple model of disasters and savings behavior ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We develop an analytical model to illustrate the conflicting tendencies towards saving and dissaving that result from the psychological effects of an earthquake, emerging from the opposing “precautionary savings” and “no-tomorrow” tendencies. The model is rooted in the expected utility framework, in the spirit of Sandmo (1970) and several prominent publications that followed (Carroll, 1997; Deaton, 1991; Hubbard et al., 1994, 1995). Where δ is a discount factor, U a standard “well-behaved” utility function (increasing, twice differentiable, concave), and Ct is consumption at time t and depends on the probability of losses. The horizon is theoretically infinite, but in practice an agent will eventually die, after which their consumption is reduced to zero (the second term in the brackets). These inequalities show that the optimal amount of savings increases with the risk of suffering losses. This is the “precautionary savings” effect. On the other hand, the optimal amount of savings decreases with the risk of death, the “no tomorrow” effect. The two effects work in opposite directions. Fig. 1 provides graphical illustration of the “precautionary savings” and “no tomorrow” effects under specific values for ρL, and ρD.4 When both ρL and ρD are zero (case a: no earthquake) the agent simply splits total wealth in half, to consume equal amounts in the two periods. When ρL = 1 (case b: losses, but not death, are anticipated with certainty and equal to L), the agent saves an additional L/2 in the form of “precautionary savings”. When only ρD > 0 (case c: death is possible) the agent will consume more in period 0, up until the limit case ρD = 1 (death is certain), when the agent has no reason to save for the next period and consumes everything immediately (the “no tomorrow” effect). All intermediate cases with ρD > 0 and ρL > 0 will combine the two tendencies, and the balance between the two will determine whether the agent will save more, or less, than a no-risk equal split. We can use this conceptual framework to guide our analysis of the savings behavior of agents in Sichuan after the earthquake. Although the true values of ρL and ρD may not be affected by the earthquake, it is how agents perceive those probabilities that matters. That perception may change after experiencing an earthquake. Witnessing the quake adds information about the possibility of such an event (and the possibility of aftershocks), leading the agent to increase the perceived values of ρL and ρD.5 This would be true for all agents who did not know an earthquake could occur, or who underestimated how destructive or deadly an earthquake could be. Furthermore, even if agents had perfect information and knew the true probabilities of earthquakes prior to the event, behavioral economics and prospect theory suggest that agents act irrationally by placing excessive weight on more recent events, due to “salience” or “availability bias” (Tversky and Kahneman, 1973). People are more aware of disaster risk after a disaster, regardless of what the probability truly is (Asgary and Levy, 2009). Finally, ρL and ρD are not limited to the probability of earthquake damages/death, but all types of adverse events. Experiencing the earthquake leads to better information and greater salience of the unpredictable and fleeting nature of life in general, which also increases perceived ρL and ρD. Thus, only agents that are both perfectly rational and perfectly informed would not increase their perceived probabilities of ρL and ρD.6 Excluding economic impacts of the quake, the psychological impacts would lead an agent to react in one of three ways: If the agent is entirely rational and perfectly informed, they will not change their perceived probabilities ρL and ρD, and should not alter their savings rate. If the agent increases perceived probabilities and the effect of ρL dominates, they should engage in “precautionary savings” and increase their savings rate. If the agent increases their perceived probabilities and the effect of ρD dominates, they should adopt a “no tomorrow” attitude and decrease their savings rate. Which effect will dominate is influenced not only by the values ρL and ρD, but also by the size of the potential losses and other idiosyncrasies of the agent’s situation. The rest of the paper is dedicated to showing empirically that, in the case of our data, the “no tomorrow” effect was dominant. Empirical framework and identification strategy ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Three considerations are worth mentioning. First, we exclude the households which suffered direct damage in our main analysis. The destruction caused by the earthquake requires reconstruction expenditures that will force households to save less through sheer economic necessity – that is not the impact channel we are concerned about. The goal of this paper is to document the shifts in savings and other behaviors that are driven by psychological and emotional channels. We solve this problem by restricting our sample to those households whose homes did not suffer damages and who did not suffer injuries nor death. Such households do not incur reconstruction expenditures, but likely observed destruction around them, which could still trigger responses through the “precautionary savings” and “no tomorrow” impact channels, as they update the perceived probabilities of losses or death occurring. The same goes for health-related expenditures, funeral expenditures, etc. Another way to deal with this problem is to change the left-hand-side variable from the savings rate to something less affected by reconstruction costs. We estimate impacts on the frequency of playing majiang and on alcohol expenditures, which are less obviously related to the need to rebuild or reinforce a house. Secondly, our identification strategy helps isolate the psychological impacts, but it also comes with some drawbacks. One concern is the possibility that selecting the sample in that way introduces selection bias. We dedicate a part of the results section to ruling out these effects by running various tests and working with the full sample. Thirdly, in regard to our choice of distance-to-epicenter as a measure of earthquake intensity. At first glance, an obvious choice for an intensity variable would be the damage to someone’s house – but this raises the “Three Little Pigs” endogeneity issue discussed in the introduction. Destruction at the village level poses less of an endogeneity problem (though it does not fully solve it, because people with similar behaviors may tend to live in the same villages). An agent’s beliefs and attitudes are also influenced by the intensity of damage to the village in general (for instance, if a neighbor’s house was destroyed or the local school collapsed), so village-level damages are a better intensity measure for our purposes. We use village- level damages, deaths, and injuries as robustness checks and for certain specifications. Distance to epicenter, however, is arguably most likely to be exogenous, which is why we make it our preferred intensity measure.7 The Wenchuan earthquake ~~~~~~~~~~~~~~~~~~~~~~~ The province of Sichuan is located in southwestern China. At the point of contact between the Tibetan plateau and the eastern fertile basins, the region’s topography is highly uneven. Tectonic tensions between the Indian and Eurasian plates created the Longmenshan Fault, responsible for the seismic activity that resulted in the 2008 disaster. The 2008 earthquake took place on May 12, with its epicenter in the county of Wenchuan. The location of the epicenter is shown in Fig. 2. The original seism reached a magnitude of 8 on the Richter scale. It was followed by strong aftershocks spreading toward the northeast along the fault line, many of which were also of considerable magnitude (also on Fig. 2). The damage of the earthquake spread through the entire region and even affected Gansu, the province to the north. The Sichuan earthquake of 2008 was among the most destructive earthquakes in recorded history: it killed at least 69,000 people, and left between 4.8 and 11 million people homeless (Bulte et al., 2018). The extent of the damage captured national and worldwide attention. The collapse of school buildings and the number of child casualties exacerbated the emotional shock, fueled anger against officials involved in school construction, and triggered a corruption scandal. The severity of the disaster prompted the Chinese government to launch a swift reconstruction effort of unprecedented scale. Reconstruction was fully complete in less than four years (Yang, 2012). The full economic impacts of the earthquake, however, may have been longer lasting, including through the psychological channels we set out to document. Data sources and summary statistics ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Data sources. Because natural disasters are difficult to predict, it is rare for economists to have the pre-disaster data necessary for solid econometric estimations. Our work was made possible by constructing a unique dataset which merges three household surveys: the Sichuan component of the annual rural panel collected by the Chinese Ministry of Agriculture’s Research Center for the Rural Economy (RCRE), a supplementary survey administrated jointly by the RCRE and the International Center for Agricultural and Rural Development (ICARD), and the Sichuan Rural Household and Migration Survey (SRHMS) administrated jointly by the Shanghai University of Finance and Economics and ICARD. The design of the RCRE and SRHMS surveys has enough in common that merging them presents only minor compatibility concerns. The RCRE dataset is a national rural panel collected annually since 1984, and one of the most widely used data sources on rural China. It includes about 800 household observations in the Sichuan province, collected from 16 villages. While all of those villages experienced some level of earthquake (it was felt as far as Beijing), 5 of them were designated as directly affected by the earthquake according to government classification, and 1 as severely affected. This may limit explanatory power for earthquake-related questions. The SRHMS was started in 2007 and gathered information on about 800 households from 6 villages. The SRHMS focused specifically on Mianzhu County, which was later to become one of the most severely affected counties in the 2008 earthquake. The SRHMS was repeated after the earthquake, in 2009 and 2012. By merging 2007, 2009, and 2011 data from the RCRE, the supplementary RCRE data, and the SRHMS, we obtain a unique panel dataset that spans all of Sichuan, has a considerable number of observations in severely affected areas, and includes a pre- earthquake baseline.8 A full description of the datasets exists in Jin et al. (2013), and the location of villages is shown in Fig. 2. Our full sample contains 22 villages and 1306 households.9 The restricted sample of households whose homes were not damaged, suffered no physical injuries nor deaths as a result of the earthquake, is 589 for each year (1767 observations in total). We used Global Positioning System (GPS) coordinates to compute distances between villages and the epicenter of the earthquake (shown in Fig. 2). Coordinates for villages were obtained from the online mapping website map.baidu.com. All epicenter coordinates came from the United States Geological Survey (USGS).10 We also used mapping software to match each village with seismic intensity estimates publicly available from the USGS website.11 We culled consumer price index (CPI) values from the Sichuan Statistical Yearbook (SBSP, 2007, 2009, 2011) using the CPI in the nearest county where data are available. Explained variables. Household surveys seldom collect data on cash amounts saved by households, such that savings need to be computed as the difference between yearly income (Y) and yearly consumption expenditures (C). A common practice introduced by Deaton (1977) is to approximate the savings rate as r = ln(Y) – ln (C) = ln(Y/C), the natural log of the income-consumption ratio. This measure has the advantage of limiting the power of outliers (Chamon and Prasad, 2010; Deaton and Paxson, 1994, 2000; Wei and Zhang, 2011), and we use ln(Y/C) as our main measure of the savings rate.12 In the appendix we also show that results are similar when using s = (Y – C)/Y as the dependent variable, but this measure is less preferred due to outliers (see violin plot in Appendix Fig. A1). Other explained variables we use include expenditures on alcohol (expressed in Chinese renminbi, or RMB), and majiang frequency (expressed as the summed number of times per month any member of the household played majiang). All variables used in the analysis are presented in Table 1. The mean savings rate ln(Y/C) approximation in the sample is 0.62, the mean expenditure on alcohol is 63 RMB per month, and the mean frequency of majiang 3.1 times/month.13 For the dependent variables we also compare means in different years in Table 2: only alcohol expenditures dropped enough in 2009 to show significance in the t-tests. Earthquake variables. Our measure for earthquake intensity is the distance between a village and the epicenter of the seism. It was computed from GPS coordinates using geographic information systems (GIS) software in the World Geodetic System (WGS) 1984 Universal Transverse Mercator (UTM) Zone 48 N coordinate system. The average distance to epicenter was 236 km (sd = 95 km). We also use the village-level shares of deaths and injuries within the sample resulting from the earthquake, and the share of collapsed houses, as measures of how deadly/destructive the earthquake was locally. Control variables. The bottom of Table 1 provides statistics for the control variables used in the analysis, for the sample of non-affected households. All our specifications control for head gender, age, age squared, party membership, and education; household size, size squared, and landownership. In addition, the key observables we control for are household incomes, job loss in the household, consumer price index, amounts gifted, loaned, and received as government transfers. Appendix Table A1 shows how those variables vary by year. Incomes grew somewhat over the period, as did prices. There was an expected increase in government transfers received (significantly higher than the base year for both 2009 and 2011). Gifts to other households increased significantly, but only in 2011. The average age of the household head increased (also expected for a panel), and landholdings are slightly smaller in 2011 than in 2007. Most other differences in the table, such as the slight decrease in number of laborers per household, are of moderate magnitude and insignificant according to means-equality tests. Impacts on the saving rate ~~~~~~~~~~~~~~~~~~~~~~~~~~ The central result of our analysis is that households who were not directly affected by the earthquake reduced their savings rate, even after controlling for confounding factors. This result suggests that (1) the change in saving and spending behavior after the earthquake is at least in part due to psychological motivations; and (2) that the “no tomorrow” effect dominates the “precautionary savings” effect. We first illustrate this result graphically. We chart out the average change in the savings rate in the two years after the earthquake, by distance to epicenter, in Fig. 3. A negative value means the savings rate decreased in 2009 compared to 2007. The observations display a clear rising pattern, with a drop in savings rate in areas closer to the epicenter. The figure displays only “unaffected” households, whose houses were not destroyed, who reported no reinforcement expenditures, and whose members suffered neither injuries nor bodily harm. The figure shows that unaffected households closer to the epicenter tended to reduce their savings rate. Since unaffected households face no direct economic shocks to alter their savings behavior, the change in savings rate we observe closer to the epicenter may instead be driven by psychology (the “no tomorrow” effect). However, before we can claim to have estimated this psychological effect, we first need to rule out all other confounding factors and competing explanations in multivariate regression analysis. Table 3 shows the impacts of proximity to epicenter on the savings rate of households, using the framework specified in Eq. (10).14 Samples are restricted to unaffected households. All columns present the results of ordinary least squares (OLS) regressions with household fixed effects, thus differencing away any potential effect of time-invariant household characteristics. Each specification is run with the ln(Y/C) approximation of the savings rate as the explained variable.15 All specifications contain controls for household characteristics that may vary over time: household head’s gender, age, age squared, party membership, and education, as well as household size, size squared, landholdings – those coefficients are not presented in the table in the interest of space, and excluding them from the regression does not alter results.16 The explanatory variables of interest are the interaction terms between the distance to epicenter and the year dummies. They measure the effect of earthquake on change in savings rate relative to the base year, 2007. The intensity*2009 interaction term can be thought of as a short-term impact. The 2011 interaction term reveals medium-term impacts, at a time when much of the reconstruction effort was complete. Column (1) of Table 3 shows that the coefficient for the 2009 interaction terms is positive and significant (0.156). After the quake, households closer to the epicenter reduce their savings rate. How can we explain this earthquake-induced drop in savings rate for households who suffered no losses? We know that the significant coefficients in column (1) cannot be due to expenditures on reconstruction/repair of their home or other belongings, nor health or funeral expenditures. It cannot be due to the loss of children or dependents either (households who lost a dependent could anticipate a reduced need for savings), as households in the sample did not suffer human harm either. We also control for household size, so we can also rule out explanations based on dependents leaving the home. However, there are other alternative impact pathways that could explain this result. We identify four such plausible hypotheses, and attempt to rule them out in columns (2) through (5) in Table 3. The first alternative explanation is the “income hypothesis”. The literature generally agrees that there exists a relationship between incomes and savings rates, i.e. richer households tend to save more as a proportion of their income (Lawrance, 1991). Facing a negative income shock, households may choose to lower their savings rate in order to maintain consumption. Similarly, if the household lost certain sources of income or employment it could alter its savings rate. We test this “income hypothesis” in columns (2) (and the following columns) by including household incomes as a control variable, as well as a control for whether the household lost a worker since the last survey round. This hypothesis is borne out with a strongly significant coefficient (0.172) on the income variable: households whose incomes rose more also saved more as a proportion of that income. Further, once we account for incomes, the 2011 distance coefficient also becomes significant (0.227), suggesting that the impact of the earthquake persisted into the medium run. The R-squared measure also increases dramatically once we add incomes, confirming that this hypothesis holds explanatory power – but not enough to explain away the savings rate effect (Table A2). In addition, we show in appendices Fig. A2 and Table A4 that the savings rate impact we measure for unaffected households is driven by expenditures, not incomes. Fig. A2 shows graphically that distance to epicenter appears unrelated to the pre-post difference in household income. Expenditures, on the contrary, rose more in households closer to the epicenter, as demonstrated by a downward sloping trendline with highly significant coefficients. Table A4 confirms these results in a regression framework, and shows they are robust to all control inclusions. This further suggests that the savings rate effect we document is driven by the expenditure side, rather than by any income effects. The next possible explanatory factor is price inflation, which reduces incomes in real terms (or, symmetrically, increases the cost of consumption). Under this “inflation hypothesis,” households closer to the epicenter would be saving relatively less because their purchasing power for tradeable goods decreased. We test this by adding a measure of local consumer price index as a control variable, starting in column (3), but find it barely alters the other coefficients. Another possible explanation may be that unaffected households helped their affected friends or relations: the “altruism hypothesis”. Yet another possibility is that households in more affected areas received a greater amount of government assistance, even if their homes were not destroyed (some programs are location based). Households may tend to treat transfer income differently from earned income (the “nonfungibility hypothesis”), which could explain their lower savings rate. We try to rule out these two options by including measures of gifts and loans extended by households in column (4), and government transfers they receive in column (5). None of this influences results. Finally, the drop in saving rates could be due to a change in psychology, the “live like there is no tomorrow” or “carpe diem” hypothesis. Unfortunately, we do not have access to any variable that would allow us to directly measure this. However, since our results suggest we can rule out the income, inflation, altruism, and nonfungibility hypotheses, this would suggest that the “no tomorrow” hypothesis might be at play. It seems to affect the savings rate significantly, and to persists in the medium term. An additional concern regarding our results may arise from the fact that the number of villages in our data is small. All our regressions use cluster-robust standard errors, which is superior to uncorrected standard errors, but nevertheless imperfect. Specifically, it has been shown that one might over-reject null hypotheses when the number of clusters is small (Cameron et al., 2008, 2011). Therefore, we also performed wild bootstrapping tests on the two interaction terms coefficients as a post-estimation procedure (Cameron et al., 2008; Canay et al., 2018). The p-values derived from this procedure are included at the bottom of the table, and confirm our result: all of the 2009 and 2011 coefficients are strongly significant. The size of the impacts we measure is economically significant. They can be interpreted as follows: being one-percent closer to the epicenter is associated with a 0.17 percentage point decrease in the savings rate on average in 2009, and a 0.22 percentage point decrease in 2011. This translates to a distance-to-savings elasticity of about 0.3. In terms of the common earthquake intensity measure used in the media, the savings rate drops 20 percentage points for each degree on the Richter scale.17 Not only is this effect economically meaningful, it persists in the medium term, with a slight increase. Physical destruction versus human harm ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We further test the insights from our analytical model by focusing on the extent of death and destruction at the village level. If the model is correct, households in areas which experienced more death should exhibit more pronounced “no-tomorrow” behavior. Areas that experienced destruction, on the other hand, would likely exhibit more tendency for precautionary savings. Naturally, both effects will be at play in both types of villages, since death and destruction are correlated. In addition, with the existence of media, all households are likely aware of the extent of losses and casualties in the region. Nevertheless, we can expect households in villages where no deaths occurred to exhibit a muted “no tomorrow” response compared to those who witnessed deaths first-hand. We run OLS fixed-effects regressions on the savings rate measure using rates of death, injuries, and damages as measures of earthquake intensity in Table 4. All specifications include all the previously identified relevant controls. Column (1) shows that the share of deaths at the village level is significantly related to lower savings rates in 2011 (−0.177), though it is not significant in 2009. In columns (2) we see that the same is true of the share of injuries, though the magnitude of coefficients is much smaller (−0.03). In column (3), we see that a higher share of collapsed houses also leads to a lower savings rate, this time with significance both in 2009 and 2011, but again a smaller magnitude compared to the death rate. Finally, we re-ran the specification with collapsed houses, but limiting only to villages where no death occurred. This unfortunately leaves us with a rather small sample and we should not overinterpret the result, but nevertheless we find a positive and significant relationship in both 2009 and 2011. These findings bolster the hypothesis that witnessing death exacerbates the “no-tomorrow” channel, while witnessing damage favors “precautionary savings”. In the next section we move beyond the savings rate, and estimate impact on other outcome variables. We look at alcohol expenditures, and the frequency of playing majiang, both of which reflect “no tomorrow” attitudes. In addition, since these behaviors are not straightforwardly related to house rebuilding, the results also serve as a further robustness check supporting the notion that the change in behavior we observe is not driven by earthquake damages but rather by psychology. Impacts on alcohol expenditures ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Agents with a “living like there is no tomorrow” attitude can be more likely to engage in behaviors that bring about immediate reward but may cause delayed harm, because they expect they might die before those bad outcomes occur. One such behavior is the consumption of alcohol, which is known to be detrimental to health, but with symptoms that tend to develop over the long run.18 Under our “no tomorrow” hypothesis, we can expect agents closer to the epicenter to increase their alcohol intake relatively more. Table 5 presents the results of regressing the logged monthly alcohol expenditures on distance to epicenter (column (1)) using OLS and household fixed effects, with all the basic controls (age, education, etc.). We then include all the additional controls that help rule out the alternative explanations to the “no tomorrow” hypothesis in column (2). Column (1) shows negative and significant signs on the distance interaction terms for 2009 and 2011: households closer to the quake epicenter increased their alcohol expenditures more than those further away from it. Results persist in column (2), ruling out alternative explanations. Given the large number of nondrinkers in the sample, a Tobit specification left-censored at zero may be more appropriate to capture the alcohol expenditure decisions.19 We present those results in columns (3) and (4), mirroring the specifications of the first two columns. We find negative and significant coefficients in both years and both models. This result echoes that of Hanaoka et al. (2018), who found that Japanese men increased their drinking behavior after an earthquake. The OLS coefficients of interest, ranging from −19.95 to −25.61, can be interpreted as follows: being one percent closer to the epicenter increases the monthly expenditure on alcohol by about 0.2 RMB. This is about 0.3% of average expenditures, which is far from negligible. It is noteworthy that, similarly to the savings rate results, the effect on alcohol expenditures seems to intensify in the medium term (2011), with all 2011 coefficients being larger and more strongly significant than the 2009 coefficients. This may be due to the addictive nature of alcohol consumption. Addictions formed or reinforced in the aftermath of the earthquake may be harder to shed than other behaviors. Impacts on playing majiang ~~~~~~~~~~~~~~~~~~~~~~~~~~ The game of majiang (or mah-jong) is a popular pastime throughout Sichuan, as in much of eastern Asia. While the game often involves some financial stakes, those are usually small enough to be considered symbolic, such that playing majiang is primarily a form of entertainment rather than gambling. Either way, agents displaying a “no tomorrow” attitude would be expected to engage in it more often, either because they are more willing to gamble with their money, or simply because playing games with friends is a way to “enjoy life while it lasts”.20 The 2007 and 2011 surveys asked how frequently household members played majiang, in number of times per month. This part of the questionnaire was not administered in 2009, which limits us to estimating only the medium-term effects, three years after the earthquake. Table 6 presents the results of regressions of majiang frequency on the distance from epicenter. Majiang frequency is expressed as the average number of times someone in the household played majiang in a month. Because the frequency can be thought of as a count variable, so we also ran the regression as a Poisson model with fixed effects (columns 3 and 4). We can also think of it as a continuous variable truncated at zero, so the fifth and sixth columns use a Tobit model left-censored at zero with random effects. We obtain negative and significant coefficients across the board. Households further away from the epicenter increased majiang frequency less than those closer to it. This is robust to the inclusion of additional controls for income levels, prices, altruism and transfers, thus bolstering the “no tomorrow” hypothesis. Again, it is noteworthy that we find a strongly significant impact in 2011, suggesting that this behavior change existed in the medium term. Ruling out selection bias ~~~~~~~~~~~~~~~~~~~~~~~~~ Most results in this paper are based on the model in Eq. (10), which minimizes endogeneity issues by restricting the sample to unaffected households, ridding us of the effect of reconstruction expenditures. However, this model is not immune to criticism. The chief concern is that by restricting the sample to unaffected households, we may be introducing some selection bias. In this section we show that this is unlikely the case. The important question is whether the selection criterion (being an unaffected household) picks households differently across areas. In other words, are households selected into the sample more unusual when they are close to the epicenter than when they are far away? To test whether that is the case, we computed, for all variables used in the analysis, the z-score for each household within their village (z = (x − µ)/σ, where x is a household observation and µ and σ are the mean and standard deviation of that variable in a given year). We then regressed, for selected households only, each of these z-score variables on the distance from epicenter, the year of the survey, and the distance*year interaction terms. Table 7 shows results for the three explained variables we use in the paper and the four main control variables. Table A6 in the appendix shows results for all the other variables. None of the regressions show any significance on the distance*2007 coefficient (bolded). This suggests that, prior to the earthquake, the households that ended up selected into the sample were no more unusual close to the epicenter than those further away. In fact, there is also no significance for the distance*2009 coefficient, suggesting that even after the earthquake the close-by households were still not more unusual than those further away. The second way we can verify that selection bias does not drive our results is by running specifications using the full sample. Table A5 repeats the same specifications as Table 3, but does not use the selected sample. Using the full sample, we obtain nearly the same results as using the selected sample, suggesting selection is not the driver here. In Table 8, we use the specification in Eq. (11), controlling for all the previously discussed covariates, and also adding controls for the amount borrowed for house reconstruction, which could influence expenditure patterns. Unfortunately, for lack of data we are not able to control for whether the loan is still being repaid or not. However, Jin and Chen (2014) find that less than a third of rebuilding and reinforcement costs were financed with loans, minimizing the extent of this issue. Results in all three columns confirm almost all results presented above: both 2009 and 2011 coefficients are significant for the savings rate, alcohol expenditures rose significantly in 2009 (insignificantly in 2011), and majiang frequency also increased in 2011, consistently with our “no tomorrow” hypothesis. Ruling out price transmission ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Prices influence consumer behavior, which is why our regressions control for price levels by including the CPI as an explanatory variable. However, we may still worry that CPI is too coarse a measure of local price fluctuations. Unfortunately, we do not have local panel data that would allow us to include specific price levels (say, of alcohol) in the regressions we present. What we can do is to test whether, in general, rural markets throughout Sichuan are well integrated. If they are, then we should be reassured that geographic differences in price fluctuations are likely not what explains the changes in consumer behavior that we document.","Even when their home was not affected by the tremors, households who live closer to the 2008 earthquake reduced their savings rate, increased their purchases of alcohol, and played majiang more often after the disaster. Our estimates are significant in 2009, and most persist into 2011. Results are not explained by incomes, prices, altruism, or non- fungibility of income sources, suggesting that these changes may be reflecting the psychological impacts of the earthquake. The results are broadly consistent with a model under which agents lean towards a “no tomorrow” attitude, rather than towards “precautionary savings.” Those who lived in villages with higher death rates also reduced their savings, but the same is not true for rates of physical damage, further supporting the notion that the “no tomorrow” channel may be at play. We can only conjecture what the precise mental and emotional pathways might be that underpin our results. Some agents might be making the rational calculation that they may not live long enough to make use of their savings. Others might prefer not to accumulate assets that can be destroyed unexpectedly. Agents may be focusing on present enjoyment, or, conversely, be suffering from depression. A number of villagers may be affected by PTSD, which has also been associated with increased spending.24 The exact mental pathways that led to our results may differ between agents in the sample, and further generalization requires more research. While we cannot verify which mechanism is most prevalent in our data, all are broadly consistent with the “no tomorrow” model. The impacts we document are economically significant: the savings rate drops by about 0.17 percentage points with each percent closer to the epicenter, or about 20 percentage points for each degree on the Richter scale of earthquake magnitude. A population that starts saving less, consuming more, taking on risks, and generally tending toward “living like there is no tomorrow” is likely to substantially alter its economic path. In the medium run, we see that households closer to the epicenter save less, spend more on alcohol, and play more majiang even 3 years after the earthquake. Through these changes in attitudes, the consequences of natural disasters may continue shaping life and economic growth in affected areas long after reconstruction has been completed.25 Our analysis helps reconcile some mixed findings in the literature. Viewed as a tug-of-war between “precautionary savings” and “no-tomorrow” tendencies, it becomes clear why some studies find positive, others negative impacts of disasters on risk preferences and risky behavior. It is possible that the prevalence, or salience, of either destruction or death during a disaster can explain these conflicting findings, as suggested by our Table 4. The Sichuan earthquake cost more than 80,000 lives. The vivid TV images of deaths may have a much stronger psychological effect than less deadly disasters. The study of the differential effects of deadly and non-deadly disasters features highly on the future research agenda. Our work also brings to light another conundrum. Economists have documented higher saving rates in more disaster-prone areas (Skidmore, 2001). Yet we find persistently lower savings rates after the quake. This could reflect the fact that not all disasters produce the same rates of damages and human casualties. It could also mean that while the “no tomorrow” channel may start dominating, it may gradually recede in the long run, and let “precautionary savings” take over. This would be consistent with insights from behavioral economics: as time passes, the memory of the earthquake becomes less “salient”, which may alter the balance between the two channels. Assessing the longevity of these impacts thus provides a compelling avenue for future research. Some of our results seem to strengthen in the medium run. This can be highly relevant when viewed through the lens of addiction, habit formation or path dependency. Additional research is needed to fully understand the short-, medium-, and long-term impacts of living under the threat of disaster."],["Inequality in wealth among elderly households, and in particular the prevalence of very low wealth holdings, can be an important consideration in the design of social insurance programs. This paper examines the incidence and determinants of low levels of financial and total wealth using repeated cross-sections of the Health and Retirement Study (HRS) and a small longitudinal sample of HRS respondents observed both at age 65 and shortly before death. Most of those who report very low wealth holdings at the end of their life had little wealth at the traditional retirement age of 65. There is strong persistence over time in reports of very low wealth, and more generally relatively little evidence that wealth is drawn down in the first 15 years of retirement. The age-specific probability of reporting low wealth increases slowly after age 65. Low lifetime earnings are strongly predictive of low wealth at retirement and at the end of life. The post-retirement onset of a major medical condition, and, for married women, the loss of their spouse, are both associated with small increases in the probability of reporting very low wealth, but they account for a small fraction of low-wealth outcomes. Low levels of wealth accumulation before age 65, rather than gaps in the safety net after 65 or rapid spend-down of accumulated assets, appear to be the primary determinant of low levels of wealth just before death. --------------------------------------------------------------------------------","Alvaredo et al. (2016) review the primary sources of information on wealth holdings for all the but very richest households. These are administrative (tax) data on estates at death, which can be used to estimate the wealth of the living by applying (the inverse of) mortality multipliers differentiated by age, sex and wealth class; administrative data such as tax data on investment income, which can be “grossed up” to estimate the associated wealth distribution; and household surveys, like the HRS. Tax evasion and avoidance can make the first two sources problematic, while low response rates and under- reporting of wealth at the top of the distribution can make surveys unrepresentative. The HRS response rate, between 81 and 91%, is unusually high for a household survey. As with most large cross-section surveys, the assets of the very wealthy tend to be underreported.3 This is not a major concern for the analysis of low wealth holdings among the poorest elderly. The HRS data have many strengths but they also suffer from several limitations. First, the HRS samples each respondent at two-year intervals. With respect to end-of-life wealth measures, if a respondent dies just after completing an interview, the last recorded wealth value is a timely estimate of wealth in the last weeks of life. For those who die many months after their last survey, however, wealth balances “at the end of life” are measured with error. Because expenditures associated with declining health are often substantial in the last few months of life, the reported balances in the last interview before death are likely to over-estimate wealth at the time of death.4 Second, there are data outliers. Some may be accurate, but others may be the result of misreporting. To minimize their impact, we exclude records for 153 persons reporting more than $10,000,000 or less than −$1,000,000 of total wealth. We also focus much of our analysis on the probability that respondents report wealth below a threshold value. Measurement errors that do not move respondents across this threshold will not affect our findings. The HRS is a longitudinal survey that currently includes five cohorts defined by the year in which respondents are first surveyed. The original HRS cohort surveyed respondents between the ages of 51 and 61 in 1992 and the Asset and Health Dynamics of the Older Old (AHEAD) cohort surveyed respondents aged 70 and older in 1993. Subsequent cohorts include the War Babies (WB) cohort, first surveyed in 1998 when respondents were between the ages of 51 and 56, the Children of Depression (CODA) cohort first surveyed in 1998 when respondents were between the ages of 68 and 74, and the Early Baby Boomers (EBB) cohort that includes respondents aged 51 to 56 in 2004. All cohorts were surveyed every second year through 2012.5 Our primary sample includes HRS respondents from all cohorts who are known to have died and who were at least 65 years old in the last survey wave prior to their death. Of the 33,316 individuals who were alive in the HRS at some point between 1996 and 2012, 9215 died during this sample period. Of them, 7848 were age 65 or older at death. For some purposes, we also analyze a much smaller set of 1073 married respondents who were observed at age 65, the date we consider traditional retirement, and who also died during our 16-year sample period. We refer to this as our “longitudinal sample” because it allows us to track the full evolution of wealth and financial assets from age 65 to death. We define wealth as the sum of home equity, the net value of other real estate, business assets, and net financial assets. We convert asset balances to $2012 using the CPI-U, and measure them net of outstanding liabilities; both wealth and net financial assets can be negative. Financial assets include IRA and Keogh balances, as well as 401(k) and other defined contribution balances associated with the respondent’s current job.6 Balances in accounts sponsored by previous employers are not included.7 Our unit of observation is the individual, but for those who are married, we associate household wealth with each member of the couple. It can be difficult to assign ownership of assets, such as housing or jointly held financial assets, to specific household members. For some tabulations, we stratify results by the distribution of household lifetime earnings at age 65. The sub-sample used to produce these results includes the roughly two-thirds of HRS respondents who approved linking their survey responses to earnings and benefit histories from the Social Security Administration.8 The prevalence of low wealth in late life ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The upper panel of Fig. 1 shows the cumulative distribution of total wealth ($000 s) and of financial assets ($000 s) from the last survey wave prior to death for the 7848 HRS respondents over the age of 65 who died during our sample period. The figure shows the distribution for individuals, both married and single, with balances between -$50,000 and $1,000,000. We excluded 554 respondents with wealth values outside this range. Among those who died, the median wealth when last observed was $115,000 ($2012). More than three- quarters of respondents had financial assets less than that value. Four percent (including persons with balances less than −$50,000) had negative net worth, and 7% reported net worth of zero. About 8% of those who died were in households with negative financial assets, inclusive of 401(k) and IRA balances; 14% reported a zero balance. The median financial asset balance was $18,500. Most decedents were in households with relatively limited financial assets. The lower panel focuses on the 1073 married individuals who are in the longitudinal sample. The distribution is similar to that in the upper panel, but the respondents in the longitudinal sample have somewhat higher wealth at most quantile rankings. This is largely because the longitudinal sample is limited to married individuals, and married individuals on average have more wealth than elderly singles. Defining “low wealth” as wealth below $100,000 would include roughly half of the decedents in our sample. Somewhat arbitrarily, we use two definitions of “low wealth”: wealth less than $100,000 and less than $50,000. We also consider two measures of low financial assets: less than $50,000 and less than $25,000. In our full sample, 46.5 (33.8) percent of decedents fell below the $100,000 ($50,000) total wealth threshold. The low level of wealth for many households in their later years is not unique to the United States. Atkinson and Sutherland (1993) report that in the U.K., a substantial fraction of elderly households has little wealth and no private retirement support, and are therefore depends primarily on public support. The evolution of wealth at older ages ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The rate at which households draw down their wealth after retirement has long been an active research question. Particular attention for public policy purposes focuses on whether some individuals draw down their wealth rapidly after retirement, reaching old age with very limited wealth holdings. Fig. 2 plots the three-year moving average of the age- specific wealth of all HRS respondents (upper panel) and the married respondents who constitute the longitudinal sample. The figure shows average household wealth, in $2012, of all HRS individuals of a given age in any year of our data sample. The figures show the mean level of wealth, which is sensitive to the top wealth households, and the 25th, 50th and 75th quantiles. The slope of the profiles reflects both age and cohort effects, since those who older are, on average, from cohorts that were born earlier and had less lifetime income than their younger counterparts. The upper panel suggests gradual wealth decline, perhaps with slight acceleration in the decline as respondents age. Mean wealth declines by more than $100,000 between ages 65 and 75, but the rate of decline in wealth is much slower than the rate of decline in remaining life expectancy. The absolute rate of decline is greatest for the mean and 75th percentile wealth holdings. For example, the mean ratio (wealth at 80/wealth at 65) is 0.81. This value is 0.80 at the 75th percentile, 0.84 at the 50th, and 0.92 at the 25th percentile of the wealth distribution. Fig. 2 suggests relatively modest wealth decline at each point in the distribution that we consider. Second, there is a substantial group of individuals – at least those up to the 25th percentile – who have very low wealth holdings throughout the retirement period. The second panel in Fig. 2 shows wealth holdings for the individuals followed from age 65 to death. There is a notable difference between upper and lower panels: there is almost no downward slope for individuals in the lower panel. This is true for the mean and for all of the quantiles we consider. This raises the possibility that a mixture of cohort and age effects in the figure confounds estimating the age pattern of draw-down. One factor that may confound the wealth-age profiles is housing equity. Previous studies, such as Venti and Wise (2004), have found that older households are slow to move from the homes that they have lived in for many years, and that they are also reluctant to tap the equity in their homes to support other consumption. It is therefore possible that the age-financial assets profile might differ from the age-wealth profile, particularly given the importance of housing equity in the portfolios of older households. Fig. 3 addresses this issue. The upper panel shows the average financial assets of HRS respondents of each age, and the lower panel presents comparable information for the “longitudinal” sample of married individuals who die during the sample period. The results are very similar to those in Fig. 2: for the full HRS sample, there is a clear negative slope to the age-financial assets profile for the median and the 75th percentile financial asset value. The lower quantiles are much more stable across ages. For the longitudinal sample, there is very little change between age 65 and age 79 in the level of financial assets. These results also suggest that the negatively-sloped age-wealth profile from the simple tabulations may be spurious. Fig. 4 shows the three-year moving average of the estimated age coefficients (γi) from Eq. (1) for specifications in which total wealth and financial assets are the dependent variable. The age-wealth profile for total wealth shows some decline with age, but the decline is much more gradual than for the full sample age-wealth profile in Fig. 2. The decline in average wealth between ages 65 and 88 in Fig. 4 is about one quarter the decline in Fig. 2. For financial assets, the wealth decline in Fig. 4 is much less pronounced than that in Fig. 3. This suggests that part of the declining age-wealth profile in simple HRS tabulations is due to time effects that confound the age profile. This factor appears to dominate the countervailing effect of dynamic sample selection: individuals in higher socio-economic strata live longer on average, so the survivors at older ages is disproportionately drawn from higher wealth individuals. This would lead to a positive bias between age and the measured wealth effects, since the lower wealth members of a given birth cohort would be likely to die at younger ages, imparting a positive bias to the slope of the age-wealth profile. Our findings provide new evidence on the decades-long debate about the extent to which the accumulation and draw-down pattern suggested by simple life cycle models can characterize observed age-wealth profiles. The finding of relatively slow drawdown of total wealth – from an average of about $550,000 at age 65 to less than $500,000 by age 88 – is consistent with models in which retirees are husbanding their resources for potential late-life expenses than with models in which they are drawing down assets as soon as they reach retirement age. It supports a number of earlier studies, including Blundell et al. (2016), DeNardi et al. (2016), and Love et al. (2009), that find relatively slow draw-down of wealth after retirement. These results underscore the importance of including bequest motives, precautionary saving motivated by stochastic late-life expenditure needs, or other factors that can rationalize the slow draw-down of wealth in models that explain post-retirement wealth dynamics. Determinants of wealth at retirement ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The relatively modest age-related changes in wealth that the foregoing figures exhibit suggest that the distribution of wealth near the end of life may be largely determined by wealth at age 65. We now turn to its determinants. Table 1a reports the fraction of individuals in households with total wealth below either $25,000 or $100,000 at age 65, and it summarizes how the probability of falling in these low wealth categories varies with lifetime earnings and education. Among married 65-year olds, 9.3 (22.4) percent have household wealth less than $25,000 ($100,000). The fractions of single persons below each threshold are greater, 39% and 57.5% respectively. The wealth a household has accumulated by age 65 depends on its lifetime labor income as well as decisions about how much to save and how that saving was invested, which affects the rate of return it has earned. Table 1a stratifies the age-65 population in three ways: by education, by the presence or absence of a pre-retirement health condition, and by lifetime earnings quintile. Each of these factors may affect retirement wealth. Regardless of whether low wealth as defined using the $25,000 or $100,000 threshold, the variation across range across levels of education is similar to the range across lifetime earnings quintiles. For example, using the $25,000 threshold for married persons, about 23% of those in the lowest education group and 3% of those in the highest fall into this “low wealth” category. The pattern is similar for singles. Table 1b presents tabulations similar to those in Table 1a, but for financial assets. Single individuals have, on average, less wealth and fewer financial assets than those in married couples. Nearly 57% of single persons have household financial assets less than $10,000 at age 65, compared with 28.8% of individuals in married couples. Similarly, 68.2% of singles have less than $50,000 of financial assets, compared with 44.5% for married individuals. Table 1b shows that the percentage of persons with low financial assets varies dramatically by lifetime earnings quintile and by level of education. Using the $10,000 threshold, married persons in the lowest education group are 7.5 times more likely to have low financial assets than those in the highest education group. Married individuals in the lowest lifetime earnings quintile are 5.7 times more likely to have low financial assets than those in the highest quintile, and those in the lowest education and earnings quintiles are 16.1 times more likely to have low financial assets that those with high education and earnings. Education may have an indirect association with wealth at retirement through the effect of education on lifetime earnings. However, education also has an association with wealth that is independent of lifetime earnings. Within each earnings quintile, there are sharp differences in the probability of reporting low wealth. For married persons in the highest earnings quintile, for example, the probability of reporting less than $10,000 in financial assets is 32% for those with less than a high school degree, compared with 5% for those with at least a college degree. These differences could reflect a direct effect of education on saving rates and retirement preparation, such as an effect of education on financial literacy, or a spurious correlation between time preference rates and educational attainment that is manifest in different levels of retirement wealth. Tables 1a and 1b also explore the relationship between the pre-retirement onset of a major health condition and the likelihood of reaching retirement age with low wealth. For married persons, the probability of reporting low wealth is higher for those who experienced a major health condition. For example, using the $25,000 threshold (top panel of 1a), 7.6% of those who did not experience a major health condition had low wealth, compared with 12% of those who did. For singles, the percentage below the $25,000 threshold is 32.8% for those who did not experience a major health event and 48.3% for those who did. Not only is the probability of falling below this threshold higher for singles, the derivative effect of poor health for singles is larger than for married individuals. This casts doubt on the empirical significance of intra-household insurance mechanisms against chronic health shocks and other financial difficulties. Across lifetime earnings quintiles, the differences in the probability of reporting low wealth conditional on a health issue, even conditioning on lifetime earnings quartile, suggest that the effect of poor health on retirement wealth is not due only to its impact on earnings. Poor health can lead to reduced labor supply as well as higher levels of health-related spending. Dobkin et al. (2018), using several merged administrative data sets, find that hospitalizations among those under 65 are associated with reduced earnings and elevated debt levels. Our findings are supportive, and suggest that even after conditioning on lifetime income quantile, there are negative effects of a chronic health condition on wealth at retirement. This underscores the possibility of a non-earnings channel, such as out-of-pocket medical expenses that draw down wealth.","The slow rate of post-65 wealth decline suggested by the foregoing figures implies that wealth at 65 is likely to be a key determinant of wealth at the end of life. We now explore this relationship in more detail, first comparing wealth at 65 and at the end of life in repeated HRS cross sections, and then studying the subset of respondents who are observed both at age 65 and just before death.10 Repeated cross-section evidence ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ 3546 HRS respondents have linked earnings histories and are observed either at age 65 or 66. 2841 respondents have linked earnings histories and who die within our sample period. We can compare the wealth at retirement and wealth when last observed of these two groups – which have 1073 married respondents in common – to explore the differences between wealth at retirement and at death. Table 2 reports these comparisons. More HRS respondents have less than $100,000 in total wealth in the last survey before death than at age 65. For married individuals, the chance at age 65 of having wealth (financial assets) below $100,000 ($50,000) was 22.4 (44.5) percent. We reject the null hypothesis of an equal percentage of respondents below these thresholds at 65 and when last observed at the 95% confidence interval for the whole sample and for several subgroups stratified by lifetime earnings.11 The same pattern, but weaker statistical significance, is observed for singles. For some subgroups reported in Table 2, the probability of low wealth is higher at retirement than at death, but the null hypothesis of equality is only rejected in one of these cases. Most individuals who report low late-life wealth were also low lifetime earners. Nearly half – 47.8% – of the married individuals with wealth of less than $100,000 in the last survey before death were in the lowest quintile of lifetime earnings, as were 40.1% of those with less than $100,000 of total wealth at age 65. The pattern is similar for financial assets: 40.7% of those with less than $50,000 in financial assets at death were in the lowest lifetime earnings quintile; 31.5% of those with this level of financial assets at retirement were in the lowest earnings quintile. These data suggest that low lifetime earnings are a key predictor of low wealth at both retirement and end- of-life. Evidence from HRS cohorts, retirement through death ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The entries in Table 2 compare individuals at age 65 and at the end of life, but relatively few of the individuals in these two samples – 1073 married and 481 single respondents – are observed at both 65 and at the end of life.12 To provide longitudinal information on wealth trajectories, we now focus on this subsample of married individuals, and explore the relationship between wealth at 65 and at the end of life. We caution that this sample disproportionately includes individuals who died at young ages, because our sample spans only 16 years. No one in this longitudinal sample died at an age beyond 82. Our findings thus apply to the draw-down patterns of the “young elderly” but may not generalize to older groups. Fig. 5 graphs total wealth at age 65 and total wealth at death for the longitudinal sample; it shows only the 932 individuals with total wealth less than $1,000,000 at each date. The number of years that elapse between the two wealth measurements can be anything from two years for those who die in their mid-60s to nearly 16 years for those who die in 2012 and who were first surveyed in 1996. Those with flat trajectories of wealth in retirement will fall close to the 45-degree line. The figure shows substantial numbers of respondents with trajectories that place them above or below the 45-degree line. There percentage for whom the difference between the two wealth values exceeds $100,000 is relatively small, however: 30.8%. At low levels of wealth, most individuals are close to the 45-degree line, suggesting little change in wealth holdings between age 65 and the last HRS observation before death. Wealth declined for 54.9% of the longitudinal sample, was recorded as zero in both surveys for 1.0%, and increased for 44.1%. One shortcoming of the longitudinal sample is that the maximum number of years between age 65 and the last wave before death is about 14 years, so we are not observing wealth trajectories for those who die after age 79. It is likely that if we had longitudinal data for an even longer period, we would observe larger differences between wealth at age 65 and at death. Measurement error may explain some of the wealth movements that we observe. Wealth is self-reported and accurate reporting of a household's financial circumstances is a challenging task that may become more difficult as respondents age and cognitive skills decline. Researchers have long recognized the potential for measurement error to confound studies of household wealth, particularly when the analysis focuses on wealth differences. Hill (2006), Venti (2011), and Meijer et al. (2013) all report evidence of significant measurement error in the HRS, but none offer suggestions for resolving it nor a calibration that would enable us to assess the potential impact on our results. For example, Hill (2006) finds that call-back verification reduced the variance of wave-to-wave asset changes about 50%. Hurd et al. (2016) demonstrate that greater cross-wave validation effort reduces the variation across waves and improve the quality of data imputation. The HRS data file that we analyze incorporates the insights of this analysis. We focus our empirical strategy on measures of whether wealth falls below a threshold value in part to reduce the impact of measurement error. For respondents whose true wealth value is far from the threshold, even a large and transitory measurement error in one year will not bump the respondent across the low-wealth indicator. Tables 3 and 4, which present core transitions between wealth at age 65 and wealth when last observed before death. The top panel shows the percentage of married persons in various total wealth intervals at age 65 who were in various intervals when last observed. The off- diagonal entries reflect movements in wealth or financial assets. For example, 31.7% of the respondents who reported between $10,000 and $50,000 of total wealth at age 65 were in this same wealth category at death. There is greater persistence in the top and bottom categories: 73.6% of those who reported less than $10,000 in wealth at age 65 were in this wealth interval in the survey before death; 76.7% of those with more than $250,000 in assets at age 65 were similarly categorized when last surveyed. Some individuals make significant transitions in wealth holdings. 7.4% of those who reported more than $250,000 at age 65 had less than $100,000 when last surveyed. While we would like to explore what happened to the respondents in this group, the number of respondents with large declines is too small. The patterns are broadly similar for financial assets (lower panel), and if anything, low values are more persistent in that case. Table 4 reverses the conditioning: it stratifies married individuals in various total wealth intervals in the last year observed before death by their reported wealth at age 65. It shows column percentages rather than the row percentages shown in Table 3. For example, 48.2% of those last observed with less than $10,000 of total wealth also had less than $10,000 of total wealth at age 65. Persistence is particularly strong for persons dying with substantial wealth: 81.6% of those who have more than $250,000 when last observed also had more than $250,000 at age 65. At low wealth levels, of those with less than $10,000 of total wealth at death, 18.6% had more than $100,000 entering retirement, while over 50% of those with between $10,000 and $50,000 of wealth when last observed had wealth of more than $50,000 at age 65. This suggests some “downward mobility” between age 65 and death among those who had accumulated substantial assets at retirement. The bottom panel of Table 4 suggests similar patterns for financial asset mobility.13 Tables 3 and 4 suggest that about one third of those who are observed with low wealth at death entered this state between age 65 and death. While large movements in wealth or financial assets between age 65 and death are relatively uncommon for those who die before the age of 80, there are some households that draw down their assets and are poorer at death than at they were at retirement age. Most of those with low wealth at death, however, had low wealth at retirement.","We now examine two potential shocks – the onset of a health condition and death of a spouse - that could lead some individuals to fall into the low-wealth category. We in particular ask if these shocks could explain the wealth trajectory of the subset of HRS respondents who did not have low wealth at age 65, but did at the time of death. Adverse health events ~~~~~~~~~~~~~~~~~~~~~ The probability of ever having experienced a major health condition – cancer, heart disease, lung disease, or stroke – rises as individuals age. In our HRS sample, this probability is about 35% at age 65. It is nearly 65% by age 90. The probability of experiencing one of these conditions for the first time rises from about 5% per two-year interval (between HRS waves) at age 65 to between 6 and 7% at ages above 80. It is relatively stable for individuals of the ages we consider in our longitudinal sample. We explore the association between the onset of a major health condition and low wealth on the eve of death by comparing the wave-to-wave change in the probability of low wealth for those first reporting the onset of a major health condition to the wave-to-wave change for those who have never experienced any major health condition. The “never experienced” group is one of several comparison groups that could be used for this analysis; it is attractive because it does not raise confounding issues about the lagged effect of past major health conditions. There are 3832 respondent-years for our married person sample in which a new health condition is reported, and 17,063 respondent-years in which a survey participant who has never reported a major condition.14 Including a set of indicator variables for the level of educational attainment controls for the variation in the baseline risk of reporting low wealth that is education-related. Table 5 presents the results of estimating Eq. (2). The onset of a major medical condition is associated with an increase of about 0.8 percentage points in the chance that total wealth is below $25,000 for married individuals. The point estimates of the effects are larger but less precisely estimated for singles and we do not reject the null hypothesis of no effect. We also estimated effects over longer time periods – two and three waves – for both married and single respondents, and found that the standard errors rose and we could not reject the null hypothesis of no effect. Our analysis differs from the many previous studies that focus on the change in wealth at the time of a new health condition in that we focus on the likelihood of reporting low wealth, which as we explained at the outset, is particularly relevant for a number of policy issues. Death of a spouse ~~~~~~~~~~~~~~~~~ Fig. 6 reports the percent of married persons in the HRS over the age of 65 in each wave whose spouse died before the next wave. The horizontal axis is the age of the surviving spouse at the beginning of the two-year interval. Women tend to have older spouses and men tend to have younger spouses, so the age at which one partner becomes a widow/widower (on the horizontal axis) may be an imperfect indicator of the age of the partner at their death. The probability that a partner will die in a two-year interval increases from about 2% at age 65 to a little over 3% at age 70 and to almost 9% by age 80. Table 6 reports a specification similar to that in Eq. (2), but the central explanatory variable of interest is now the death of a spouse between wave t-1 and wave t. These estimates use a longer sample, but a similar approach, to the study by Sevak et al. (2003). The findings suggest that when a husband dies, the probability that his wife will report total wealth below $100,000 rises by about 1.6 percentage points, and the probability that she will report wealth below $25,000 rises by about 2 percentage points. For men, we do not reject the null hypothesis that the loss of a spouse has no effect. Some point estimates for financial assets, although none that are statistically significantly different from zero, suggest that loss of a spouse is associated with a rise in financial assets. Could there be any circumstances under which this might occur? Death-related payouts, such as life insurance benefits could lead to this outcome, especially because even a modest payout could move someone out of the low-wealth status. Financial assets could also rise if caring for a declining spouse leads to sale of a house and an associated set of balance sheet transfers. We explored the role of life insurance, and did not find any consistent evidence that the survivors of spouses with insurance were more likely to exit the low wealth state than the survivors of uninsured spouses. Even for the survivors of spouses without insurance, there were some reductions in the probability of reporting low wealth. We also explore the relationship between changes in wealth between age 65 and death, and new health conditions or the death of a spouse, in our longitudinal sample. 59.1% of those who experienced a major new health condition between age 65 and death reported a decline in wealth, compared with 52.7% of those who did not report but died within the sample. This difference is statistically significantly different from zero. For financial assets, the values are 59.2 and 54.2%, with a t-statistic of 1.58 for the difference. 63.2% of those who lost a spouse and also died within the sample reported a decline in wealth between 65 and death, compared with 54.3% of those who died but did not report losing a spouse (this difference is also statistically significantly different from zero). Reversing the conditioning, among those who experienced a decline in wealth (financial assets) between age 65 and death, 45.7% (45.4%) were diagnosed with a new health condition. This corresponds to 39.3% (40.5%) for those whose assets increased. Those whose wealth declined were also more likely (14.6 vs 10.6%) to have lost a spouse (t-statistic of 1.98). These findings do not control for potential differences between the various groups, but they suggest that there may be cumulative effects of both adverse health shocks and loss of a spouse.","Low lifetime wealth accumulation, which results in low wealth at retirement, is the most important factor contributing to low wealth in late life. Nearly two thirds of HRS respondents who were observed both at age 65, and about one year before their death, and who had net worth of less than $50,000 at death, also had similarly low wealth levels at age 65. Lifetime earnings and educational attainment are important determinants of wealth at age 65, and education is strongly correlated with wealth even after controlling for lifetime earnings. Just over 50% of high school graduates have low wealth, compared with only 6.4% of those with a college degree. Forty-five percent of married persons in the lowest quintile of the distribution of lifetime earnings have net worth of less than $100,000 at age 65, compared with only 7% of those in the highest quintile. These findings document an association but do not explain the mechanism linking education and wealth at retirement. One possibility is that education increases awareness of the need to save, or it makes individuals wiser investors who can earn a higher rate of return on their savings. Reverse causality is another possibility: those with more education may have wealthier parents and may have received larger bequests or other transfers. Still another possibility is that some third factor, perhaps time preferences, may have similar effects on financial and educational investments, inducing the strong positive correlation between education and wealth. Identifying a causal mechanism underlying the strong association merits further attention. In contrast to concerns that some households will draw down their retirement wealth at a rapid rate in the years following retirement, and exhaust their wealth before they die, we find relatively few households dropping to low wealth at death from modest wealth at retirement. In our longitudinal sample, 34.3% of married persons had wealth of less than $100,000 at age 65, compared with 40.1% just prior to their death. 67.1% had financial assets worth less than $50,000 at retirement, compared with 70.1% just prior to death. 55% reported lower wealth at death than at retirement. Both a decline in health, and the loss of a spouse, raise the likelihood of reporting low wealth, but the effects are modest. Onset of a major health condition is associated with an increase in the fraction of married persons with wealth below $100,000 from 21.3 to 23.8%. Death of a spouse is associated with a rise from 29.7 to 30.9%. The findings regarding health shocks are consistent with the view, described for example in Barcellos and Jacobson (2015) and Dobkin et al. (2018), that Medicare protects most of the over-65 population from substantial burdens associated with health shocks. This may be particularly true for those low in the wealth distribution. Poterba and Venti (2017) find that the negative wealth effects of three health-related shocks – hospitalization, admission to a nursing home, and use of a home health aide – are all much greater for those with net worth above $500,000 than for those with net worth below $100,000. This can reconcile the possibility that wealth declines, on average, in response to these shocks, while there is only a small effect on the probability of falling into low wealth. Households near the low wealth threshold may experience relatively less draw-down in wealth in response to these shocks than those higher up in the distribution. We have implicitly treated health shocks as creating mandatory expenditure needs that must be met, but there is another channel through which such shocks could influence the draw-down of wealth. Households experiencing such shocks might change their spending patterns, and hence their wealth trajectories. If adverse health events cause individuals to reduce their estimate of their longevity, they might respond by increasing their consumption spending and the rate at which they draw down wealth. Because the change in wealth reflects the return on wealth, plus other income, less consumption and health expenditures, an increase in consumption would appear as a decline in wealth for those experiencing adverse health shocks. Exploring competing mechanisms for the wealth health linkage is a topic for future work. Most of our results are based on a sample of individuals who died before age 80. It is possible, as Lee and Kim (2008) argue, that adverse health shocks at later ages are costlier than similar shocks at younger ages. More generally, wealth dynamics at older ages may differ from those at younger ages. As longer HRS longitudinal samples become available, it will be possible to extend our findings; they may change."],["We investigate the short-term effects of fiscal adjustment on economic activity in 20 OECD countries from 1970 to 2009. We compare two approaches: the traditional approach based on changes in cyclically adjusted primary balance (CAPB) and the narrative approach based on historical records. Proponents of the latter argue that it captures discretionary fiscal adjustment more accurately than the traditional approach. We propose a new definition of CAPB that takes account of fluctuations in asset prices and reflects idiosyncratic features of fiscal policy in individual countries. Using this new definition, we find that fiscal adjustments always have contractionary effects on economic activity in the short term; we find no evidence of expansionary (non-Keynesian) fiscal adjustments. Spending-based fiscal adjustments lead to smaller output losses than tax-based fiscal adjustment. These results are in line with the literature using the narrative approach, suggesting that the CAPB, when correctly specified, can be used as a measure of fiscal adjustments. --------------------------------------------------------------------------------","The recent global economic and financial crisis and the associated fiscal austerity in the Eurozone and elsewhere have resulted in a renewed interest in the relationship between fiscal reform and economic growth. Although there is widespread agreement that reducing public deficits and debt has important benefits in the long term, there is less of a consensus regarding the short-term effects of fiscal adjustment. In part, this is because of the oft-quoted examples of Denmark and Ireland, which experienced improved growth performance after periods of strict fiscal austerity in the 1980s.1 Their experience defies the conventional Keynesian theory which predicts negative short-run economic effects of restrictive fiscal policy. Subsequently, Giavazzi and Pagano (1990), Alesina and Perotti (1997), Alesina and Ardagna (1998, 2010) and others sought to find other examples of such expansionary fiscal adjustments and argued that fiscal adjustment can stimulate economic growth even in the short term, in a phenomenon referred to as ‘Non- Keynesian effects’ or ‘expansionary fiscal contraction’. Typically, the empirical studies seeking evidence of expansionary fiscal adjustments rely on observing changes in the cyclically adjusted primary balance (CAPB). The CAPB is an indicator that captures discretionary fiscal policy and other noncyclical factors by excluding the automatic effects of business cycle fluctuations (through transfers, taxes and interest payments) on the budget.2 However, Guajardo et al. (2014) criticize this approach on the grounds that using the CAPB can yield misidentified fiscal adjustment episodes. In particular, they argue that improvements in the CAPB due to stock-market booms and/or one-off fiscal revenues can be misclassified as fiscal reforms. They favor a narrative approach, based on careful study of historical documents to identify fiscal adjustments episodes. Applying this to a sample of OECD countries, they fail to identify any expansionary fiscal adjustments. In this paper, we build on and extend the work of Guajardo et al. (2014). We consider 20 OECD countries so that our scope is similar to that of their paper. In contrast to their approach, we use the CAPB instead of the narrative approach. However, we modify the CAPB to take account of the problems that Guajardo et al. (2014) point out. In particular, we construct the CAPB measure so that it reflects fluctuations in asset prices and takes into account the heterogeneity of fiscal policy across countries. Using this new measure of fiscal adjustment, we obtain results that are very similar to those that Guajardo et al. (2014) obtain with the narrative approach. With our modified CAPB measure, we find no evidence of Non-Keynesian effects. Nevertheless, our results confirm that spending-based fiscal adjustments have more beneficial macroeconomic effects than tax- based fiscal adjustments, which is in line with the previous theoretical and empirical evidence. This paper is organized as follows. Section 2 reviews the theoretical and empirical literature on the effects of fiscal adjustments. In section 3, we explain our new measure to identify fiscal adjustments and list the fiscal adjustment episodes that we identify. Section 4 outlines the empirical framework and presents the results. Section 5 examines the robustness of our results. Finally, section 6 concludes. Theory ~~~~~~ There is a general agreement that fiscal consolidation and the resulting reduced government debt contribute to long-run economic growth. However, there is disagreement about the short-term effect. Keynesian economics predicts that a cut in government spending or an increase in taxes reduces the aggregate demand and income directly, which leads to negative multiplier effects on the output in the short term. As a result, government debt to output ratio may not be reduced much or at all because both output and tax revenues fall due to the contractionary effects of the fiscal adjustment. In contrast, Neoclassical economics predicts that fiscal adjustments can stimulate the economy with an increase in private consumption and investment through several transmission mechanisms even in the short term. These mechanisms entail both demand and supply side effects. First, on the demand side, wealth effects or credibility effects are suggested to be at work. Blanchard (1990) proposes a model in which consumers react to two kinds of effects. One is the intertemporal tax redistribution effect by non-Ricardian agents in a Keynesian model where an increase in taxes decreases consumption. The other is due to distortionary effects of taxes, whereby an increase in taxes eliminates the need for a larger and more disruptive adjustment in the future. As a result, people can expect an increase in their permanent income due to the future reduction in the deadweight loss, and therefore consume more. Blanchard argues that if people exhibit little myopia and the fiscal adjustment is made from a high debt level, consumption can react positively. Bertola and Drazen (1993) present an optimizing model and demonstrate that if a change of fiscal policy induces sufficiently strong expectation of future policy change in the opposite direction, it can give rise to a nonlinear relationship between private consumption and government spending. If a cut in government spending induces expectation of significantly lower future taxes, it may induce an increase in the current private consumption. Similarly, Sutherland (1997) links current fiscal policy and future expected taxes. His model emphasizes the dynamics of government debt and considers consumers with finite horizons. At low levels of debt, fiscal policy has the usual Keynesian effects because people expect the debt stabilization program as something distant from their perspective. On the other hand, at high levels of debt, as a major fiscal consolidation is imminent, people react to government spending in a non-Keynesian way, expecting that they will have to pay more taxes shortly. Therefore, fiscal adjustment can give rise to positive wealth expectation effects when it occurs under high and rapidly growing debt-to-GDP ratio. Other mechanisms include credibility effects, whereby fiscal adjustment reduces the default and inflation risk via the decline in interest rates (Feldstein, 1982). When a high level of government debt affects the interest rate risk premium, a reliable fiscal adjustment can reduce the premium and, in turn, the reduction of interest rate raises permanent income. In addition, lower interest rate can also lead to the appreciation of financial assets which triggers higher consumption and investment. Expansionary fiscal adjustment may take place also on the supply side via the labor market and investment (Alesina et al., 2002, 2008). If fiscal adjustment takes the form of a cut in public spending, especially in the area of public employment, rather than an increase in taxes, it can lead to a reduction of overall wage pressure in the economy and stimulate private employment and investment. Empirical considerations ~~~~~~~~~~~~~~~~~~~~~~~~ There is a large empirical literature seeking to find evidence of expansionary fiscal adjustments (non-Keynesian effects) since Giavazzi and Pagano (1990) suggested, based on the examples of Denmark and Ireland in the 1980s, that large and decisive fiscal adjustment can increase private consumption. Fiscal adjustment is usually identified as an improvement of CAPB in excess of a chosen threshold over a given period. Two aspects of fiscal adjustments receive attention: the factors that ensure that fiscal adjustments are expansionary or successful, and the effects of fiscal adjustments on macroeconomic outcomes. The studies interested in the former classify the episodes according to the definition of expansionary or successful fiscal adjustment3 and then perform a descriptive analysis of the characteristics of fiscal components and other related macroeconomic variables such as GDP and interest rate before, during, and after the fiscal adjustment period (Alesina and Ardagna, 1998, 2010, 2012; Alesina and Perotti, 1995, 1997; Giudice et al., 2007; McDermott and Wescott, 1996). These studies tend to find that fiscal consolidations based on spending cuts rather than on tax increases are more likely to be both expansionary and successful. Other papers use binary dependent variable models such as logit or probit to analyze which factors determine the success of fiscal consolidations (Afonso et al., 2006; McDermott and Wescott, 1996) and their expansionary effects (Alesina and Ardagna, 1998; Giudice et al., 2007). McDermott and Wescott (1996) argue that the success in reducing the debt ratio depends on the size and composition of fiscal adjustment. They show that fiscal adjustments based on spending cuts are more likely to be successful than tax-based ones. Furthermore, the greater the magnitude of fiscal adjustment, the more likely it is to succeed. On the other hand, they show that fiscal adjustments are more likely to fail during a global recession. Afonso et al. (2006) assess fiscal consolidations in Central and Eastern European countries and suggest that spending- based adjustments were more successful. Giudice et al. (2007), finally, conclude that fiscal consolidation is more likely to promote economic growth during periods of below potential output and when the fiscal adjustment is based on spending cuts. The latter focus (on macroeconomic effects of fiscal adjustments) is less common. Using panel data of industrial and developing countries, Giavazzi et al. (2000) analyze the relationship between fiscal policy and national savings and conclude that their relationship can be nonlinear when the fiscal impulse is sufficiently large and persistent, similar to the previous studies on fiscal policy and private consumption (Giavazzi and Pagano, 1990, 1996). Ardagna (2004) studies the determinants and channels through which fiscal adjustment affect growth. She shows that whether a fiscal adjustment is expansionary depends largely on the composition of fiscal policy, and that spending cuts can lead to higher growth rates via the labor market rather than through agents' expectation. On the other hand, Burger and Zagler (2008) analyze the relation between U.S. growth and fiscal adjustments in the 1990s and argue that non-Keynesian effects prevail through an increase in consumption because of improved consumer confidence and an increase in investment via the labor market and financial market. Afonso (2010) assesses expansionary fiscal adjustment in European countries and finds that fiscal consolidations tend to have long- term expansionary effects, but no significant effects in the short-run. Thus, although there are some differences among these empirical studies in the factors affecting expansionary fiscal adjustment such as the size, composition and the initial conditions, overall, the empirical literature based on fiscal adjustment episodes identified by the changes in the CAPB supports the existence of non-Keynesian effects. Several papers, however, take issue with the results of the aforementioned CAPB-based studies. First, they argue that the results can be plagued by selection bias, measurement error, spurious correlation, or simultaneity issues when identifying fiscal consolidation episodes. Using the same panel data as Giavazzi et al. (2000), Kamps (2006) challenges their finding that non-Keynesian effects are a general and easily exploitable phenomenon by showing that the nonlinear effect disappears when cross-country heterogeneity is taken into account. Park and Song (2010) and Hernández de Cos and Moral-Benito (2011) raise the possibility of endogeneity of the fiscal consolidation decision and find that fiscal adjustments have negative effects on GDP when taking this into account. In line with these criticisms, IMF (2010), Guajardo et al. (2014) and Devries et al. (2011) suggest an alternative way of identifying fiscal consolidations instead of using the CAPB. Their approach is based on reviewing historical documents, similar to Romer and Romer (2010), to identify discretionary fiscal changes motivated by the desire to reduce the budget deficit. They then compare their episodes with those of Alesina and Ardagna (2010) and show that the episodes identified using historical data have contractionary effects on GDP, while the CAPB-based episodes are associated with a rise in GDP. Hence, using the CAPB is likely to be biased in favor of non-Keynesian effects. The narrative literature, furthermore, identifies a number of problems related to using the CAPB. First, using a statistical concept such as the CAPB can be marred by non-policy related developments such as booms or contractions in the stock market.4 Second, the CAPB-based method is likely to ignore the motivation behind fiscal changes. For example, the rise of CAPB can be aimed at restraining economic overheating, not at reducing the budget deficit.5 In addition, it can omit some episodes of fiscal adjustment followed by an adverse shock and discretionary fiscal stimulus.6 Third, the CAPB data cannot exclude some cases of offsetting positive changes in the CAPB caused by large one-off accounting operation in the previous year that are unrelated to fiscal adjustment measures, such as the capital transfers of Japan in 1998 and the Netherlands in 1995.7 Based on their new dataset, they conclude that fiscal adjustments have contractionary effects on economic activity, and argue that large spending-based fiscal consolidations cannot be expansionary. On the other hand, Alesina and Ardagna (2012) re-estimate the effect with new episodes identified based on the persistence of CAPB changes rather than on their size. They make a somewhat intermediate conclusion that spending-based adjustments cause smaller recessions than tax-based ones. They also argue that an expansionary fiscal adjustment is possible when it is combined with accommodating monetary policy. Last but not least, fiscal adjustments may affect the economic activity and vice versa. Countries are more likely to consider implementing fiscal adjustment to reduce the debt-to-GDP ratio when they are experiencing relatively favorable economic growth. Therefore, expansionary fiscal adjustments can be the result of self-selection and the decision to implement a fiscal adjustment is endogenous.8 Moreover, as many theoretical studies argue, if wealth effects and expectations are the main channels by which the fiscal adjustments affect economic activity, the episodes identified by the narrative approach based upon announced plans for deficit cuts can capture the fiscal adjustments and their effects better and more correctly than those identified by the CAPB based on actual fiscal outcomes. The main upshot of the preceding discussion is that while the narrative approach seems superior in identifying fiscal adjustment correctly,9 using the CAPB has the advantage in its simple and easy application. Therefore, if the construction of the CAPB could be improved to reflect the problems pointed out by the narrative approach, the CAPB could be a useful indicator of fiscal policy. That is what we set out to do in the remainder of the paper. Data ~~~~ We use an unbalanced panel of OECD countries covering the period from 1970 to 2009. All fiscal and macroeconomic data are obtained from the OECD Economic Outlook No.88.10 The sample includes 20 countries for which we have data for 20 years or more: Australia, Austria, Belgium, Canada, Denmark, Finland, France, Germany, Ireland, Italy, Japan, Korea, Netherlands, New Zealand, Norway, Portugal, Spain, Sweden, United Kingdom, and United States. To identify the instances of discretionary fiscal adjustment, we use cyclically adjusted primary fiscal variables. In particular, we use primary fiscal variables which exclude interest payments because the fluctuations in interest payments cannot be regarded as discretionary. To make the cyclical correction, we follow the method proposed by Blanchard (1993) and used also by Alesina and Perotti (1995, 1997) and Alesina and Ardagna (1998, 2010, 2012). It is simpler and more transparent than the more complicated official measures such as those of the OECD and IMF which utilize estimates of potential output and fiscal multipliers (Alesina and Ardagna, 1998, 2010). The basic principle of this method is that since the government spending depends negatively on GDP due to unemployment benefits and the revenues respond positively to GDP due to tax receipts, the changes in the cyclically adjusted fiscal variables can be calculated from the difference between the predicted current-year value (which would prevail if unemployment had not changed since the previous year) and the actual value in the previous year.11 The CAPB, however, can also be affected crucially by changes in asset prices. For example, a stock market boom can increase the cyclically-adjusted tax revenues because of capital gains taxes as well as increased tax receipts due to higher private consumption and investment. As identified by Guajardo et al. (2014), such measurement error, i.e., the correlation between the change in the CAPB and the error term (in the regression capturing the effect of fiscal policy on economic activity) is likely to be positive, and this could downplay the contractionary effects of fiscal consolidation in using the CAPB. Among other studies that highlight the importance of asset price changes to fiscal policy outcomes, Morris and Schuknecht (2007) find that stocks and real estate wealth, which account for a significant share of household and corporate wealth and whose value can change considerably over a relatively short period of time, are particularly relevant for tax receipts.12 Tagkalakis (2011a, 2011b) also finds that financial markets have a significant impact on the fiscal positions and suggests that higher asset prices improve fiscal balances and contribute to initiating a successful fiscal adjustment.13 Therefore, we use a share price index as an additional variable determining the CAPB.14 The impact on fiscal balance, especially tax revenues, can be different according to the types of asset price and tax systems (Morris and Schuknecht, 2007; Tagkalakis, 2011a). Therefore, when considering asset price variables as a business cycle factor, it would be ideal to include other types of asset prices such as equities and property prices. We use only the share price index due to data availability and its particular relevance to tax revenues. This can be deemed a limitation of our methodology, but we believe this index is representative of the other asset price movement.15 Guajardo et al. (2014) criticize the CAPB using the example of Ireland in 2009. In that instance, the CAPB to GDP ratio, used by Alesina and Ardagna (2010), fell because of the decline in tax receipts due to the sharp fall in stock and house prices. They argue that this shows the inaccuracy of the CAPB, as the Irish government was implementing austerity reform at the time. Indeed, our new measure that takes account of the fluctuations in asset price has the Irish CAPB improving by 1.3% of GDP in 2009. Definition of fiscal adjustment ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The literature identifies fiscal adjustment episodes as large and lasting changes in the CAPB. Table 1 summarizes the definitions used in the literature: the criteria of size and persistence are considerably different across the various studies, and even a little arbitrary. In addition, although different studies impose different thresholds, the same threshold is always applied to all countries. In other words, they do not allow for country-specific heterogeneity in discretionary fiscal shocks and the private sector responses to them. Since the expectations and confidence of the private sector are key factors for the transmission of fiscal shocks, past fiscal record should be considered. For example, for a country which has seldom experienced large changes in discretionary fiscal policy, a small fiscal adjustment can send a strong signal of the government's willingness to reduce the budget deficit. However, for a country that has had large fluctuations of fiscal policy in the past, a similarly sized fiscal adjustment can prove insufficient to elicit any response from the private sector. For example, while Burger and Zagler (2008) and Guajardo et al. (2014) identify several fiscal adjustment episodes in the U.S., Alesina and Ardagna (2010, 2012) identify none for that country. Therefore, one should consider the idiosyncrasy of fiscal policy in each country. For this reason, we base our definition of fiscal adjustment on the country-specific average (µi) and standard deviations (σi) of the changes in the CAPB. Our definition for identifying fiscal adjustment episodes has 4 criteria, incorporating size, persistence and country-specific heterogeneity, as follows: These criteria are chosen for the following reasons. First, as already explained, the different cut-off values reflect the heterogeneity of each country's fiscal policy, as embodied in the average (µi) and standard deviation (σi) of the changes in the CAPB: the standard deviation (σi) during 1970–2009 ranges from 3.73% of GDP in Norway to 0.88% in the U.S. Second, the 3rd criterion ensures that episodes when the CAPB improves little or deteriorates temporarily, but are offset in the following year, are also counted. Third, the 4th criterion excludes sharp increases in the CAPB due to one-off accounting operations such as one-time capital transfers. As in the other studies, there is also an element of arbitrariness in our definition, in particular in choosing the multiples (1, 1/3, 4/3, 2) of standard deviation. In the robustness section, we use alternative rules and thresholds in order to check whether the results are sensitive to these values. According to our definition, we identify 199 instances of fiscal adjustment in 20 OECD countries from 1970 to 2009. These consist of 66 distinct episodes, as summarized in Table 2 and Figs. 1 and 2.16 These episodes include only those that led to a sufficiently large improvement in the CAPB. This list includes several well- known episodes such as Denmark (84–86) and Ireland (82–84, 86–88). Importantly, the episodes that are used by Guajardo et al. (2014) to illustrate the discrepancies between the two approaches are identified correctly.17 Note that applying country-specific thresholds for identifying fiscal adjustments need not necessarily imply that the thresholds are generally lower. When we use Alesina and Ardagna's (2012) definition of fiscal adjustment but account for the stock market, this results in identifying fewer episodes for some countries but more episodes for others (results available upon request). As Fig. 1 shows, most episodes are of short duration. Of the 66 episodes, 11 last for one year, while 19 episodes take two or three consecutive years. The longest episode is 9 years: Japan from 1979 to 1987. Fig. 2 shows that the episodes of fiscal adjustment appear most frequently during the 1980s and 1990s. In particular, we can observe instances of concentrated fiscal adjustment of relatively short duration in the EU countries. This is likely to be related to the Maastricht treaty in 1992 which set criteria for euro area membership (Guichard et al., 2007).18 Determinants of fiscal adjustment ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To assess to what extent the incidence of fiscal adjustment depends on contemporaneous or past growth as well as other macroeconomic or fiscal factors, in this subsection, we analyze the determinants of the decision on undertaking fiscal adjustment. To this effect, we estimate a panel logit model where the dependent variable is equal to one during periods of fiscal adjustment and zero otherwise. Although the multi-year fiscal adjustments can be regarded as a single episode, the fiscal authority decides not only on initiating fiscal adjustment, but also on its continuation in the subsequent years. Therefore, the dependent variable takes value of one for each year during an episode of fiscal adjustment.19 Initial conditions such as the economic and policy environments can influence the decision on fiscal adjustment. Therefore, the explanatory variables are classified under three categories: macroeconomic, fiscal and political. The resulting marginal effects are reported in Table 3.20 The first column contains all variables, while the subsequent columns omit first the fiscal variables and then also the macroeconomic ones (except contemporaneous or lagged growth).21 The result of this exercise is that current growth is never significant unless there are no further fiscal or other macroeconomic variables left in the regression, in which case current growth becomes negative and significant (column 9). Lagged output gap, when included in the regressions, is significant and negatively affects the probability of fiscal adjustment. When lagged output gap is omitted, lagged output growth becomes significantly negative instead. As for the coefficient of current output growth being significant in column (9), it is not clear whether this indicates endogeneity or omitted-variable bias, given that it only appears when no other macroeconomic or fiscal variables are present in the regression. Note also that none of the political variables in column (9) are significant either, which suggests that the current output growth term in effect captures only the impact of fiscal and other macroeconomic variables. For the sake of comparison, we replicate also the analyses of Alesina and Ardagna (2010, 2012) and Guajardo et al. (2014) with our data: these results are not reported but are available upon request. When we do so, the current growth rate is significant and positive when using the fiscal adjustment episodes identified by Alesina and Ardagna (2010, 2012), and insignificant when using the fiscal adjustment episodes of Guajardo et al. (2014). In other words, fiscal adjustments identified with the standard CAPB-based method are more likely to occur when the growth performance is good. Our results are thus more similar to those obtained with the narrative approach. Hence, the episodes identified with our definition are less at risk of being endogenous than those of Alesina and Ardagna (2010, 2012). However, since lagged output gap or lagged growth do appear significant, the episodes identified with our definition may demonstrate less evidence for exogeneity than the narrative episodes of Guajardo et al. (2014).22 We turn now to the impact of our complete set of variables on the probability of fiscal adjustment, as portrayed in column (1) of Table 3. The inflation rate has a positive effect on the decision on fiscal adjustment, but only at the 10% significance level. The long-term interest rate plays a significant role in prompting fiscal adjustment at the 1% significance level: high long-term interest rate imposes greater burden on government debt, so that it is likely to encourage fiscal adjustment. As for the fiscal variables, the primary balance of the previous year plays a significant role. A rise in the initial primary balance by 1% of GDP decreases the likelihood of fiscal adjustment by 2.2%. In contrast, the initial debt-to-GDP ratio is only weakly associated with fiscal adjustment: it is significant only at the 10% significance level and the size of the effect is very small. Finally, most political variables turn out insignificant. Specifically, there is no evidence supporting the ‘political budget cycle’ story, whereby the incumbent adopts expansionary fiscal policy in an election year to stimulate the economy so as to increase the chances of re-election. Instead, we find that the probability of adopting fiscal adjustment does not decrease significantly in the year of general election. This may be because our data are composed of only OECD countries with a higher level of development, democracy and greater transparency.23 Finally, federal nations are more likely to undertake fiscal adjustment.","With respect to the estimation, we follow the same methodology as Guajardo et al. (2014) and Alesina and Ardagna (2012). First, we estimate equation (6) by ordinary least squares and then compute the estimated cumulative responses of real GDP and its components to a shock of 1% point change in the CAPB-to-GDP ratio for the first three years in order to measure the response on the level of real economic activity variables in the log terms.25 We calculate the standard errors of the impulse response using the delta method.26 Table 4 presents the estimated coefficients of the changes in the CAPB on the economic activity variables in our baseline model. The first column reports that growth responds negatively to contemporaneous changes in the CAPB, but positively to its lagged change. As the negative effect of contemporaneous fiscal adjustment is much larger than the lagged positive effect,27 the fiscal adjustment is found to have contractionary effect in the short term. In other words, non-Keynesian effects, or expansionary fiscal adjustments, do not occur. In that, our results are very similar to those of Guajardo et al. (2014) based on the narrative approach. The response of growth is mirrored in those of the individual components of GDP. The effects of current and lagged fiscal adjustment on private consumption and investment are very much in line with those on growth (columns 2 and 3). As for the labor market, the effect on the real wage is negative but insignificant. On the other hand, the effect on unemployment rate is large and positive at the 1% significance level, both contemporaneously and with a lag. This shows that fiscal adjustment reduces output and raises unemployment in the short term. Columns (6) and (7) show the impacts of fiscal adjustment on interest rates. Interest rates fall when a country's fiscal position improves, which is consistent with the finding of Ardagna (2009). Table 5 shows the corresponding impulse-responses resulting from an improvement in the CAPB by 1% of GDP for three years following fiscal adjustment, based on the results in Table 4. The growth rates are cumulated to obtain the estimated impact of fiscal adjustment on the level of economic activity, following Guajardo et al. (2014) and Alesina and Ardagna (2012). Fiscal adjustment has statistically significant effects on GDP, private consumption and other macroeconomic variables with the peak contractionary effect occurring within 1 or 2 years. In particular, a fiscal adjustment equal to 1% of GDP reduces real GDP by about 0.3% in the year of fiscal adjustment. These results are very similar to those of Guajardo et al. (2014), despite the different definition of fiscal adjustments, specification and data. Fig. 3 compares the responses of GDP to a fiscal adjustment shock between our baseline and Guajardo et al.'s (2014) baseline. Although the timing of peak contractionary effects is different, both sets of results report negative effects on GDP sustained for three years and diminishing gradually over time.28 In summary, our results suggest that fiscal consolidations have a significant contractionary effect in the short term. In addition, although Guajardo et al. (2014) raise some issues with respect to using the CAPB, our own CAPB-based measure, which takes account of those criticisms, shows results which are very similar to theirs.","Next, we present the results of several alternative approaches to test the robustness of our results. First, given that both our definition for computing the CAPB and the criteria for identifying fiscal adjustment have an element of arbitrariness in them, we consider alternative definitions. Second, as discretionary fiscal policy cannot be entirely exogenous to the state of the economy, we address the issue of endogeneity in our model. Third, we investigate the role played by the composition of fiscal adjustment in terms of tax increases and spending cuts. Fourth, we check the robustness of our findings to the inclusion of other variables to control for monetary and exchange rate policies, as well as the institutional and political environment. Finally, we also investigate the sensitivity of our results across country groups. Alternative definitions of fiscal adjustment ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our CAPB measure and the resulting definition of fiscal adjustment are different from the extant literature using the CAPB approach. In particular, the thresholds applied are arbitrary to some extent. Therefore, we perform additional analysis to assess whether changes in the threshold affect critically the results. First, we use alternative thresholds with respect to the standard deviation of CAPB changes. Second, since the average and standard deviation of the changes in the CAPB for each country can be affected by outliers, we drop the largest positive and negative values representing the changes in the CAPB. Third, we replace the share price index with the house price index.29 Fourth, we use the CAPB reported by the OECD instead of computing it ourselves.30 Finally, it may be argued that the assumption of unemployment proxying business-cycle fluctuations is a strong one. Therefore, we replace it with the output gap.31 Tables 6 and 7 show that the baseline results (column 1) are robust to a series of alternative criteria for the definition of fiscal adjustment and also to alternative CAPB definitions. In particular, in column (2), we apply a lower threshold for identifying fiscal adjustment while columns (3) and (4) apply stricter thresholds. In each case, fiscal adjustment has a similarly sized negative effect on growth. Dropping the outliers (column 5) similarly does not lead to substantially different results. The results obtained when the house price index is used instead of stock prices (column 6) are not much different from the baseline either. The results with the OECD CAPB data show an insignificant negative effect (column 7). This difference is likely to be due to the different assumptions and methodology. As Alesina and Perotti (1995) and Alesina and Ardagna (1998) point out, the OECD method depends on measures of potential output, which are regarded as highly arbitrary, and a set of elasticities of taxes and expenditures.32 In addition, although the OECD purports to eliminate one-off transactions from the primary fiscal balance, this adjustment may not be perfect because one-off transactions in its methodology are derived from the deviation from trend in net capital transfers, not from individual records. For instance, the Netherlands in 1996 is one of the cases that historical records indicate as having a one- off transaction in the previous year, but is included in the fiscal adjustment episodes according to the OECD CAPB version.33 Given the methodological differences, we feel that the fact that we obtain different results with the OECD CAPB figures actually strengthens our case. Next, we replace unemployment with output gap as a proxy for business cycle fluctuations. In column (8), we use the output gap instead of unemployment both for government revenue and spending, while in column (9), we use unemployment for spending and output gap for revenue. The results are again broadly in line with the baseline.34 Table 7 compares the impulse-responses based on the estimation results. The stricter the thresholds, the more negative the effect of fiscal adjustment on GDP. When dropping outliers, the negative effects are smaller than in the baseline, but are still significant. While the response in case of OECD CAPB is not significant, most estimates indicate a decline of GDP similar to the baseline result for three years. Using output gap instead of unemployment as a proxy for business-cycle fluctuations results in a stronger (more negative) output response, especially when we use the output gap for both government revenue and spending. Finally, in the preceding results, we used a single time trend for the whole period, 1970–2009. Alesina and Ardagna (1998), in contrast, use two time trends: 1960–1975 and 1976–1994. While we do not have enough data to estimate a separate trend for the 1970s (to account for the oil-price shocks that took place during that decade, which was the motivation for the two time trends employed by Alesina and Ardagna, 1998), we re- estimated our results with two time trends based on the adoption of the Maastricht Treaty in 1992: 1970–91 and 1992–09. The impact of this change is minimal and the results are very similar to those obtained with a single time trend (not reported but available upon request). Endogeneity of fiscal adjustment ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Second, we estimate the effects of fiscal adjustment on growth under the assumption that the decisions on fiscal adjustment and its size are endogenous to the state of the economy. This means that since the cyclical correction cannot remove the automatic changes of fiscal variables in response to output entirely, some of the discretionary fiscal changes are related to the fluctuations of contemporaneous output. As a result, the fiscal adjustment variable (ΔFAi,t-j) can be correlated with the contemporaneous error term (νi,t): E(νi,t |ΔFAi,t-j) ≠ 0. Therefore, similar to Hernández de Cos and Moral-Benito (2011), who also take the endogeneity into consideration, we estimate the effect of fiscal adjustment via two-stage least squares (2SLS).35 We select the fiscal adjustments based on the narrative approach by Guajardo et al. (2014) as the first instrument because it is more likely to be exogenous given that the identification is based on historical records. In addition, we use lagged long-term interest rate which we showed to be strongly correlated with fiscal adjustment in the logit analysis of section 4 and is predetermined but not strictly exogenous to the contemporaneous error term. The results are reported in Tables 8 and 9 . First, when including the changes in the CAPB during normal periods, the magnitude of the coefficient of fiscal adjustment (Table 8) and the response of GDP to a fiscal adjustment shock (Table 9) are somewhat larger than in the baseline model, but otherwise the results change little: we still observe contractionary effects which are very similar to the baseline. Importantly, the changes in the CAPB that are not associated with fiscal adjustment (NFA) do not have any effect on growth, as expected. This means that the assumption of fiscal adjustment being different from other changes in the CAPB in normal periods is reasonable. Next, when using instrumental variables to control for potential endogeneity of fiscal adjustment, the effect on growth is stronger (more negative) than in the baseline. This pattern appears regardless of the instruments used. Table 10 reports the results of first stage regressions, confirming the validity of the instruments considered. Both instruments have strong relation with fiscal adjustment. However, the test results indicate that the narrative fiscal adjustment of Guajardo et al. (2014) has more explanatory power than the lagged long-term interest rate. When we use both instruments at the same time, the results are also rather similar to those obtained with each instrument alone. In addition, according to the Durbin–Wu–Hausman test for endogeneity of fiscal adjustment, the null hypothesis that the fiscal adjustment can be treated as exogenous is rejected at the 5% significance level. Therefore, we can conclude that fiscal adjustment identified by the changes in the CAPB is not strictly exogenous to growth. Nevertheless, the results corrected for endogeneity of fiscal adjustment support our baseline finding of contractionary effects. Does composition of fiscal adjustment matter? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Many studies analyze whether the effects of fiscal adjustment depends on its composition. They generally agree that fiscal adjustments focusing on the spending side are more likely to have expansionary effects on GDP than those on the tax side. Therefore, in this subsection, we investigate what role the composition of fiscal adjustment plays in the response of economic growth. First, we divide the fiscal adjustments episodes into two types: ‘spending-based’, in which the change in the CAPB is mainly (by at least 50%) due to spending cuts, and ‘tax-based’, in which the change in the CAPB is mainly due to revenue increases (Guajardo et al., 2014 and McDermott and Wescott, 1996, apply the same criterion). Next, we implement another categorization of fiscal adjustments, splitting them into three types: ‘pure spending-based’ ones where the improvement in the CAPB is entirely due to spending cuts, ‘pure tax-based’ ones which are totally due to revenue increases, and ‘mixed’ cases that combine the two types of adjustment. Fig. 4 shows the estimated effects of fiscal adjustment according to its composition. First, although spending-based adjustments do not have a significant expansionary effect, they also do not have significantly negative effect on GDP and private consumption except in the year of fiscal adjustment. When compared with tax-based adjustments, spending-based adjustments are less contractionary and can even offset the large negative effects of tax-based adjustment because the response of the baseline is in between the responses associated with the two types of fiscal adjustment. This result is consistent with Alesina and Ardagna (2012). On the other hand, tax-based fiscal adjustments have a contractionary and statistically significant effect on GDP with a peak negative effect of −0.68% and on private consumption with a peak negative effect of −0.71% within three years. When the composition of fiscal adjustment is classified into three types, as shown in column (2) of Fig. 4, the results do not differ much. While the results for mixed adjustments are almost the same as the baseline, pure tax-based fiscal adjustments decrease GDP significantly whereas pure spending-based fiscal adjustments appear contractionary but the effect is not statistically significant even in the year of fiscal adjustment. An alternative way is to identify fiscal adjustments based on large changes of fiscal variables rather than by looking at changes of fiscal balance: as an increase in cyclically-adjusted revenues or a decrease in cyclically-adjusted spending. Although this method is different from the conventional method based on fiscal balance, it has several advantages. First, we can capture some episodes of fiscal adjustment which might be otherwise excluded. This is the case when the fiscal adjustment on spending (revenue) side is offset by a counter- balancing change of revenue (spending). Second, we can reduce the risk that the results are driven by a particular threshold (e.g. 50%) chosen to identify tax-based and spending- based adjustments. Therefore, we re-identify fiscal adjustments based on large changes of cyclically-adjusted revenues and spending respectively.36 The former are denoted as ‘tax side’ and the latter denoted as ‘spending side’. Then, we replace ΔFA in the baseline specification with these two types of fiscal adjustments to estimate their effects on GDP and private consumption. Fig. 5 shows the estimated effects of fiscal adjustment according to its composition. While fiscal adjustments based on an increase in revenues have a largely contractionary and statistically significant effect on GDP and private consumption, fiscal adjustments based on a decrease in spending have a small expansionary but not statistically significant effect on GDP and negligible effects on private consumption. Therefore, we still cannot find any firm evidence of expansionary effects even in the case of fiscal adjustment based on large spending-cuts. However, this reconfirms that spending-based adjustments are less contractionary than tax-based ones, which is consistent with the previous results of compositions of fiscal adjustment based on the CAPB. Guajardo et al. (2014) argue that a possible reason for the different effects depending on the compositions of fiscal adjustment is that monetary policy is more favorable in case of spending cuts. They suggest that central banks conduct monetary stimulus more actively following spending cuts than after tax hikes so that the policy rate increases in response to tax hikes and decreases in response to spending cuts.37 Therefore, we investigate the response of short-term interest rate to the two types of fiscal adjustment. As Fig. 6 shows, the response of the short-term interest rate is significantly different according to the two types of fiscal adjustment only in the year of fiscal adjustment. After one year, the short term interest rate falls significantly in both cases. Therefore, this result can partially support the argument of Guajardo et al. (2014) that the different effects depending on the composition of fiscal adjustment are ascribed to different monetary policy stances. The role of the economic environment ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Next, we check the robustness of our findings by including the short-term interest rate and the real effective exchange rate among the regressors.38 These two additional control variables are aimed at accounting for monetary policy and exchange rate policy respectively. Table 11 and Fig. 7 show the results. The fit of the regression improves when we include variables relating to the economic policy. The coefficient of fiscal adjustment, as well as that for tax-based adjustment, remains significantly negative, although they are smaller than those without controlling for policy variables. Similarly, spending-based fiscal adjustment has a smaller negative coefficient, but is still statistically insignificant. The change of effects can be attributed to monetary policy in that the short-term interest rate has a significantly negative effect on growth, as expected, but the exchange rate is not significant. Fig. 7 confirms that fiscal adjustments have less contractionary effect on GDP when we control for monetary policy than in the baseline. Therefore, monetary policy can affect the response of GDP to fiscal adjustment shocks. If the short-term interest rate falls, it leads to an increase in GDP. Therefore, if fiscal adjustment coincides with a large reduction in the short-term interest rate, this may stimulate the economy in the following periods. However, even in this case, this result is to be attributed not to the fiscal adjustment, but to the lax monetary policy. In regard to the effects of composition of fiscal adjustment, Fig. 7 shows that the response of GDP is larger in tax-based fiscal adjustments than in spending- based ones. Therefore, as Fig. 6 in the previous subsection shows, if the discretionary monetary policy responds differently according to the type of fiscal adjustment, this could help account for the different effects depending on the composition of fiscal adjustment. However, it cannot be a decisive factor, contrary to the argument of Guajardo et al. (2014), in that when we control for the short-term interest rate, there is still a large difference between the effects of tax-based and spending-based fiscal adjustment on GDP. Furthermore, there can be other omitted factors that are likely to influence the effects of fiscal adjustment on economic activity. The omission of relevant variables could bias the response of output in estimating the effects of fiscal adjustments. Therefore, we add additional variables into the baseline model one by one. The initial government debt is considered because a high debt level can lead to expansionary fiscal adjustment via the wealth effect. International factors such as the exchange rate regime and financial openness can be taken into account as another potential factor. Furthermore, Ilzetzki et al. (2011) find that the degrees of exchange rate flexibility and openness to current and capital account transactions are critical determinants of the size of fiscal multiplier. Therefore, we include the exchange rate regime and financial openness index.39 We also include dummies for Euro Area membership and for participation in IMF programs. Political environment can similarly play a role and therefore we include a dummy for countries with a federal system of government and a measure of ruling party ideology (the latter is obtained from Potrafke, 2009). Finally, we also include lagged inflation and the long-term interest rate as two potentially important macroeconomic variables. Tables 12 and 13 show the regression results obtained when we control for the impact of these additional factors. The results are very similar to the baseline results, as is clear from column (10) of Table 12, where all the variables are included in the same regression. Here GDP growth (T-1), ΔFA (T) and ΔFA (T-1) remain highly significant, as in the baseline, and the signs as well as magnitudes of their coefficients remain remarkably similar. Moreover, all of the additional controls turn out insignificant in column (10), with the exception of gross debt, inflation and long-term interest rates, all of which have a negative effect on growth, which is intuitive. As regards columns (1)–(9), where one variable at a time is added to the baseline, we find that only the exchange rate regime and the long-term interest rate are significant. Aside from the fairly comprehensive list of variables considered, regulatory reform such as labor and product market institutions, and structural reforms should be considered as significant and relevant factors influencing the estimated effects of fiscal adjustment on economic activity. Several studies investigate the interactions between fiscal adjustment and these market institutions and structural reforms and show that these regulatory policies can play a significant role in initiating fiscal adjustment and determining its success (Guichard et al., 2007; Hauptmeier et al., 2006; Tagkalakis, 2009).40 However, when these are controlled for, the qualitative effects on economic activity of fiscal adjustment does not change (Alesina and Ardagna, 2012; Hernández de Cos and Moral-Benito, 2011). Although we do not address the effects of labor and product market institutions and structural reforms during the fiscal adjustment episode in this paper, they can affect the responses of output to fiscal adjustment in the long term as well as quantitatively via employment and investment behavior. Effects of fiscal adjustment across country groups ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The effects of fiscal adjustment on the economic activity may be different according to the sensitivity of private agents, which in turn depend on the past trajectory of fiscal policy. While our definition of fiscal adjustment controls for country-specific heterogeneity, in this subsection, we explore this issue further by dividing the 20 countries into two groups on the basis of two criteria: the frequency of fiscal adjustments and the volatility of discretionary fiscal policy. For the first criterion, high and low frequency groups include 10 countries each. Similarly, high and low fluctuation groups each consist of 10 countries according to the standard deviation of changes in the CAPB.41 Table 14 reports the estimated responses of GDP and private consumption to a fiscal adjustment shock. Interestingly, for the high group in terms of both frequency and fluctuation, economic activity displays a significantly negative response only in the year of fiscal adjustment. On the other hand, the low groups in frequency and fluctuation alike display a strong response to fiscal adjustment in all years. This finding supports the notion that economic agents respond more sensitively to unexpected or unusual shocks. When fiscal policy undergoes frequent changes, agents become accustomed to such changes and their responses become smaller.","This paper investigates the short-term macroeconomic effects of fiscal adjustment in 20 OECD countries over the period 1970–2009. This issue has been studied in many previous contributions already. Recently, it has become more central in academic and policy circles due to the rising fiscal deficits and public debts during the current global crisis. Much of the literature argues that fiscal adjustments can promote economic output even in the short term. However, after identifying fiscal adjustment episodes from historical documents, Guajardo et al. (2014) conclude that fiscal adjustments are always contractionary. They also criticize the CAPB-based measures used in the rest of literature as being imprecise and biased toward overstating the potential expansionary effects of fiscal adjustments. This paper reconsiders the CAPB-based measure in order to identify the fiscal adjustment episodes more accurately, taking into account the problems identified by Guajardo et al. (2014). The main features of our new measure of fiscal adjustment are as follows. First, we consider the effect of asset price fluctuations on tax revenues when cyclically correcting the fiscal balance. Second, our criteria for selecting fiscal adjustment episodes allow for the heterogeneity of individual countries in fiscal policy, contrary to the uniform approach in the previous literature. Third, our criteria account for temporary one-off transactions and temporary adverse shocks which can undermine the accuracy of the CAPB. Although Guajardo et al. (2014) argue that the CAPB is an unreliable guide regarding fiscal adjustment, our new criteria identify fiscal adjustment episodes that largely overlap with their narrative-based ones. Based on the fiscal adjustments identified, we estimate the effects of fiscal adjustment on economic activity. Our key result is that a fiscal adjustment has contractionary effects on economic activity in the short term. This provides little support for the expansionary fiscal adjustment hypothesis: such ‘non-Keynesian effects’ are very limited and probably occur only under specific conditions, not generally. This is consistent with the results of Guajardo et al. (2014). As for the role of the composition of fiscal adjustment, spending-based fiscal adjustments lead to smaller reductions of output than tax-based fiscal adjustments. This finding is in line with most of the literature regardless of the approach used. Further work could explore in more depth the effects of fiscal adjustments. First, as for the reasons behind the different effects of tax-based and spending-based adjustments, more detailed disaggregation of fiscal spending and taxes could be used for the analysis. Second, most of the literature on fiscal policy has studied developed countries such as the OECD because of data limitations. However, as the data for developing countries have become more easily available lately, the fiscal adjustment in developing countries also needs to be investigated for the comparison with developed countries' results. Another possible extension would address the anticipation effects by private agents through comparing the narrative data based on announced plans with the CAPB-based data based on actual outcomes. However, to capture the timing of fiscal adjustment accurately, quarterly rather than annual data may be required."],["We examine the evolution of market potential and its role in driving economic growth over the long twentieth century. Theoretically, we exploit a structural gravity model to derive a closed-form solution for a widely-used measure of market potential. We are thus able to express market potential as a function of directly observable and easily estimated variables. Empirically, we collect a large dataset on aggregate and bilateral trade flows as well as output for 51 countries. We find that market potential exhibits an upward trend across all regions of the world from the early 1930s and that this trend significantly deviates from the evolution of world GDP. Finally, using exogenous variation in trade-related distances to world markets, we demonstrate a significant causal role of market potential in driving global income growth over this period. --------------------------------------------------------------------------------","What has been the trajectory of market potential over the long twentieth century? And is there a causal relationship between market potential and global income growth? Here, we contribute to a long-standing literature along these lines by developing a structural measure of market potential which is fully comparable across countries and across time and which comes with fairly minimal data requirements. As in the preceding literature, we model market potential as a summary measure of both external and internal demand which explicitly takes into account the costs of transaction and transport associated with the exchange of goods. But in contrast to much of the preceding literature, we seek to exploit the wide variation in the evolution of the global economy over the long twentieth century to investigate the causal role of market potential in shaping global income growth over this period. Our goal, then, comes in assessing the relationship between globalization and growth in the long run by: (1) developing a theoretically- derived measure of market potential appropriate for historical use rather than relying on narratives derived from “data-as-given” time series such as aggregate exports, allowing us to relate globalization and growth in a more disciplined fashion; (2) collecting a new dataset on aggregate exports, bilateral trade, and GDP for 51 countries; (3) constructing our proposed measure of “structural market potential”, as well as charting and decomposing its evolution through time; and (4) exploiting exogenous variation in trade-related distances to world markets in order to determine the causal role of market potential in driving global income growth over this period. Of course, we are far from the first to consider the theme of market potential and its role in the growth process. In the very first contribution to this literature, Harris (1954) was motivated by the question of why, with only 12% of the United States by area, the Northeast produced fully 50% of its manufacturing output and employed 70% of its industrial labor force in 1950. His informal model is one in which firms balance production versus trade costs in determining their location and in which the presence of deep input and output markets influence this decision. His paper also marks the first usage of the term market potential which Harris defines as “an abstract index of intensity of possible contact with markets” and is calculated as the sum of markets accessible (often proxied by income or population) to a given point over distance-to- markets from that point. Krugman (1991, 1992) resurrected this notion of market potential but explicitly grounded it in a spatial general equilibrium model, thereby setting off an expansive body of work in the new economic geography literature. The basic structure of the Krugman model was then extended by Helpman (1998), enhancing its tractability in empirical work (e.g., Hanson 2005). In addition, Fujita et al. (1999) gave rise to the workhorse model of the new economic geography literature. Importantly, these modeling approaches rely on a common set of elements, typically in the form of CES consumption, simple production functions and monopolistically competitive firms. Symmetry in preferences and technology yield a structural link between market potential and standards of living. For our purposes, however, the most important contribution to this literature comes from Redding and Venables (2004). Motivated by the wide dispersion in cross-country manufacturing wages and incomes, they concentrate on two mechanisms which may potentially explain such disparities: (1) the distance of countries to markets in which their output is sold; and (2) the distance of countries to markets from which capital and intermediate goods are purchased. Thus, the presence of trade costs means that more distant countries face a penalty on their sales as well as additional costs on imported inputs. As a consequence, firms in these countries can only afford to pay relatively low wages, translating into lower levels of GDP per capita. This result holds even if technologies are the same across countries. Liu and Meissner (2015) recently considered the theme of historical market potential in the context of the Redding and Venables (2004) model. Using cross-sectional data for 27 countries in 1900 and 1910, they establish that market potential was a significant determinant of GDP per capita in the early twentieth century. They also raise the prospect that the United States did not necessarily benefit from a natural lead in market potential as its greater domestic market size was counterbalanced by its greater distance to other—in particular, European—markets. Finally, Head and Mayer (2011) consider panel evidence for the role of market potential in driving differences in GDP per capita for the period from 1965 to 2003. Thus, they are able to establish a broader consistency with the results of Redding and Venables (2004). However, we argue that there is a complication with Redding and Venables' approach in a panel setting, which makes its use in a historical context potentially problematic. Namely, one ideally needs a full matrix of bilateral trade flows for every year, imposing a large cost in terms of data collection. This is due to the fact that the construction of market potential in Redding and Venables (2004) and Head and Mayer (2011) relies on a set of exporter and imported fixed effects pertaining to all countries in the world, based on a standard gravity model of bilateral trade. Without the full matrix of global trade flows, estimates of these fixed effects can shift substantially.1 Our proposed solution, then, comes from exploiting a link between the model of Redding and Venables and structural gravity models that allows us to bypass exporter and importer fixed effects. The rest of the paper proceeds as follows. Section 2 lays out the relationship between market potential and structural gravity. It does so first by revisiting the work of Redding and Venables (2004) on market potential and then by relating it to the work of Anderson and van Wincoop (2003). This results in a new solution for the measure of market potential that is less data-intensive and therefore particularly suitable for historical settings. We refer to it variously as “structural market potential” or, simply, market potential. Section 3 introduces the underlying data, presents our new evidence on market potential over the long twentieth century, and provides a comparison to existing formulations of market potential. Section 4 first relates our new measure to global growth in the context of standard wage equations drawn from the existing literature and then exploits variation in trade-related distances to world markets in order to establish a causal relationship between market potential and global income growth. Finally, Section 5 concludes.","We first outline the basic setup of the Redding and Venables (2004) new economic geography model. We then relate it to the structural gravity framework by Anderson and van Wincoop (2003). As a departure from the existing literature, this allows us to derive an analytical solution for the market potential measure mainly in terms of directly observable and easily estimated variables. The new economic geography model ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Apart from the demand-side aspect captured by market capacity, the right-hand side of eq. (2) also contains supply-side variables in the form of nipi1−σ, net of bilateral trade costs tij. Redding and Venables (2004) refer to this term as the supply capacity of country i, si ≡ nipi1−σ. It consists of an extensive margin measure ni for the number products originating in i as well as their price competitiveness embodied by pi. Redding and Venables (2004) provide further details for the supply side of the model. For instance, they impose a Cobb-Douglas technology with an immobile factor (e.g., labor), an internationally mobile factor (e.g., capital), and a composite intermediate good with price Pi. They introduce increasing returns by way of a fixed input requirement.3 It turns out, however, that the supply-side details are not essential for the aggregate gravity relationship that emerges from the model as the basis for the empirical analysis.4 Exploiting the link with structural gravity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ All else being equal, MAi increases in global income. Intuitively, if the global economy grows, demand for individual country i's output rises. In contrast, SAi decreases since the growth of production in the world represents more competition and thus a decline in relative supply capacity. Not surprisingly, growing yi increases both market and supplier access since it represents both rising availability of supply to customers elsewhere as well as rising demand for foreign products. Higher domestic trade costs tii work in the opposite direction since they hamper the domestic economy. We can think of tii as the cost of reaching domestic customers and sourcing domestic supply. The role of domestic trade flows xii is perhaps less obvious to understand. Holding output constant, due to market- clearing a rise in xii means less trade with foreign countries, which implies that bilateral trade costs tij must have risen. A rise in bilateral trade costs is associated with more isolation from global markets, which in turn hurts demand prospects as well the ability to obtain the supply of goods emerging from partner countries. We draw two conclusions for our empirical analysis. First, since the market and supplier access measures are proportional, for a given cross-section they do not contain independent information. Therefore, we proceed with a single measure corresponding to the expression for market access in (12). We variously call it “structural market potential”, or, simply, market potential (MPi). Second, in contrast to the previous literature, we do not necessarily require estimates of or information on bilateral trade costs across countries to compute our market potential expression (12). Instead, it is a simple function of variables related to the domestic economy and a global constant.9 Moreover, the variables in eq. (12) are for the most part given by the data. That is, income yi as well as global income yW are directly observable while domestic trade flows xii can be constructed from the data. Domestic trade costs scaled by the elasticity of substitution, tiiσ−1, can be constructed based on estimates from a standard gravity regression using domestic trade cost proxies such as internal distance as we do here (details below). The dataset ~~~~~~~~~~~ We collected a large annual dataset for 33 countries over the period from 1910 to 2010 which is comprised of aggregate exports and GDP and for an additional 18 countries from 1950 to 2010. We chose to begin data collection in 1910 in order to maximize the cross- section of countries at our disposal. This data includes newly collected trade observations, in particular, for the periods spanning the World Wars. We provide details on our sources in Appendix III while Fig. 1 summarizes the sample graphically. Countries in black (n = 33) are those for which the full complement of output and trade data is available from 1910 while those in grey (n = 18) are those for which consistent data is available only from 1950. The sample countries represent roughly 75% of world GDP in 1910, roughly 85% of world GDP in 1950, and roughly 90% of world GDP in 2010. Constructing market potential ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We construct our preferred measure of market potential as given in eq. (12). This approach does not require estimation of the entire term as we simply insert the data directly into the right-hand side expression. Thus, the data on income yi and global income yW are readily available. We construct domestic trade as the difference between income and total exports, xii = yi – xi, where xi denotes total exports. As income is measured as GDP and is cast in value-added terms, it is in principle not consistent with exports as a gross- value measure. However, as a robustness check, we are able to use total gross manufacturing production instead of GDP for the years from 1980 to 2006. This leaves our main results materially the same, a point we discuss in fuller detail in Appendix II.11 As a caveat, we stress that a shortcoming of the measure for tii is that a number of components that arguably matter for domestic trade costs, such as domestic infrastructure, are left out. Given that the distance and contiguity measures do not change over time, the changes in domestic trade costs are driven by time-varying gravity coefficients. It would be preferable to have a more detailed, country-specific time-varying measure of domestic trade costs, but limited data availability poses restrictions in that regard. Measuring domestic trade costs is an active area of research (e.g., Ramondo et al. 2016), and better measures might improve market potential measures such as ours in the future. Fig. 2a shows the average of the log values of market potential for two samples, indexed to a value of 100 in 1950. The first sample is comprised of all 33 countries for which we have a complete set of aggregate export and GDP data from 1910. The second sample is comprised of the same plus the 18 countries for which we have a complete set of aggregate export and GDP data from 1950. While Fig. 1 might suggest that there may be non-random sample selection across the start dates of 1910 and 1950, Fig. 2a indicates this is likely a non- issue as the correlation between the two series is in excess of 0.99. In general, there is a clear upward trend driven by the growth of the world economy, with periods of global depression and recession in the early 1930s, the early 1980s, the early 1990s, and the late 2000s registering as troughs in the series. Underlying these global patterns is substantial heterogeneity with large and persistent—albeit unreported—differences in the levels of market potential across continents, e.g., Latin versus North America or Asia versus Europe. At the same time, Fig. 2a plots the log value of world GDP, also indexed to a value of 100 in 1950. We do so to assuage concerns that our measure of market potential captures nothing more than the evolution of world economic activity over the long twentieth century. For sure, the various series for market potential and world income exhibit a somewhat similar upward trajectory, but it is clear from Fig. 2a that there is very little variation in world income growth from year to year. In contrast, our measures of market potential register significant divergence from world income. And it is precisely this variation which we will use below to identify the causal effect of market potential on economic growth at the individual country level. Based on eq. (14), we can also extract and plot the implied price index Pi by removing world income from the market potential measure. Here, we assume a value for the elasticity of substitution of σ = 5. In Fig. 2b, we plot this implied price index for two key economies, India and the United Kingdom, normalized to 100 in the year 1910. How should we interpret this implied price index? Consider the following benchmark case. If trade costs did not change and the world experienced uniform income growth across all countries, then the price indices would not change.14 In that case, market potential would follow exactly the same trend as global income over time. By contrast, higher trade cost levels serve to increase these price indices. This is precisely what we observe from 1910 to 1930, reflecting rising protectionism in the interwar period. More specifically, the price index rose by 92% for India and 70% for the United Kingdom. This rise is then followed by falling price indices, reflecting a long-run trend of declining trade barriers and increasingly open economies. Overall, the implied price indices can be interpreted as an inverse proxy of our “structural market potential” measure. We stress that these price indices are not the same as conventional consumer price indices (and therefore not directly observable) since they may also capture non-pecuniary trade frictions such as information barriers (see Anderson and van Wincoop 2003). In this vein, we note that there is significant variation in market potential—particularly in relative terms—across individual countries. As an example, Fig. 3 speaks to this issue by considering the trajectories of the log of market potentials for India and the United Kingdom over the long twentieth century. There, it is apparent that while much of the variation in the two series is shared in common—again, driven by the evolution of world income—there is still scope for differential rates of growth in market potential in the long run. This is seen most clearly in the ratio of the two series (UK:IND). It rises up to 1930 when the United Kingdom's lead attains its maximum and then consistently falls into the present day where Indian and UK market potential stand nearly at par.15 To further our understanding of the underlying spatial correlations, we compute Moran's I for various years. This measure takes on a value of −1 in the case of perfect dispersion (negative spatial autocorrelation), a value of close to 0 in the case of random spatial arrangement, and a value of +1 in the case of perfect positive spatial correlation. We compute Moran's I for the logarithmic values of our market potential measure for the sample from 1950 (n = 33), and also for logarithmic GDPs as a comparison.16 We follow the common approach of giving a weight of 1 to neighbors in the spatial weights matrix and a weight of 0 otherwise (i.e., we use the contiguity indicator variable in the spatial weights matrix). The resulting values of Moran's I for log market potential are 0.94, 0.94, and 0.88 for the years 1950, 1980, and 2010, respectively. The corresponding values of Moran's I for log GDP are 0.59, 0.53, and 0.51. We make two observations. First, market potential is more strongly spatially correlated than GDP. This finding is intuitive given that many large economies tend to be spatially clustered, e.g., in Western Europe. Second, the degree of spatial correlation declines over time but only in a minor way. This finding can be explained, for instance, by the shift of global economic power away from Europe and North America towards Asia. A comparison to Redding and Venables (2004) and Harris (1954) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The previous literature constructs market and supplier access measures (3) and (4) by estimating eq. (8) for xij where supply capacity si and market capacity mj are taken as exporter and importer fixed effects, respectively. Redding and Venables (2004) follow this procedure for a single cross-section in 1994. Head and Mayer (2011) have panel data for the period from 1965 to 2003 and estimate the fixed effects year by year. The use of exporter and importer fixed effects implies a specific normalization due to the omitted exporter/importer category. For instance, Redding and Venables (2004) omit the exporter fixed effect for the United States as a normalization and also omit the constant in their specification such that no importer fixed effect has to be dropped. In contrast, our method of constructing market potential through eq. (12) does not directly rely on exporter and importer fixed effect estimates and thus avoids the year-by-year normalization. While we readily acknowledge that Redding and Venables were only concerned with market potential across the countries of the world at a given point of time, one benefit of our measure is that it allows us to more consistently compare levels of market potential over time.17 Fig. 4 follows Fig. 3 by considering the trajectories of the log values of the market access measure, MAi, for India and the United Kingdom for the period from 1910 to 2010. As in Fig. 3, there is a fairly consistent, upward trend throughout the second half of the twentieth century and the first decade of the 21st century. However, in Fig. 4, we also observe two sharp increases in market access during the first half of the twentieth century to the extent that the average (log) values for market access for the United Kingdom in 1919 and 1946 exceed those for 2010. Taken purely at face value, this would seem to be an implausible result given what we know about global macroeconomic history, in particular the role of the World Wars in disrupting global trade flows (Jacks et al. 2011). There is also the related issued that the relative value, UK:IND, is remarkably flat, hovering around a value of 1.05 throughout the long twentieth century. Thus, while the approach of Redding and Venables (2004) is very useful for a given cross- section of data, our results suggest caution in blindly using it in repeated cross- sections particularly during periods when international trade flows are heavily distorted by global conflict. For our purposes of both charting and understanding the trajectory of market potential over the long twentieth century, we therefore prefer the measures presented in Figs. 2a and 3.18 Decomposing the growth of market potential over time ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To understand the decomposition in eq. (17), it is useful to consider the hypothetical benchmark of income growing by the same uniform rate across all countries. In that case, the income and domestic trade shares in the square brackets would not change, and market potential would be driven exclusively by overall global income growth through the last term. If one country grew faster than the otherwise uniform rate, its market potential would rise more quickly than elsewhere. In Table 1, we present the results of decomposition (17), constructing the right-hand side variables as described in section 3.2. We use our sample of 33 countries that we group by five regions (Asia, Australia/New Zealand, Europe, Latin America, and North America). We present a decomposition for the full period from 1910 to 2010 as well as separate decompositions for the periods from 1910 to 1960 and from 1960 to 2010. Overall, market potential grew by 305% across countries on average over the full period. Perhaps not surprisingly, this growth is rather similar across regions as global income growth serves as a common factor in driving market potential. However, this comparison of 1910 versus 2010 heavily masks important differences across sub-periods. In particular, countries experienced only very moderate growth in market potential prior to 1960. This was a period marked by isolationism and war with an associated rise in trade costs as well as domestic trade shares (see Jacks et al. 2011). In contrast, the period after 1960 was characterized by sizeable (positive) contributions to the growth in market potential stemming from declining trade costs and increasing openness. In particular, Asia experienced above-average growth in market potential due to its expanding share of global income while the opposite was the case for Europe.","Here, we return to one of the motivating questions for this paper, namely whether there is a causal relationship between market potential and global income growth. In what follows, we first establish that our new measure of “structural market potential” delivers results from so-called wage equation regressions which are consistent with those found in Redding and Venables (2004) and Hanson (2005) among others. However, in their work, the expression for market access is conveniently separable into two constituent components, domestic and foreign market access. Thus, the latter of these two strips away any domestically- determined elements of demand. This contrasts with our measure of “structural market potential” which will clearly be endogenous in light of the presence of domestic output in eq. (12). In order to break this mechanical link in between income (per capita) and market potential, we then use exogenous variation in trade-related distances to world markets as an instrument, finding an economically and statistically significant role for market potential in driving global income growth over the long twentieth century. Wage equation regressions in levels ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ An appropriate starting point is provided by Redding and Venables (2004). In their work, they derive what is known as a wage equation, i.e., an equation that structurally relates the price of the immobile factor of production (or wage) to a country's market access/market potential. Based on their model, the same wage equation would arise in our context. Redding and Venables demonstrate a strong association between GDP per capita (their proxy for wages) and both domestic and foreign market access in the cross-section. This association remains even after conditioning on a large number of covariates and controlling for potential endogeneity. Head and Mayer (2011) run an analogous set of panel regressions, finding results consistent with those of Redding and Venables. However, with our new measure of structural market potential, it is an open question whether this empirical regularity remains intact. Table 2 first tries to establish the simple association between the log of GDP per capita and the log of market potential. Standard errors are clustered on countries here—and in all regressions—to control for within- country serial correlation of arbitrary form.19 The coefficient reported in column (1) is precisely estimated and comparable in magnitude to that reported by both Redding and Venables (2004) and Head and Mayer (2011). Of course, there are many other potential determinants of GDP per capita, and the specification in column (2) thus controls for both common patterns over time and fixed, unobserved country-level characteristics. This estimation then relies upon variation within countries over time which is not determined by global shocks or trends such as the evolution of world GDP over time. That the coefficient actually increases in magnitude is a reassuring sign of our measure's salience. The next two columns repeat the regression for different samples. The full sample in columns (1) and (2) includes 33 countries with observations on market potential and GDP per capita from 1910 to 2010 and 18 countries with observations on market potential and GDP per capita from 1950 to 2010. A brief review of Fig. 1 suggests the former are predominantly developed nations in North America and Western Europe while the latter are mainly developing nations in Africa and Asia. Column (3), which is based on the balanced sample dating from 1910 only, shows a point estimate which is smaller than that in column (2) but which is still highly statistically significant. Given the countries that join the sample in 1950 and that are part of the sample for column (2), this might suggest that the link between market potential and GDP per capita may have become stronger over time and/or is stronger for developing nations. Column (4) excludes observations spanning the two World Wars, specifically the years from 1914 to 1919 and from 1939 to 1949. These observations may be problematic if these years entailed a breakdown in normal economic relationships or suffered a deterioration in terms of data quality. The magnitude of the elasticity between GDP per capita and market potential is virtually unaffected. In any case, we still favor the results in column (4) as it addresses the most serious concerns related to sample selection across countries and years and data quality. The final two specifications average our measures of market potential and GDP per capita over (non-overlapping) five- and ten-year periods, respectively, and represent our preferred specifications. This approach of aggregating over time can be thought of as reducing the role of measurement error in particular years as well as diminishing the potential role of domestic and global business cycles in driving the results. Across columns (5) and (6), the values of the coefficients are stable and broadly similar to column (4), again pointing to a tight—but not necessarily causal—relationship between levels of GDP per capita and market potential throughout the long twentieth century. Wage equation regressions in differences ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 3 follows the regressions of Table 2 but considers a different set of dependent and independent variables. Instead of considering the logs of GDP per capita and market potential, we follow Hanson (2005) by estimating the wage equation in log differences. This allows us to better account for potential serial correlation in the specifications of Table 2 and is closer in spirit to this paper's theme of economic growth and market potential. Comparing Table 3 to Table 2 across the various specifications in columns (1) through (6) reveals that the estimated elasticities remain statistically significant. However, they are smaller in magnitude, suggesting a plausible role for country-level, time-varying omitted variables in driving the earlier results. With respect to the specifications in columns (5) and (6) in Table 3 compared to those in Table 2, the results are more encouraging in this regard. The point estimates in Table 3 are smaller in magnitude as before, but they are statistically indistinguishable from those in Table 2. Honing in on our preferred specification in column (6), the results suggest that every percent change in market potential was matched with a roughly 0.65% change in GDP per capita. Taking this result at face value suggests that a significant share of global growth over the long twentieth century could be attributed to changes in market potential in the long run. Wage equation IV regressions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Of course, there are good reasons why these results should be approached with caution. Above all, there is clear endogeneity in any wage equation regression given the way our measure of market potential is constructed in eq. (12) as a function of domestic output. Facing a similar problem, Redding and Venables (2004) as well as Head and Mayer (2011) instrument market potential with measures of geographic centrality, namely a country's distance from Belgium, Japan, and the United States. Naturally but unfortunately, such measures do not vary over time, a condition which underlies many other possible instruments for market potential. Faced with this prospect, we instead draw inspiration from a series of papers by Feyrer (2009a, 2009b). In Feyrer (2009b), the author begins with the observation that historically the vast majority of international trade by value has been conducted via sea routes and that to this day the vast majority of international trade by physical volumes continues to be conducted in this manner. However, presently, a very large share—upwards of 40%—of international trade by value is conducted via air routes as improvements in aircraft technology and logistics have enhanced the air industry's importance in this regard. Thus, over time countries with shorter air routes to its trading partners relative to its sea routes (e.g., India) have benefited more from this exogenous technological change than those with relatively similar air and sea routes (e.g., Canada). As Feyrer notes, “[such] heterogeneity can be used to generate a geography based instrument for trade that varies over time.” In a similar vein, Feyrer (2009a) exploits the shock to the global economy embodied by the closure of the Suez Canal from 1967 to 1975. While many trade routes remained unaffected, some did not and found the distances separating markets increasing significantly. For instance, Feyrer reports that India nearly led the pack with a 30.6% increase in its trade-weighted distance to foreign markets while a country like Canada only experienced a 0.2% increase. Using this exogenous variation in distance over time, Feyrer goes on to separately estimate the effect of distance on trade and the effect of trade on income. Here, we combine both approaches. In particular, we use the great circle distances from the CEPII GeoDist database (see Appendix III) to represent distances on air routes to Japan, the United Kingdom, and the United States, critical nodes of the world economy throughout the long twentieth century. We also collect the corresponding distances for sea routes reported in Philip (1935). Conveniently, this source also delineates which sea routes utilized the various major canals of the world, e.g., the Kiel, the Panama, and the Suez.20 This information allows us to incorporate changes in the distances of sea routes introduced by the various closures and openings of these canals over the period from 1910 to 2010.21 The final step is in constructing a series on the share of US imports by value which are transported by air over this period based on Hummels (2007) and various reports of the International Air Transport Association. Table 4 reports the results of this exercise using 4128 annual observations for GDP per capita and market potential (our original sample of 4431 observations minus the 303 observations associated with Japan, the United Kingdom, and the United States). The top panel of column (1) represents the first stage regression results. In order of magnitude, effective distances to the United States, then Japan, and finally the United Kingdom all register as statistically significant. Quantitatively, these three instruments explain a significant amount of the variation in our measure of market potential, with the R-squared of the regression registering a healthy 0.25. The regression also passes a standard test of joint significance (Angrist-Pischke F test) where the null hypothesis is that the endogenous regressors are jointly insignificant and a standard test of under-identification (Angrist-Pischke underid. test) where the null hypothesis is that any particular endogenous regressor is unidentified. The bottom half of column (1) represents the second stage regression results. There, the elasticity between market potential and GDP per capita is estimated to be 0.41, or about half the size of the equivalent estimate reported in column (2) of Table 2. However, this elasticity is precisely estimated and, in combination with the fixed effects, captures a majority of the variation in GDP per capita across space and time. Furthermore, the second stage passes a standard test of over-identification (Hansen J statistic) where the null hypothesis is that the included instruments are uncorrelated with the error term and that the excluded instruments are correctly excluded from the estimating equation. Again, we replicate the same set of results as in previous tables by using full versus restricted samples (columns (1) versus (2) and (3)) and by averaging dependent and independent variables over increasingly large periods of time (columns (4) and (5)). All of the coefficients are precisely estimated, fall within the range of 0.41 and 0.47, and are smaller than their OLS counterparts, suggesting a potential role for endogeneity in driving our previous results. At the same time, across all specifications a significant portion of the variation in GDP per capita is explained by the instrumented version of our market potential variable. Naturally, standard concerns regarding the exogeneity of our instruments and the exclusion restriction may remain. We therefore prefer to interpret these results as suggestive and not definitive. Nevertheless, they add to a growing body of literature that provides evidence of causal effects arising from changes in market potential. Apart from Feyrer (2009a, 2009b), this literature includes the contributions by Hanson and Xiang (2004) who examine home market effects in the exports of OECD countries across industries as well as Redding and Sturm (2008) who exploit the division of Germany after World War II and its subsequent reunification.","We develop a new approach to the old notion of market potential. Developing a structural gravity model of trade, we show that market potential can be expressed as a function of directly observable variables such as domestic trade flows and output and easily estimated proxies for domestic trade costs. We derive this expression by solving for multilateral resistance price indices across countries. These indices indirectly capture bilateral trade costs and therefore contain variation that is essential for computing market potential. Our approach has two key advantages. First, our measure is straightforward to compute. As we do not need to add up exporter and importer fixed effect coefficients, it offers an alternative to the more onerous construction of traditional market potential measures. Second, our measure of market potential naturally lends itself to comparisons over time, not only in the cross-section. On the empirical side, we construct market potential measures for 51 countries over the period from 1910 to 2010. We find that market potential exhibits an upward trend across all regions of the world from the early 1930s and that this trend significantly deviates from the evolution of world GDP. Finally, we also show that our measure of market potential is closely associated with average incomes, both in the cross-section and over time and across various specifications. Most importantly, we exploit exogenous variation in trade-related distances to world markets generated from changes in logistics and transport technology along with geopolitical events in order to assign a causal role for market potential in driving global income growth over this period."],["I develop a tractable macro model with endogenous asset liquidity to understand monetary-fiscal interactions with liquidity frictions. Agents face idiosyncratic investment risks and meet financial intermediaries in competitive search markets. Asset liquidity is determined by the search friction and the cost of operating the financial intermediaries, and it drives the financing constraints of entrepreneurs (those who have investment projects) and their ability to invest. In contrast to private assets, government bonds are fully liquid and can be accumulated in anticipation of future opportunities to invest. A higher level of real government debt enhances the liquidity of entrepreneurs' portfolios and raises investment. However, the issuance of debt also raises the cost of financing government expenditures: a higher level of distortionary taxation and/or a higher real interest rate. A long-run optimal supply of government debt emerges. I also show that a proper mix of monetary and fiscal policies can avoid a deep financial recession. --------------------------------------------------------------------------------","Asset liquidity captures the ease with which financial assets can be traded without strongly affecting their prices. The recent long-lasting world-wide financial crisis has shown that liquidity fluctuations in asset markets can have a huge impact on asset prices and the real economy.1 In fact, empirical evidence points to procyclical variation in the liquidity of a wide range of financial assets.2 Understanding the frictions that generate these fluctuations and the possible policy responses seems thus crucial. This paper studies monetary–fiscal interactions when movements in asset prices and the real economy are driven by endogenous fluctuations in liquidity. Although the research on monetary–fiscal interactions at least go back to Sargent and Wallace (1981), Leeper (1991), and Sims (1994),3 the consideration of endogenous liquidity frictions is usually missing. To fill this gap, I incorporate these frictions into an almost standard New Keynesian model with sticky prices. In the model, government bonds are fully liquid and represent public liquidity, while privately issued claims, traded on competitive search markets, are only partially saleable and have a bid-ask spread. Private claims represent private liquidity and carry a liquidity premium, a premium that buyers will demand when the assets cannot be easily converted into consumption goods at the fair market value. I show that the government faces a tradeoff between the benefit of public liquidity provision and the cost of financing government expenditures. Therefore, an optimal long- run debt-to-GDP ratio can exist. When the economy is hit by large adverse financial shocks, a proper mix of monetary and fiscal policies is the key to stabilize the economy. In the model, private agents face idiosyncratic investment opportunities and there is no insurance market for idiosyncratic risks. When an investment opportunity arrives, an agent becomes an entrepreneur. Otherwise, she is a worker earning wages. The entrepreneur can finance investment by issuing financial claims to the new capital stock and/or by liquidating existing claims. Issuing or reselling private claims needs to go through competitive financial intermediaries, which implement a costly search and matching technology. The model׳s innovative aspect is the asset market structure. Asset liquidity is endogenously determined in a competitive search environment where agents meet financial intermediaries. Two important results emerge by endogenizing asset liquidity. First, the model generates the positive co-movement between asset saleability and asset prices, as observed in the data. The key aspect is that asset demand is endogenous and determines saleability, tightness of financing constraints, and asset price jointly. Second, government policies can affect the market structure and asset liquidity. As market- and bank-based financial intermediation both share the essential feature of matching savers and borrowers, our framework admits both interpretations of the intermediation process, as in De-Fiore and Uhlig (2011). The financial frictions, induced by costly search, reduce the amount of resources that are transferred to entrepreneurs with investment projects. Therefore, if the liquidity of private claims is too low, private agents cannot fully self-insure idiosyncratic risks. When fully liquid government bonds circulate, private agents will purchase them as a buffer stock in case a good investment opportunity comes up. Such precautionary savings reduce the real return from bonds and further raise liquidity premium. I analyze a class of monetary and fiscal policy rules with long-run targets suggested by actual policy practice. The monetary authority sets the nominal interest rate as a function of current inflation and output. The fiscal authority chooses a level of tax rate that depends on the quantity of real government debt; it can also adjust government expenditures as a function of current debt and output. I first study the long-run optimal supply of (real) government debt. The government faces a key tradeoff when supplying government bonds. On the one hand, more (real) government debt implies more public liquidity, which can be accumulated by agents awaiting investment opportunities and enhances entrepreneurs׳ ability to carry out investment when the opportunity arises. The financing constraints are less tight, thus reducing the liquidity premium of private claims and further encouraging buyers entering the financial market. On the other hand, more (real) government debt also implies a higher real interest rate and/or a higher tax rate, raising the cost of financing government expenditures. After analyzing the long-run economy, I turn to equilibrium dynamics after financial shocks. First, I study a simple class of policy rules in which the monetary authority sets the nominal interest rate as an increasing function of inflation only and the fiscal authority only adjusts the tax rate but not government expenditures. Adverse financial shocks, modeled as rises in the search costs, drive away the asset demand, reduce the saleability of private claims, and push down their price. The financial shocks affect the bid-ask spread, as in Ajello (2012) and Bassetto et al. (2015). Adverse financial shocks drive up the hedging value of liquid assets and push workers towards them. The liquidity premium rises and the demand for private claims falls: a flight to liquidity occurs. Entrepreneurs who issue private claims are even more financing constrained. Therefore, demand for goods drops and firms produce less. Such output contraction lowers agents׳ income expectation further, pushing them even more towards liquid assets. We thus see persistent falls in investment, consumption, and output. Because of less demand for physical investment after adverse financial shocks, inflation falls and nominal interest rates drop. When the adverse financial shocks are large, the nominal interest rate could drop to the zero lower bound (ZLB). With nominal frictions, I show that large adverse financial shocks combined with the ZLB can generate a much deeper recession, compared to the case when the ZLB is not a hard constraint. Finally, I consider a more sophisticated class of policy rules. I allow the nominal interest rate to be a function of both inflation and output, and allow the fiscal authority to adjust government expenditures. The optimal policy responses to financial shocks, within this class of monetary and fiscal rules, suggest an anticipated expansionary fiscal contraction under which the ZLB does not constrain nominal interest rates. In particular, an anticipated reduction of government expenditures “crowds in” resources, relaxes entrepreneurs׳ financing constraints, and raises aggregate demand immediately. Compared to the simple rule, such expansionary fiscal contraction (in response to financial shocks) implies that the fiscal authority should be more accommodating while the monetary authority needs to be more responsive to inflation. Relationships to literature: This paper analyzes financial frictions based on Kiyotaki and Moore (2012) and Shi (2015).4 The large family structure for simple aggregation in this paper also follows Shi (2015). The presence of liquidity constraints opens up the possibilities for government bonds or fiat money to circulate, which at least goes back to Holmström and Tirole (1998). That is, if private liquidity is not enough, public liquidity can be added to achieve efficiency.5 This paper provides a novel channel in which public liquidity provision is costly due to distortionary taxation. Therefore, an optimal supply of public liquidity emerges. Importantly, asset saleability and asset price are endogenously generated from costly search and matching, thus avoiding the (counterfactual) negative co-movement discussed in Shi (2015) and Bigio (2012) when asset saleability is exogenous. A drop in asset demand reduces saleability of assets and their prices, thus tightening financing constraints. This effect is similar to that from the random search framework of Cui and Radde (2015), who focus on the fragility of financial markets. For an survey of the literature using search in asset markets, see Rocheteau and Weill (2011) or Lagos et al. (2016). Since the last financial crisis, nominal interest rates in many developed economies have stayed at the ZLB. I show that why with liquidity frictions adverse financial shocks can generate deflationary pressures and push nominal interest rates to the ZLB (when the monetary authority follows a simple Taylor rule). In this regard, this paper is a compliment to Eggertsson and Krugman (2012) and Buera and Nicolini (2014). Two sets of policy tools have been proposed to deal with the ZLB. The first is unconventional monetary policy or “Quantitative Easing” (QE).6 Del Negro et al. (2011) demonstrate the large impact of liquidity fluctuation on investment and asset prices and how unconventional monetary policies might stabilize the economy. The second proposal, such as in Christiano et al. (2011) and Woodford (2011), argues that raising government expenditures will not crowd out private consumption when the ZLB binds. I show that even without the consideration of “QE”, the monetary–fiscal interactions are important when the economy features liquidity frictions. Further, after adverse financial shocks, fiscal expansion might not be better than fiscal contraction which crowds in resources, relaxes financing constraints of some private agents, and raises aggregate demand. Government debt provides liquidity service and has the “crowding-in” effect, similar to Woodford (1990). This aspect is in contrast with Aiyagari and McGrattan (1998) in which government debt is a perfect substitute to private assets (or capital stock). Government debt relaxes agents׳ borrowing constraints but also has the “crowding-out” effect on capital accumulation. In Angeletos et al. (2013), government bonds also “crowd in” resources as they have a higher level of exogenous pledgeability as collateral than capital assets. This paper, however, features asset saleability (and a bid-ask spread) instead of asset pledgeability. In addition, I show that public liquidity provision can also alter the liquidity of privately issued claims through the endogenous search and matching.","I construct a growth model with liquidity frictions and nominal price stickiness. The economy is populated by a continuum of similar households (with measure one), firms, financial intermediaries, and a government. Time is discrete and infinite. Each period is divided into four sub-periods. (1) The households׳ decision period. The shocks to aggregate productivity At (TFP shocks) and to intermediation search cost κt (financial shocks) are realized. All members in a representative household equally divide the household׳s financial assets consisting of government bonds and privately issued financial claims on capital stock. The household instructs its members on the optimal type-specific choices to be carried out after individual types (explained below) have been realized. (2) The production period. Each member receives a status draw, becoming an entrepreneur (type i) with a probability χ and a worker (type n), otherwise. Workers supply labor hours, while only entrepreneurs have access to investment projects transforming consumption goods into capital stock one-for-one. Intermediate goods firms, in a monopolistically competitive market, rent capital stock Kt and labor Nt from the household to produce intermediate products. The after-tax rental rate and wage rate are rt and Wt. Final goods firms produce numeraire consumption goods by aggregating intermediate products. The profits (Dt) from intermediate goods firms are equally distributed among members. (3) The investment period. There is no insurance among household members, and they keep separated until the consumption stage. Because of the equal division of assets in the beginning, entrepreneurs are financing constrained in investment. They seek outside financing. Financial markets open in which entrepreneurs offer financial claims for sale and workers purchase these claims through financial intermediaries, which implement a costly search and matching technology. Search frictions imply that private financial claims are only partially liquid. In contrast, government bonds are fully liquid as they can be traded on a frictionless spot market. Financial intermediaries are competitive and search in financial submarkets. Entrepreneurs and workers choose the best financial market subject to the participation constraints of financial intermediaries. The competitive search process shares similar features with that in the labor search literature such as Moen (1997). (4) The consumption period. Entrepreneurs and workers reunite again in their respective households, pool all assets together, and share consumption goods across all members. Throughout the paper, I focus on the type of equilibrium in which entrepreneurs are financing constrained and both private claims and government bonds circulate. I verify the existence of such type of equilibrium in the numerical analysis. A representative household ~~~~~~~~~~~~~~~~~~~~~~~~~~ Note that there is no currency, which is another form of public liquidity in practice. One can give fiat money an extra role, for example, by having money in the utility of the household. Nevertheless, a higher level of money or government debt does not necessarily mean more liquidity. Public liquidity here means real liquidity which has to be eventually backed by the government׳s primary surplus. What is important is the real value of liquidity, whether it is in the form of government bonds or in the form of money. For this reason and for simplicity, I mix government bonds and money. Asset price and asset liquidity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ There is a continuum of competitive financial intermediaries, each of which chooses on which submarket to collect and match quotes at per-quote costs of κ units of consumption goods. The probability of filling a buy quote is fm, while the probability of filling a sell quote (or asset saleability) is ϕm. Firms ~~~~~ The production sector consists of intermediate goods firms and final goods firms, as in a standard New Keynesian model such as in Leeper et al. (2012) and references therein. To simplify, I only incorporate sticky prices using the Calvo type of price adjustment. Government policies and the transformed economy ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The government follows policy rules which are taken as given by private agents. I restrict to a class of policy rules that have long-run targets and induce a balanced growth path. For this reason, the “long-run equilibrium”, the “steady state”, and the “balanced-growth path” are used interchangeably. Recursive equilibrium ~~~~~~~~~~~~~~~~~~~~~ Now, I close the model by defining the recursive equilibrium. Readers who are not interested in the technical definition can skip this subsection.","In this section, I characterize the equilibrium analytically. The government can provide public liquidity (real government debt) to mitigate the liquidity frictions. When private liquidity is in shortage due to the search frictions, public liquidity can be used as an alternative. A higher level of public liquidity means that entrepreneurs with government bonds become richer and can invest more, according to the investment equation (11). It can also reduces the liquidity premium through the endogenous search market structure, as entrepreneurs rely less on the outside financing. Then, I show that public liquidity provision is, however, not free because of the distortionary taxation. When the government sets long run targets, it faces the tradeoff between the benefit of public liquidity provision and the cost of government financing. Liquidity premium and asset price ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Now, I study the relationship between asset saleability and asset price. One shall see that whether a lower level of the equilibrium saleability ϕ is associated with a higher or a lower asset price qi depends on the relative strength of asset supply and demand effects. On one hand, a lower level of ϕ implies tighter financing constraints and less supply relative to demand on the asset search market as shown by Shi (2015). Then, the shadow value of private claims rises. As intermediaries have to offer better conditions to attract scarce supply, this should be reflected in a higher equilibrium asset price. On the other hand, lower equilibrium asset saleability implies that private claims are less effective investments to hedge future funding needs, which would reduce demand, increase the equilibrium liquidity premium, and compress the asset price. Costly public liquidity provision ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Since private claims do not have enough liquidity, the government can provide public liquidity to mitigate liquidity frictions. But public liquidity provision is costly, as there is a tradeoff between the benefit of public liquidity provision and the cost of government financing. Intuitively, since private claims carry a liquidity premium and government bonds are scarce, precautionary entrepreneurs tend to push down the real interest rate compared to the case without liquidity frictions. Consequently, the (real) price of bonds are unnecessarily high, and the government needs to supply more real government debt. Given a qn, this public liquidity provision also reduces the liquidity premium in (34), as entrepreneurs rely less on the financial market but more on public liquidity. In later numerical analysis, I will show that these effects still hold in general equilibrium. The above discussion suggests that liquidity premium has a negative impact on financing investment. On the contrary, if government expenditures need to be financed at the same time, some degree of liquidity premium is preferred by the fiscal authority. One possible way to see this is that, a higher real interest rate implies that government expenditures become more expensive to finance. With a higher real interest rate, the government with a high level of real government debt will have to either cut government expenditures (which reduces the social welfare that depends on public spending), or raise more distortionary taxation. Calibration ~~~~~~~~~~~ The model is calibrated in quarterly frequency to match several US long-run statistics. The calculation of steady state can be found in Appendix A. The optimal debt-to-GDP ratio ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ With more public liquidity, workers search less private liquidity such that the equilibrium asset saleability falls. But at the same time entrepreneurs are less financing constrained, and the representative household is thus more willing to consume. As consumption rises with the debt-to-GDP ratio, leisure increases (hours worked fall) because they are complements. One interesting feature is that output is non-monotone. The drop of labor hours reduce output initially, but output rises with the debt-to-GDP ratio because the cut in expenditures “crowds in” investment and production. Importantly, although the public liquidity provision mitigates financial frictions, the cut in government expenditures will eventually reduce the social welfare with the rise of debt- to-GDP ratio (for a given ψg). The consumption-equivalence welfare gains are about 1.1% if the debt-to-GDP ratio is 0.6 (the calibrated level) and about 0.6% if the debt-to-GDP ratio is 1.0. Therefore, the parameter ψg is chosen such that welfare is maximized at the calibrated debt-to-GDP ratio. In sum, facing liquidity frictions, the monetary authority prefers a higher real interest rate and public liquidity provision, as it will attenuate liquidity frictions. Nevertheless, financing government expenditures becomes more costly with a higher real interest rate. The fiscal authority thus prefers a lower interest rate. The conflicts then imply both an optimal real interest rate and an optimal level of government debt that balance the tradeoff.","First, I study a special and simple class of rules in (24), in which the monetary authority sets the nominal interest rate only as a function of inflation, while the fiscal authority adjusts taxes in response to the level of real government debt outstanding. The simple rules are fairly standard in the literature. Then, I move to discuss the scenario when the government is able to follow the full set of rules in (24). A simple class of monetary and fiscal rules ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ This division is similar to Leeper (1991), who points out that both monetary policy and fiscal policy can be labeled as either active or passive. Being active means that the authority is forward looking and is not constrained by the current government budget constraint; it is free to choose a rule. Being passive means that it is backward looking and is constrained by the active authority and the private agents׳ behaviors. An economy system has a locally unique stable equilibrium if and only if one authority is active and the other authority is passive. Therefore, region I is the active monetary–active fiscal (AM + AF) regime, while region II is the passive monetary–passive fiscal (PM + PF) regime. Region III combines both active monetary–passive fiscal (AM + PF) and passive monetary–active fiscal (PM + AF) regimes. The reason for such non-linear region boundaries is that liquidity frictions together with sticky prices generate non-trivial inflation and liquidity premium dynamics after shocks. The monetary–fiscal interaction is more complicated than a model without liquidity frictions. Next, I discuss the details. Simple policy rules and shocks ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ On the contrary, when there are only negative TFP shocks, the return from investment is persistently low and the need for investment drops. Therefore, negative TFP shocks reduce liquid assets׳ hedging value for future investment (given that there are no financial shocks). The difference can be clearly seen in the liquidity premium dynamics after the two shocks. Adverse financial shocks push up liquidity premium, while negative TFP shocks do the opposite. After the financial shocks, flight to liquidity pushes up liquidity premium which dominates the increase of marginal product of capital such that asset price falls. This channel avoids the rise of asset price generated by an exogenous fall of ϕ as noted by Bigio (2012) and Shi (2015). TFP shocks, however, generate smaller declines in asset saleability in the initial periods. This is because a persistent lower TFP, albeit pushing down liquidity premium, reduces rental rates of capital and also generates persistent falls in the demand for private claims and asset price. Note that asset price and asset liquidity affect investment, since entrepreneurs use the financial market to leverage their net worth for investment projects. The drop of asset price qi and saleability ϕ tightens entrepreneurs׳ financing constraints. That is why financial shocks have a significant impact on investment (with a 5% initial drop) compared to TFP shocks (with a 3% initial drop). Note that the consumption movement after financial shocks is affected by nominal rigidities. In the previous literature without sticky prices, investment and consumption usually move in opposite directions after negative financial shocks (Kiyotaki and Moore, 2012 and Shi, 2015). The drop of output is small and only limited to the forgone capital accumulation. With nominal rigidities, “flight to liquidity” reduces inflation and increases real wages. Intermediate goods firms thus find it optimal to reduce labor hours. Then, output falls immediately together with a sharp decline of investment. That is why consumption does not need to rise and move in the opposite direction of investment (except for the first two periods as we have a moderate degree of price stickiness). After financial shocks, workers invest more in liquid assets which are not backed by real investment. Therefore, capital accumulation significantly slows down and the total net worth of the economy falls, which further exacerbates the financing situation of entrepreneurs in the future. Anticipating this, workers further fly to liquidity which causes persistent reduction of investment, consumption, and output. The impact of zero lower bound after financial shocks ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The previous exercise suggests that the nominal interest rate could drop below zero, when adverse financial shocks are large enough to generate a high deflationary pressure. Nevertheless, the zero lower bound (ZLB) seems to be a crucial technical constraint, as agents will prefer to hold money if the nominal rate is below zero. In my model, the ZLB will be important as product prices are sticky. If constrained by the ZLB, which is a hard constraint,9 the nominal interest rate can further affect policies and the real economy activities. Importantly, the ZLB amplifies the effect of financial shocks: the constraint is binding in a given period, and agents expect it to be binding in the future. This belief lowers expected future income and net worth and generates deflationary expectations. Such expectation leads to a rise in real rates and a fall in demand. Therefore, workers prefer to save in liquid government bonds and reduce consumption further. This again reduces the demand for private claims and their saleability, and entrepreneurs find it even harder to finance investment projects. As a result, asset price drops significantly more if the ZLB binds. Note that the ZLB brings up the real interest rate when inflation falls, and the fiscal authority needs to raise more taxes and issue more government debt to finance the interest payments. This effect further distorts the economy. Liquidity premium increases less when the ZLB binds, reflecting that the fiscal authority finds it more costly to finance government expenditures. The above exercise complements the studies by Eggertsson and Krugman (2012) and Buera and Nicolini (2014) in which they argue that the hit of ZLB is because of tightening of borrowing constraints. Here, adverse financial shocks are similar to the tightening of borrowing constraints as issuing and/or reselling assets are harder and more costly. Agents fly to liquidity, reduce aggregate demand, and push down inflation. Then, the monetary authority significantly reduces the nominal interest rate, which could stay at zero for a long period of time. Optimal simple and sophisticated policy rules ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Since the unconditional expectation of average household utility can be expressed as the weighted unconditional variances of output, inflation, consumption, investment, and government expenditures, I search for the coefficients which minimize the unconditional variances. As a comparison, I also compute the optimal policy parameters when the government is constrained to the simple rules. The difference can be seen from the optimal policy parameters in Table 3. One might view that fiscal expansion is always helpful in stabilizing inflation, since it can raise aggregate demand when inflation is low. This point is argued by Christiano et al. (2011) and Woodford (2011) when nominal interest rates are at the ZLB. Nevertheless, anticipated fiscal contraction can also raise aggregate demand when entrepreneurs are financing constrained. Agents expect that future fiscal contraction gives back resources to entrepreneurs, thus relaxing financing constraints and raising the demand for investment in the future. Workers are more willing to search for investment projects, relaxing entrepreneurs׳ financing constraints today. Therefore, entrepreneurs can invest more today, and households accumulate more capital and find themselves richer in the future, which further encourages workers to search for investment projects today. That is, in the simulation, the fiscal contraction׳s effect dominates the fiscal expansion׳s effect in raising demand, as endogenous financing constraints are powerful in amplification. The important consequence of this transmission is that inflation is stabilized even if the economy is hit by large financial shocks. Under the sophisticated policy rules, however, inflation is stable and real interest rates drop after shocks. That is why liquidity premium increases more than that under the simple policy rules while the real debt burden only increases slightly. The tradeoff between the benefit of public liquidity provision and the cost of government financing is less severe when agents fly to liquidity and government expenditures fall. The increase of the real value of debt is mostly from the voluntary drop of government expenditures, instead of from the drop of nominal price levels. By using fiscal contraction, the government maintains a healthy increase of public liquidity provision. As a result, the rise in distortionary taxation is moderate (Fig. 5). In sum, the monetary authority is more active than what the simple rules imply. However, it is unlikely constrained by the ZLB even if the economy is hit by very large financial shocks. Although the government is constrained by policy rules in the model, a careful mix of monetary and fiscal policies can avoid the deep recession generated by liquidity frictions and the ZLB. Remark: I do not analyze the optimal taxation in these policy experiments. First, fixing the tax rule is relatively simple. Perhaps, in practice, it is indeed hard to change taxation within a short period of time, while adjusting government expenditures is easier. Second, if the fiscal authority can adjust tax rules, one might view that it is better to provide real liquidity right after shocks by promising higher tax rates in the future. However, it is unlikely that such policy is better than the sophisticated rules, since it requires a higher level of distortionary taxation and is subject to more severe commitment issues.","I illustrate the importance of monetary–fiscal interactions with endogenous liquidity frictions. In particular, the tradeoff between the benefit of public liquidity provision and the cost of financing government expenditures implies an optimal long-run supply of government debt. In combating large adverse financial shocks, the proper mix of monetary and fiscal policies can avoid the zero lower bound on nominal interest rates under a simple Taylor rule. For simplicity, the paper assumes government debt as a blend of non- interest and interest bearing liabilities. Future research can give an explicit role for fiat currency and/or reserves of central banks. At the same time, I assume “private claims” as an amalgam of privately issued equity and debt. One could further analyze a more realistic capital structure and the implication for monetary–fiscal interactions. Finally, further work could also focus on the government commitment issues (for example illustrated in Bassetto, 2005) when government debt provides liquidity services."],["Expectations play a crucial role in modern macroeconomic models. We consider a New Keynesian framework under a behavioral model of expectation formation and under rational expectations. Contrary to the rational model, the behavioral model predicts that inflation volatility can be lowered if the central bank reacts to the output gap in addition to inflation. We test the opposing theoretical predictions in a learning-to-forecast experiment. In line with the behavioral model, the results support the claim that output stabilization can lead to less volatile inflation. --------------------------------------------------------------------------------","Expectations play a crucial role in modern macroeconomic theory. Standard models used for scientific research and policy analysis typically assume a representative fully rational agent. However, the assumption that all agents in an economy are fully rational and able to determine the model-consistent expectation of the underlying process governing real- world economic outcomes is highly problematic. A great deal of research has shown that humans generally do not react fully rationally to the world around them. This research ranges from providing evidence for simple biases to showing the inability of humans to work with probabilities and to forecast future economic behavior (Tversky and Kahneman, 1974, and Grether and Plott, 1979, are seminal early contributions, and many have followed since; see Camerer et al., 2011, for an overview). Moreover, the claim based on evolutionary arguments that behavior deviating from the homogeneous rational expectations solution will be driven out of markets over time has not held up to scrutiny (Brock and Hommes, 1997; Brock and Hommes, 1998; De Grauwe, 2012a; see also Arthur et al., 1997). In this paper we consider a standard macroeconomic model under both behavioral and rational expectations. We examine aggregate macroeconomic behavior and policy implications arising from the alternative assumptions on expectation formation, paying particular attention to price stability. The behavioral model of expectation formation is a heuristic switching model developed over a long period of time based on (mainly microeconomic) research investigating how people form expectations and how they adapt them over time. Models of this kind perform well in describing expectation dynamics using both survey and experimental data (e.g., Branch, 2004; Hommes, 2011; Assenza et al., 2018). A key difference in outcomes between the macroeconomic models with behavioral and rational expectations concerns price stability, that is inflation volatility. Assuming rational expectations, there is a clear trade-off for a central bank between fighting inflation volatility and output gap volatility. If the central bank reacts to the output gap in addition to inflation, under rational expectations this will result in an increase of inflation volatility. The outcome is different under behavioral expectations. Starting from a situation in which the central bank does not react to the output gap at all, the central bank can simultaneously decrease inflation volatility and output gap volatility by reacting to the output gap. However, inflation volatility as a function of the extent of output gap reaction is U-shaped. This means that reacting to the output gap on top of inflation will only lower inflation volatility up to a certain point, after which inflation volatility starts to increase again. These different outcomes regarding inflation volatility can be tested in the laboratory. We design a learning-to-forecast experiment where the only difference between treatments consists in the monetary policy rule used by the central bank. In one treatment, the central bank only reacts to inflation, while in the other it also reacts to the output gap. Our experimental results support the claim that inflation volatility can be lowered when the central bank also reacts to the output gap, in line with the predictions of the behavioral model. Our results from the behavioral model and the experimental data have clear policy implications for central banks whose sole aim is to achieve price stability, such as the European Central Bank (many other central banks, including those of New Zealand, Canada, England, and Sweden, have a hierarchical mandate with price stability as the primary objective for monetary policy). Even if these banks ultimately only care about price stability, this goal is better achieved if they also react to changes in the output gap. This is important and at odds with standard macroeconomic thinking built upon full rationality. Our work mainly relates to two streams of literature. Firstly, it relates to the literature on behavioral macroeconomics and learning in macroeconomics, for example Marcet and Nicolini (2003), Orphanides and Williams (2006), Branch and McGough (2009, 2010), Woodford (2010), De Grauwe (2011, 2012a, 2012b), De Grauwe and Kaltwasser (2012), Anufriev et al. (2013), Kurz et al. (2013), Benhabib et al. (2014), and Bertasiute et al. (2018); see Evans and Honkapohja (2001) and Woodford (2013) for overviews. In particular, our paper is not the first to derive a non-monotonic relationship between inflation and output gap volatility. De Grauwe (2011, 2012a) obtains a similar result in a different macroeconomic model with simple behavioral rules of expectation formation. Moreover, Kurz et al. (2013) show that the trade-off between inflation and output gap volatility is non-monotonic under rational diverse beliefs (see Kurz, 2009, for a survey on rational diverse beliefs). Branch et al. (2009) present similar findings in sticky information economies in which the degree of attentiveness or the rate at which agents update their information is endogenized. The main contribution of our paper is the empirical test of policy trade-offs by means of a learning-to-forecast laboratory experiment. To our knowledge, our results are the first experimental evidence for the missing trade-off between inflation and output gap volatility. Secondly, our research relates to the literature on experimental macroeconomics and learning-to-forecast experiments, for example Marimon and Sunder (1993), Kelley and Friedman (2002), Lei and Noussair (2002), Arifovic and Sargent (2003), Adam (2007), Heemeijer et al. (2009), Bao et al. (2012), Kryvtsov and Petersen (2013), Cornand and M’baye (2018), Pfajfar and Zakelj (2014), Assenza et al. (2018), and Hommes et al. (2019); see Duffy (2012), Assenza et al. (2014), and Cornand and Heinemann (2014) for reviews. Most closely related to our paper are the learning-to-forecast experiments of Pfajfar and Zakelj (2014) and Assenza et al. (2018), which are framed in a setup similar to ours. The first paper studies inflation expectation formation under different interest rate rules reacting only to inflation, finding that rationality of expectations can be rejected for the majority of participants. We focus instead on the policy trade-off between inflation and output gap volatility, comparing experimental outcomes when the central bank only targets inflation and when it also targets the output gap. Another important difference concerns the experimental design. While in Pfajfar and Zakelj (2014) participants forecast inflation only, we allow subjects to forecast both inflation and the output gap, in accordance with the theoretical macroeconomic model underlying the experiment. Assenza et al. (2018) consider participants forecasting both inflation and the output gap but focus on the Taylor principle as a device to pin down inflation dynamics. This paper is organized as follows. In Section 2 we describe how we model the economy and the formation of expectations. We also show the main differences between the rational and behavioral versions. In Section 3 we first describe the experimental design and the procedures. Then we show the experimental results. Section 4 concludes.","In this section, we first describe the underlying macroeconomic model. Then we introduce the behavioral model of expectation formation. After that, we compare the outcomes of both models and describe the economic intuition behind these outcomes. A behavioral model of expectation formation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Models with rational expectations are based on the assumption that agents have perfect information and a full understanding of the true model underlying the economy. There is, however, a large body of empirical literature documenting departures from this assumption. Furthermore, in survey data there is usually a high variance of inflation forecasts (e.g., Carroll, 2003; Mankiw et al., 2003; Branch, 2004) strongly suggesting that expectations are heterogeneous. Existing literature on expectation formation shows that many people use heuristics to make forecasts of future (macroeconomic) variables. This behavior is not necessarily a consequence of agents’ irrationality; it can also be a “rational” response of agents who face cognitive limitations and have an imperfect understanding of the true model underlying the economy (e.g. Gigerenzer and Todd, 1999; Gigerenzer and Selten, 2002). Next, we introduce a behavioral model of expectation formation for such an environment. Robustness and measurement of inflation volatility The simulation results are qualitatively robust to a wide variety of changes. This includes changes in all parameters of the macroeconomic model. It also includes changes in the parameters of the behavioral model of expectation formation. More interestingly, the results are also robust to other models of behavioral expectation formation, such as a heuristic switching model with fewer and simpler heuristics or adaptive expectations without any switching involved. Also if the central bank uses a slightly different Taylor rule to smooth the interest rate, the results persist (more precisely, we model the interest rate rule in that exercise as a weighted average of the regular Taylor rule and the previous period’s interest rate). Such variations are shown in Online Appendix B. While the results are qualitatively robust to these changes within this macroeconomic framework (which is the most standard framework for macroeconomic policy analysis), it is possible in other macroeconomic frameworks to reverse the results obtained by rational expectations. That is, it is possible to obtain a reduction of inflation volatility by increasing ϕy; an example are models that only include shocks to the aggregate demand equation (1) such as technology shocks, preference shocks or variations in government purchases, but do not include shocks to the short-run aggregate supply relationship (2) (see Woodford, 2003). In such frameworks, behavioral expectations are an additional reason why inflation volatility decreases when also targeting the output gap. One of the reasons why the mean squared deviation from the target may be popular among economists is that it constitutes a welfare criterion under homogeneous expectations. However, as shown in Di Bartolomeo et al. (2016), it is not an appropriate welfare criterion when agents have heterogeneous expectations. In this case, price dispersion arises not only because of the staggered price setting mechanism but also because of the heterogeneity of prices set by reoptimizing firms in the Calvo lottery, which depends on the heterogeneity of firms’ expectations of future inflation. In our behavioral model, this heterogeneity increases in the relative changes in inflation. In general, reducing inflation volatility by means of an interest rate rule that reacts to both inflation and output gap fluctuations, is thus welfare improving not only because it reduces the price differences between optimizing and non-optimizing firms, but also because it reduces the cross-sectional variance of prices set by optimizing firms. We refrain from using the precise welfare criterion as it depends on more than inflation alone. While welfare criteria derived in particular models may have influenced the fact that price stability is now the sole aim of many central banks, these central banks now have the mandate to achieve price stability and not the aim of maximizing a model-dependent welfare criterion. In addition, using a composite welfare measure would reduce the clarity and readability of the paper. Note that both simulation results and experimental results are similar when considering the precise welfare criterion.","The only task for subjects in the experiment is to forecast inflation and output gap. These forecasts are then used to calculate subsequent realizations. The model underlying the experimental economy is the macroeconomic model described in Section 2.1 (with the same calibration of macroeconomic parameters as before). Before we describe the experiment in more detail, we now explain the treatments and hypotheses. The design of the experiment and the hypotheses can be motivated with the theory described in Section 2. Treatments and hypotheses ~~~~~~~~~~~~~~~~~~~~~~~~~ We are interested in testing the null-hypothesis (which can be derived from the rational expectations model in Section 2) that inflation volatility in T1 is less or equal to inflation volatility in T2 against the alternative hypothesis (which can be derived from the behavioral model) that inflation volatility is greater in T1 than in T2. Fig. 3 summarizes these hypotheses. Course of events and implementation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The experiment was programmed in Java and conducted at the CREED laboratory at the University of Amsterdam. The experiment was conducted with 258 subjects recruited from the CREED subject pool (43 groups of six subjects each, distributed over thirteen sessions). After each session, participants filled out a short questionnaire. Participants were primarily undergraduate students, the average age was slightly above 22 years. About half of the participants were female, about two-thirds were majoring in economics or business, and about half were Dutch. During the experiment, ‘points’ were used as currency. These points were exchanged for euros at the end of each session at an exchange rate of 0.75 euros per 100 points. The experiment lasted around two hours, and participants earned on average about 30 euros. The series of error terms used in the model equations (gt and ut in Eqs. (1) and (2)) differed across groups within each treatment, but the sets of noise series used in the two treatments were the same.8 Results ~~~~~~~ There are data of 43 different groups, 21 in T1 and 22 in T2. The groups’ actions do not influence one another in any way; thus the observations at the group level are statistically independent. The data for all groups separately including all individual forecasts can be found in the Online Appendix E. Inflation Fig. 5 gives an overview of realized inflation in all experimental economies, separately for T1 and T2. Each line corresponds to the inflation in one economy, tracked over all 50 periods of the experiment. Almost all economies are close to the inflation target after 50 periods, and in the economies with inflation still oscillating around the target the amplitude of these oscillations is decreasing.9 Groups are heterogeneous, with some groups exhibiting much larger volatility than other groups. In many groups in both treatments, inflation is within one percentage point from the target in most or all periods. Inflation generally fluctuates around its target: mean inflation is between 3.13 and 4.33 in T1 and between 2.79 and 3.82 in T2 (treatment averages of mean inflation are 3.55 in T1 and 3.41 in T2). On average, inflation fluctuates a bit less in T2 than in T1, as predicted by the behavioral model.10 The heuristic switching model as predictor of subjects’ forecasts After having analyzed the economic outcomes, we now examine the performance of the heuristic switching model used to derive the predictions in the experiment. Does this model accurately describe subjects’ forecasts? Or does the rational expectation solution or one of the heuristics alone predict subjects’ forecasts better than the switching model? In addition, we compare the prediction performance to that of a few parsimonious heuristic switching models. We report the prediction performance of the heuristic switching model (HSM), the performance of the homogeneous rational agent solution (RE) and the performance of the four heuristics involved in the switching model without any switching: adaptive expectations (ADA), weak trend-following (WTR), strong trend-following (STR), and the learning, anchoring and adjustment rule (LAA). In addition, we compare the model to four parsimonious heuristic switching models, namely one where agents switch between using naive expectations and using a trend-following rule with coefficient one (Naive+Trend), one where agents switch between forecasting the steady state and a trend-following rule with coefficient one (Fundamental+Trend), and models where agents switch between the steady state plus a constant and the steady state minus a constant (Biased Fund.; we use three different constants, 0.25, 0.5, and 1). We use the mean squared difference between the models’ two- period ahead predictions and average forecasts in a group as prediction error. The predictions are thus out-of-sample predictions. In the main heuristic switching model, we use the same parameters as in Section 2 (that is, we use a preexisting calibration that is not influenced by our experimental data; for the parsimonious switching models, we use the same parameters unless otherwise specified). To form the predictions of the heuristic switching model, the fractions of agents using the different heuristics need to be determined. For this, we use the fractions that are implied by the theoretical model with a continuum of agents. We use the squared difference between this prediction and a group’s average forecast, because there are no degrees of freedom with this method.15 Table 2 shows average prediction errors across all periods and all groups in a treatment. Table 2 shows that, across the board, the benchmark model performs much better than rational expectations. Also evident in the table is that the rational expectation solution is a worse predictor in all cases than any of the four involved heuristics alone. Furthermore, the switching model is a better predictor in all cases than any of the four heuristics alone. In general, the differences are considerable. The switching model does much better than most of the other models. There are two heuristics that do very well when employed alone: the weak trend-following rule and the anchoring and adjustment rule. Nevertheless, the switching model predicts all four forecasts better than these heuristics. The prediction errors of these two best-performing heuristics when employed alone are always at least 25% greater than the prediction errors of the switching model. In addition, the main heuristic switching model does better throughout than the smaller switching models (that also have prediction errors of at least 25% above the benchmark model). Note that we are not attempting to fit the parameters of the heuristic switching model to the data ex post, as this would only give us lower prediction errors at the cost of adding degrees of freedom and thereby make the comparison with the homogeneous rational agent solution uneven. Fractions of heuristics used In the following, we consider the fractions of employed heuristics in the model fitted to the experimental data. This can help to understand which heuristics are used most and whether there are patterns concerning the use of the heuristics over time. Fig. 11 shows the fractions of the heuristics over time for inflation and output gap in T1 and T2 (the lines represent averages across groups). One can see that all heuristics have some support in the experiment. One can also see that the graphs of inflation and output gap forecasting in T1 are very similar. The same holds for the corresponding graphs in T2. This suggests that while subjects learn and update over time the way that they form expectations, they form expectations on inflation and output gap in similar ways, and change how they form expectations on inflation and output gap in similar ways. This is not self-evident: it could well have been the case that subjects rely more on trend extrapolation for one variable while behaving adaptively when forecasting the other variable. Regarding the use of the heuristics themselves, the adaptive rule and the anchoring and adjustment rule are used more often than the trend-following rules. The use of the adaptive rule increases over the course of the experiment, partially explaining the learning observed in the experiment. Furthermore, the trend-following rules are used less and less as the experiment proceeds (for inflation and output gap alike in both treatments; there are some small upward movements toward the end of the experiment in T1, however). This also contributes to the stability of inflation and output gap in the second half of the experiment, as the trend- following rules are destabilizing. The use of the anchoring and adjustment rule follows a less-clear pattern. It increases strongly in the beginning in T1 and decreases again thereafter. In T2 the use of this rule increases more slowly; afterwards, it levels off. This rule, which has two components, has a less clear- cut interpretation than the other rules. One component is destabilizing, taking into account last trends rather than predicting a return to the anchor immediately, while the other component, the anchor itself, is stabilizing, as the long-run averages are very close to the inflation target and the steady state of the output gap. It is interesting to see that, overall, relatively little use is made of the weak trend-following rule. While this rule alone predicts group level aggregates of forecasts rather well (Table 2), when looking at it from the point of view of the heuristic switching model, it seems that this is only the case because it approximates the prevailing mixes of the whole set of heuristics.","We have conducted a learning-to-forecast experiment to test the predictions of a macroeconomic model with behavioral expectations. This behavioral model yields results that differ from those of the same macroeconomic model based on rational expectations. Namely, the behavioral model yields that inflation volatility can be reduced if the central bank reacts to the output gap on top of inflation. The predictions of the behavioral model are supported by the outcomes of our experiment, in which the only treatment variation consists in a modification of the central bank’s monetary policy reaction function. These results are relevant for monetary policy analysis. They show a different relationship between inflation and output-gap than is usually assumed. The policy implications are particularly straightforward for central banks that aim at price stability alone, such as, for example, the ECB; these central banks should react to the output gap even if they are ultimately only interested in price stability. Note that our policy recommendation has a broad foundation. We obtain it not only with our behavioral benchmark model but also assuming a variety of other specifications of expectation formation (as outlined in Online Appendix B). Moreover, the same recommendation can arise from slightly different macroeconomic models with different behavioral expectations (Branch et al., 2009; De Grauwe, 2011; De Grauwe, 2012a; Kurz et al., 2013). In addition to these theoretical findings, we present empirical evidence from a laboratory experiment leading to the same policy recommendation. We are aware that many macroeconomists are still skeptical about the idea that one can learn about macroeconomic behavior by conducting laboratory experiments (with small group sizes). However, while one cannot mirror a completely macroeconomy with all its decisions in the laboratory, it is possible to shed light on some specific macroeconomic questions. Our experiment is designed such that it abstracts from everything except for the feedback mechanism from expectations to realizations (which is altered by a one-parameter change). All decisions made by agents in the economy are computerized and correspond fully to the underlying macroeconomic model except for the forecasting of future variables. If such forecasts in the laboratory are formed in similar ways as forecasts in the outside world, the results from the laboratory experiment help to understand macroeconomic behavior. There is indeed recent evidence that students’ forecasts in the laboratory have very similar characteristics to forecasts in the outside world (Cornand and Hubert, 2018) supporting the external validity of macroeconomic learning-to-forecast experiments."],["We develop a new model of multi-product firms which invest to improve the perceived quality of both their individual products and their brand. Because of flexible manufacturing, products closer to firms' core competence have lower costs, so firms produce more of them, and also have higher incentives to invest in their quality. These two effects have opposite implications for the profile of prices. Mexican data provide robust confirmation of the model's key prediction: firms in differentiated-good sectors exhibit quality-based competence (prices fall with distance from core competence), but export sales of firms in non-differentiated-good sectors exhibit the opposite pattern. --------------------------------------------------------------------------------","What makes a successful exporting firm? This question has attracted much interest from policy makers, keen to design effective export promotion programs, and from academics, keen to understand the implications of globalization for economic growth. Two answers have been proposed. The first focuses on firm productivity. Studies by Clerides et al. (1998) and Bernard and Jensen (1999), among others, have found that firms self-select into export markets on the basis of their successful performance at home. This evidence inspired the theoretical work by Melitz (2003) where only the most productive firms find it worthwhile to cover the extra costs of exporting. The second answer focuses on product quality. A growing body of work has provided evidence that successful exporters charge higher prices on average, suggesting that quality matters.1 This study integrates these two views and shows both theoretically and empirically that firms may choose to compete on the basis of either cost or quality depending on the characteristics of the products they sell and the markets in which they operate.2 Unlike other studies which have compared the behavior of different firms, and emphasized the between-firm extensive margin, we focus on the portfolio of products sold by multi-product firms, and highlight what Eckel and Neary (2010) call the “intra-firm extensive margin”. Our theoretical innovation is to construct a model of multi-product firms in which the quality of goods is determined endogenously by the firms' profit-maximizing decisions. Because of flexible manufacturing, products closer to a firm's core competence have lower costs. As a result, firms produce more of those products, but they also have higher margins on them, and therefore higher incentives to invest in their quality. These two effects have opposite implications for the profile of prices and, depending on which effect dominates, the model implies one of two possible configurations which we call “cost-based” and “quality-based” competence, respectively. The former corresponds to the case where a firm's core products are sold at lower prices, in order to induce consumers to buy more of them. In the words of Jack Cohen, founder of the UK supermarket chain Tesco, firms “pile 'em high and sell 'em cheap”. As a result, the profile of prices across a firm's products is inversely correlated with its profile of sales. By contrast, quality-based competence corresponds to the case where the dominant effect comes from firms' investing more in enhancing the quality of their core products. As a result, these products command higher prices, and so the profile of prices across a firm's products is positively correlated with its profile of sales. Our model not only allows for different profiles of prices but also makes predictions about which kinds of goods should exhibit which profile. In particular, it predicts that a higher level of product differentiation encourages firms to invest relatively more in the quality of individual varieties than in the quality of their overall brand. As a result, quality- based competence should be more in evidence in sectors where products are more differentiated. We test this prediction using a rich Mexican data set already used by Iacovone and Javorcik (2007, 2010). Most previous empirical studies of multi-product firms at plant level have been constrained to use data on export sales only, or to combine export and production data at different levels of disaggregation.3 By contrast, a unique characteristic of our data is that it provides consistently disaggregated information on both the home and export sales of all goods produced by a large representative sample of manufacturing establishments.4 As we show, the Mexican data provide robust confirmation of the model's key prediction: comparing price profiles with sales profiles, we find that firms in differentiated-good sectors exhibit quality-based competence to a much greater extent than firms in non-differentiated-good sectors, both at home and abroad. The contrast is particularly striking in export markets, where Mexican producers in non- differentiated-good sectors engage in cost-rather than quality-based competence. Our results are robust to focusing attention on a variety of subsamples, including only those products sold both at home and abroad, only those plants which sell on the home market and also select into exporting, and only single-plant firms. Our paper builds on and extends the existing literature on multi-product firms in international trade. While there already existed a large literature on multi-product firms in the theory of industrial organization, our model is one of a number of recent trade models which is more applicable to the kinds of large-scale firm-level data sets which are increasingly becoming available.5 Within this latter tradition, existing models impose one or other profile of a firm's prices by assumption. One class of models assumes that products are symmetric on both the demand and supply sides, with the motivation for producing a range of products coming from economies of scope. As a result, all products sell in the same amount and at the same price6. A different approach, pioneered by Bernard et al. (2010, 2011), emphasizes asymmetries between products on the demand side due to exogenous stochastic factors. Before they decide to enter, firms draw their overall level of productivity and also a set of product-market-specific demand shocks. The latter determine the firm's scale and scope of sales in different markets, and imply that its price and output profiles are always positively correlated. By contrast, Eckel and Neary (2010) develop a model that emphasizes asymmetries between products on the cost side and implies that price and output profiles are always negatively correlated.7 The present paper integrates these demand and cost approaches in an endogenous way. We extend the “flexible manufacturing” approach of Eckel and Neary (2010) by allowing costs to affect the profile of investment in quality across different varieties, and develop a model which is more in line with recent work on models of heterogeneous firms that engage in process R&D: see, for example, Bustos (2011) and Lileeva and Trefler (2010) on single-product firms, and Dhingra (2013) on multi- product firms. It is even more closely related to those papers which allow for endogenous investment in quality, such as Antoniades (2009) and Kugler and Verhoogen (2012), including the view that quality is really perceived quality, which may be market-specific, so investment in quality includes spending on marketing as in Arkolakis (2010). All this work has so far focused on single-product firms only. Our specification is we believe the first to incorporate investment in quality into a model of multi-product firms, combining insights from extensive literatures in both industrial organization and marketing science. From the former, especially Stigler and Becker (1977), we take the view that firms invest in perceived quality through advertising, which enters the utility function directly in a way that is complementary to consumption itself. From the latter, notably Jacoby et al. (1971), Boush et al. (1987), and Aaker and Keller (1990), we take the view that consumers of multi-product firms are affected both by product-specific marketing and by advertising of a firm's overall brand, and that the relative effectiveness of the former is greater when products are more differentiated. This brief review of the literature on multi- product firms highlights our main interest: how the theoretical models differ in the way they model the demand for and the decision to supply multiple products. The models also differ in other ways which are of less interest in the present application. One type of difference is in the assumptions made about market structure. In particular, most recent models assume that markets can be characterized by monopolistic competition, in which firms produce a large number of products but are themselves infinitesimal relative to the size of the overall market.8 By contrast, Eckel and Neary (2010) assume in their core model that markets are oligopolistic. In this paper, we know little about the market environment facing individual firms: we do not know with which other Mexican plants in the sample they compete directly, and we have no information at all on their foreign competitors. Hence we prefer to remain agnostic on this issue, where possible deriving predictions which will hold at the level of individual firms irrespective of the market structure in which they operate. A further dimension of difference concerns the level of analysis, whether partial or general equilibrium. Some of the trade theory papers, including Eckel and Neary (2010), highlight general-equilibrium adjustments working through factor markets as an important channel of transmission of external shocks. However, with the data set we use, it is not possible to ascertain how factor prices are affected by general-equilibrium adjustments to changes in trade policy. Hence, we concentrate on testing implications of the model in partial equilibrium. Section 2 of the paper presents the model and shows how differences in technology, tastes and market characteristics determine whether a multi-product firm exhibits cost-based or quality- based competence. Section 3 describes the data and explores the extent to which they confirm our theoretical predictions. Finally, Section 4 summarizes our results and presents some concluding remarks. The Appendix supplements and extends the theoretical results of Section 2, and in particular shows that they extend to a Cournot oligopolistic market with heterogeneous firms.","As already explained, the paper extends the flexible-manufacturing model of Eckel and Neary (2010) to allow for the interaction of quality and cost differences between the varieties produced by a multi-product firm. To simplify ideas and notation, we focus in the text on a single monopoly firm, but, as we show in Appendix B, all the results extend to a heterogeneous-firm industry in which firms engage in Cournot competition. Section 2.1 introduces our specification of preferences, while Section 2.2 briefly reviews the earlier model, which allowed for cost-based competence only, showing how a firm chooses its product range, its total sales, and their distribution across varieties in a single market. Section 2.3 explores the additional complications which quality-based competence introduce and derives our main theoretical result, and Section 2.4 considers the model's comparative static properties. Preferences for quantity and quality ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As discussed in the introduction, we remain agnostic in the paper about whether this sub- utility function is embedded in a general- or partial-equilibrium model: our analysis is compatible with both approaches. All we need is to assume that the marginal utility of income can be set equal to one. This is ensured if the sub-utility function (1) is part of a quasi-linear upper-tier utility function, with all income effects concentrated on the “numéraire” good. Alternatively, as in Eckel and Neary (2010), (1) can be one of a mass of sub-utility functions without an outside good, with the marginal utility of income set equal to unity by choice of numéraire. In either case, Eq. (1) is only one of the many sub-utility functions, corresponding to separable preferences for different groups of products. In our empirical work we will allow for the possibility that key parameters, and especially the product differentiation parameter e, may vary across markets and countries. To economize on notation we do not make this explicit in Eq. (1). Cost-based competence ~~~~~~~~~~~~~~~~~~~~~ Consider next the technology and behavior of the firm in a single market, which is segmented from the other markets in which the firm operates. The firm's objective is to maximize profits by choosing both the scale and scope of production, as well as choosing how much to invest in enhancing the quality of individual varieties and of its overall brand. We begin by abstracting from the quality dimension in this sub-section, and recapping the results of Eckel and Neary (2010) for the case where the firm's competence derives from differences between varieties in production costs only. This is most easily done by setting β equal to zero in Eq. (1), so utility does not depend on quality. Though it is convenient to make explicit the variety-specific intercepts a(i) in all equations, we do not consider the implications of differences between them until the next sub- section. Quality-based competence ~~~~~~~~~~~~~~~~~~~~~~~~ Our modeling of preferences and investment in perceived quality draws on previous work in both industrial organization and marketing science. Our assumption that both perceived quality and the physical quantities of goods consumed enter the utility function is consistent with the “complementary view” of advertising in the industrial organization literature. This can be traced back to Stigler and Becker (1977, p. 84), who argue that “utility depends not only on the quantity of the good but also the consumer's knowledge of its true or alleged properties” and that “the knowledge, whether real or fancied, is produced by the advertising of producers”. Bagwell (2007, p. 1720) adds that a “consumer may value ‘social prestige’, and advertising by a firm may be an input that contributes towards the prestige that is enjoyed when the firm's product is consumed.” By distinguishing between product-specific investments in perceived quality and investments in the quality of a firm's (umbrella) brand, we extend this framework to a multi-product setting. This result has been derived the case of a single monopoly firm, but it is independent of the extent of competition which the firm faces. We show formally in the Appendix that it continues to hold in a heterogeneous-firms Cournot oligopoly market, but the intuition is straightforward. With all goods symmetrically differentiated, firms compete against each other only at the level of total output, not at the level of individual varieties. Changes in the extent of inter-firm competition affect the scale and scope of production as well as the level of investment in quality, but do not influence the within-firm profile of prices across products, which is the focus of our empirical analysis. Comparative statics ~~~~~~~~~~~~~~~~~~~ The predictions of the model for the shape of the firm's equilibrium price profile given in Proposition 1 are the ones that we take to the data in the next section. It is also of interest to explore the comparative statics properties of the model. Here we note the effects of exogenous shocks on the scale and scope of a single monopoly firm, while in the Appendix we show that our results generalize to the case of a group of firms engaged in Cournot competition.","Our theoretical model makes a number of novel predictions about the behavior of multi- product firms. One of these in particular is unique to our model, has both theoretical and policy interest, and lends itself to empirical testing with our data. This is the prediction from Corollary 1 that the profile of prices across the different goods produced by a multi-product firm is more likely to be positively correlated with the corresponding profile of outputs, thus exhibiting what we have called quality-based competence, when products are more differentiated. In the remainder of the paper we subject this prediction to empirical testing. We first describe the data and document the profiles of sales across firms' products which it exhibits; then we explain how we operationalize the prediction about price profiles; subsequent sub-sections present the results of testing it and consider various robustness checks. The data ~~~~~~~~ We begin by reviewing the data set.21 A unique characteristic of our data is the availability of plant-product level information on the value and the quantity of sales for both domestic and export markets. Our data source is the Encuesta Industrial Mensual (EIM) administered by the Instituto Nacional de Estadstica Geografa e Informática (INEGI) in Mexico. The EIM is a monthly survey conducted to monitor short-term trends and dynamics in the manufacturing sector. As we are not primarily interested in short-term fluctuations, we aggregate the monthly EIM data into annual observations. The survey covers about 85% of Mexican industrial output, with the exception of “maquiladoras”.22 It includes information on 3183 unique products produced by over 6000 plants.23 Plants are asked to report both values and quantities of total production, total sales, and export sales for each product produced, making the data set particularly valuable for our purposes. Note that the unit of observation is the plant rather than the firm: we return to this issue in our robustness checks below. Products in the survey are grouped into 205 clases, or activity classes, corresponding to the 6-digit level CMAP (Mexican System of Classification for Productive Activities) classification. Each clase contains a list of possible products, which was developed in 1993 and remained unchanged during the entire period under observation. The classification of products is similar in level of detail to the 6-digit international Harmonized System classification, though with differences that reflect special features of the structure of Mexican industrial production.24 Table 2 shows that the number of plants in the sample varies from 6291 in 1994 to 4424 in 2004. Between 1579 and 2137 plants were engaged in exporting.25 The decline in the number of establishments during the period under analysis is due to exit.26 In our empirical analysis, consistent with our theoretical model, we refer to each plant-product combination as a “variety”. The number of varieties sold ranges from 19,154 in 1994 to 12,887 in 2004, while the number of varieties exported rose from 2844 in 1994 to 3118 in 2004, reaching a peak of 4193 in 1998. Sales profiles ~~~~~~~~~~~~~~ As a first step in exploring the properties of the data through the lens of our theoretical model, we considered the patterns of sales across the varieties produced by different plants in our sample. (Details are given in a background paper: Eckel et al. (2009)). The results were consistent with the model presented in Section 2, and also broadly in line with empirical patterns found in other recent studies of multi-product firms.27 In particular, the data show that exporting plants are larger, and that larger plants produce more products. The vast majority of plants sell more products at home, and most exported products are also sold at home. Finally, the profile of sales across products is highly non-uniform, with a broadly similar ranking of products by sales in the home and foreign markets. Empirical strategy ~~~~~~~~~~~~~~~~~~ We wish to use estimates of Eq. (21) to test the prediction of Proposition 1 that a higher degree of product differentiation should make firms more likely to exhibit a price profile that reflects quality-based rather than cost-based competence. To implement this test, we need independent observations on the degree of product differentiation, and for this purpose we make use of the classification developed by Rauch (1999). He grouped goods by the Standard International Trade Classification (SITC), Revision 2, four-digit classification into three categories, “differentiated,” “traded on organized exchanges,” or “reference priced.” We combine the latter two into a catch-all “non-differentiated” category, and follow many authors in adopting the so-called “liberal” classification, which maximizes the number of goods classified as non-differentiated.31 To implement this classification with our Mexican data, we had to make a concordance between the clases in our data and the SITC system. Fortunately, this was possible without too much arbitrariness.32 We are thus able to explore how the relationship between the price and sales profiles of multi-product firms varies with the degree of product differentiation. Results for price profiles at home and away ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 3 gives the results of estimating Eq. (21) over different subsets of the data on all plant/product/year observations for which the plant in question sells at least five products. Each column gives the results of regressing the corresponding price on product- year and plant-year fixed effects and on dummy variables for the highest to fourth-highest selling products. Thus, in the first equation, the coefficient 0.091 gives the estimated price premium on the top-selling product in the home market relative to the average price on the excluded category of all products ranked fifth or lower in home sales. This coefficient is highly significant, indicating that, on average, the highest-selling product from each plant commands a price premium of 9.1% (i.e., exp (0.091)-1). This provides strong evidence of quality-based competence, in the sense in which we have used the term in our theoretical model. The other coefficients in this equation are also highly significant, and fall steadily in size, again confirming a pattern of quality-based competence. The fourth equation in the table shows that export sales exhibit a similar pattern on average, with the coefficient on the dummy variable for the top-selling product equal to 0.162 and highly significant, the coefficients on the next two not significantly different from that of the top-selling product, and all three significantly greater than the last coefficient, for the top fourth-selling product. The most interesting feature of the table is the pattern of the estimated coefficients when we disaggregate by type of product and by destination. Looking first at the second and third equations, both differentiated and non-differentiated products sold at home exhibit the same pattern of quality-based competition. However, each coefficient for differentiated products is significantly greater than the corresponding coefficient for non-differentiated ones, exactly as our theory predicts.33 This difference between the two categories of products is repeated but to an even more striking extent in the export market, as the fifth and sixth equations show. The top three coefficients for differentiated exports are all highly significant and significantly larger than in the home market, implying even higher price premia for the top-selling products in this category. By contrast, the coefficients on the top four non-differentiated export products are all negative, implying that these products exhibit cost-based rather than quality-based competence. So far we have only considered the subset of firms selling five or more products. Tables 4–7 extend the analysis to observations in which the same plant sold at least two, three or four products in the one year. In each equation the residual category is all products with ranks lower than the lowest-ranking dummy variable included, and the final equation is repeated for reference from Table 3. The advantage of focusing on plants that export fewer than five products is a considerable increase in degrees of freedom, and the results are qualitatively similar to those in Table 3. Tables 4 and 5 confirm that the pattern of quality-based competition which Table 3 showed for plants selling five or more products on the home market also applies to plants selling fewer products. For differentiated products, Table 4 shows that the implied price premium for the top product ranges from 4.1% when all plants producing two or more products are included, to 11.5% when only those producing five or more are included. In Table 5 the corresponding figures are 3.9% and 3.1%, showing once again that non-differentiated products exhibit significantly less quality-based competition than differentiated ones, as our theory predicts. Also of considerable interest is that the pattern of prices falling with a product's rank which was found in Table 3 continues to hold for the larger samples of plants selling fewer products. Not only are most coefficients of the dummy variables for second- and lower-ranking products in these equations significantly different from zero, but there is a clear and in many cases significant downward trend in the coefficients in each column. We can thus conclude that there is strong evidence that prices fall with a product's distance from a plant's core competence, so the price and production profiles are negatively correlated, implying that on average the firms in our sample compete on the basis of quality-based competence on the home market. Tables 6 and 7 show that export sales behave even more differently depending on the degree of product differentiation. Consider first Table 6, which shows that the price profile of exports in differentiated sectors is qualitatively similar to that at home. The evidence for a monotonically decreasing profile is less strong in the case of plants producing five or more products, although this may be due to the smaller number of observations in this sub-sample, and in any case the top three products command a significant price premium over products ranked fifth or lower. Moreover, the quantitative magnitude of the effects is much higher than in Table 4: the price premium for the top product ranges from 8.5% when all plants producing two or more products are included, to 22.8% when only those producing five or more are included, compared with 4.1% and 11.5% respectively in Table 4. By contrast, Table 7 tells a very different story for exports of non-differentiated products. Not a single coefficient in this table is significantly positive, most are negative, and the overall pattern is one of increasing coefficients as we move down each column. Unlike Tables 4, 5 and 6, this provides strong evidence against quality-based competence, and suggestive evidence in favour of cost-based competence for exports of non-differentiated products. Though not as overwhelmingly significant as the results for differentiated products, the results imply that the two groups of products behave very differently, and exactly in the way predicted by Proposition 1. For differentiated exports, prices fall with their distance from the plant's core competence, suggesting that Mexican exporters in these sectors compete on the basis of quality. By contrast, for non-differentiated exports, prices tend to rise with their distance from the plant's core competence, suggesting that competition in such sectors is on the basis of cost rather than quality, exactly as our theory suggests. Overall, these four tables confirm that the coefficient pattern summarized in Eq. (22) continues to hold when we consider plants that sell up to five or more products. Robustness checks ~~~~~~~~~~~~~~~~~ A possible concern with the results so far is that the sample sizes are very different in different tables, with more products produced for the home market than for exports. This is perfectly consistent with our model which predicts that higher costs of accessing a foreign market will reduce the range of products sold there. Nevertheless it might suggest a concern that the regularities we have found in our data reflect behavior very different from that predicted by our model; for example, that plants sell different products in the home and foreign markets, or that plants which select into exporting are very different from those that sell only on the home market. To address these concerns we reestimate our price profile equations first for those products that are sold on both markets, and next for the home sales of exporting plants. We also present results for various subsets of the data, including firms with only a single plant and plants that are either domestically- or foreign-owned. Table 8 addresses the issue of different sample sizes directly by reestimating the equations for only those observations on products that are both exported and sold at home, so the numbers of observations are the same at home and abroad. It can be seen that the conclusions drawn from the earlier tables survive this robustness check. There is clear evidence of quality-based competition in differentiated products in both home and foreign markets, and this contrasts with the absence of a pattern in the coefficients for non-differentiated products. Even in the latter case, the size and sign of the coefficients, though not their significance, are consistent with those from the larger samples in Table 3. We can conclude that the evidence from this smaller sample is less overwhelmingly in support of different behavior by non-differentiated product plants at home and away; but that the evidence for a difference between behavior by plants in differentiated and non-differentiated sectors remains strong in both domestic and export markets. Table 9 addresses the question of whether plants that select into exporting behave differently on the home market. It gives results for home sales by export plants in both differentiated and non-differentiated clases, and it is clear that the two behave very similarly to the corresponding samples of all plants selling on the home market, as in Tables 4 and 5. Once again, home sales of differentiated products exhibit quality-based competence, while those of non-differentiated products do not. Bearing in mind that the plants in Table 9 are identical to those whose exporting behavior is shown in Tables 6 and 7, our earlier conclusions are reinforced. Exporting plants in differentiated sectors exhibit quality-based competence in the home market, whereas those in non-differentiated sectors do not, so the very different behavior of exporters in non-differentiated sectors shown in Table 7 does not reflect any differential selection process of plants into exporting. A different robustness check addresses the concern that our theory was developed for multi-product firms, whereas our data consist of observations on multi- product plants. Treating plants as the unit of observation ignores the interdependence of decision-making within multi-plant firms. To deal with this problem empirically we would ideally like to have data on the ownership patterns of plants in all years. Unfortunately, we can only identify which plants were owned by the same firm in the penultimate year of our sample, 2003. We therefore adopt the following strategy. We retain in the sample only those plants which were single-plant firms in 2003, and consider their sales and price profiles in all years. This risks including some observations on plants which did not correspond to single-plant firms either in 2004 because of mergers and acquisitions, or in years prior to 2003 because of divestitures. However, the number of such cases is likely to be small, and this strategy seems preferable to losing many more degrees of freedom by focusing on single-plant firms in 2003 only.34 Table 10 gives the results of this robustness check, for single-plant firms selling at least five products. The evidence for quality-based competence remains very strong for both categories of home sales and for differentiated exports. For these three categories, most coefficients are highly significant, implying that products closer to the core sell for higher prices than the non-core products in the default category of each equation. As for exports of non- differentiated products, the evidence for cost-based competence is weaker than in earlier tables, though the hypothesis that these sales exhibit quality-based competence is strongly rejected. We can conclude that our earlier results are robust to excluding plants owned by multi-plant firms in 2003. A final concern we address is whether our specification is more applicable to Mexican-owned plants or to foreign-owned ones. On the one hand, we would expect the decisions of foreign-owned Mexican plants to be taken as part of the global operations of their parent multinational companies rather than on a stand-alone basis, at least for sales in their export markets, and perhaps in their home market too. This would suggest that the considerations we have highlighted in our theoretical model should apply more to Mexican-owned plants. On the other hand, we would expect multinational companies to have stronger world brands and so to exhibit more quality-based competence in all markets. It seems appropriate therefore to check whether the results hold when we consider the two groups of plants separately. Tables 11 and 12 show that the pattern summarized in Eq. (22) continues to hold for both groups of plants, though there are interesting differences between them. First, domestically-owned plants have almost flat price profiles in export markets, showing that even in differentiated sectors these plants do not compete on quality, though the results suggest that non- differentiated exports come closer to exhibiting cost-based competence. Second, foreign- owned plants compete strongly on quality even in the home market, as we might expect, leveraging their superior brands even in non-differentiated sectors. But they too compete more on cost in export non-differentiated sectors, showing that their exports of these goods do not command quality premia. Overall, the broad pattern from previous tables is confirmed for both groups of plants, and is most pronounced for foreign-owned plants.","This paper has developed a new model of multi-product production in which firms invest to improve the quality of their products as well as the quality of their overall brand. It is thus the first to integrate two important strands of recent work on the behavior of firms in international markets. On the one hand, the growing evidence that many firms, and especially most large exporters, are multi-product, has inspired theoretical and empirical work which focuses on the “intra-firm extensive margin”, changes in the range of products produced by firms, distinct from the inter-firm extensive margin which has attracted so much attention in the literature on heterogeneous single-product firms. On the other hand, an increasing number of authors have suggested that successful firms in international markets compete on the basis of superior quality rather than superior productivity. Our model integrates these two strands in a tractable framework. Crucially, it endogenizes both the choice of product range and the choice of quality, or more specifically, the choice of investment in quality, thus allowing a range of issues to be explored which have so far been little studied. The model has interesting implications for the manner in which firms compete in international markets. In particular, it throws light on the question of whether productivity or quality is the key to successful export performance, and suggests a way of reconciling these two views. Because of flexible manufacturing, firms produce more of products closer to their core competence. They also have incentives to invest more in the quality of those goods. These two effects have opposite implications for the profile of prices. On the one hand, to the extent that consumers view all products as symmetrically differentiated substitutes for each other, firms can only sell more of their core products by charging lower prices for them. Hence, the direct effect of lower production costs for core products is that firms “pile 'em high and sell 'em cheap,” implying that the profiles of prices and sales should be negatively correlated, an outcome we call “cost-based competence”. On the other hand, firms face stronger incentives to invest in raising the perceived quality of their core products, since these are the products with the highest mark-ups. Even though investment in the quality of an individual product is subject to diminishing returns, this implies that firms will invest more in the quality of their core products, so raising the price which consumers are willing to pay for them. This indirect effect of lower production costs for core products implies that the profiles of prices and sales should be positively correlated, an outcome we call “quality-based competence”. We show that both these outcomes are possible in our model, and that which of them prevails depends on a number of exogenous factors. In particular, the greater the degree of product differentiation, the more the firm faces differential incentives to invest in the quality of different products, and so the more likely is the indirect effect to dominate, giving rise to quality-based competence. We prove these results in the text for the case of a single multi-product firm, and show in the Appendix that they also hold in an oligopolistic model with heterogeneous firms. This last prediction is the one we explore empirically, drawing on a unique data set on Mexican plants already used by Iacovone and Javorcik (2007, 2010). A great advantage of this data set is that it gives detailed information on both home and foreign sales at the same level of disaggregation, allowing us to test theoretical predictions about their relative profiles. Our findings show that a two-way distinction is crucial: between home sales and exports on the one hand, and between differentiated and non-differentiated products on the other. In the domestic market, we find that both differentiated and non-differentiated products exhibit quality-based competence, with prices falling as sales value falls. However, this pattern is significantly more pronounced for differentiated products, exactly as our theory predicts. The same holds true in the export market, where the difference in price behavior between the two groups of products is considerably greater: plants in differentiated-product sectors exhibit quality-based competence in export markets, but those in non-differentiated-good sectors exhibit cost-based competence, with core-competence products selling for significantly lower prices on average. These results turn out to be robust to a slew of alternative ways of grouping our data. We find very similar results whether we consider all products or only those which are sold in both home and foreign markets; and whether we consider all plants active in either market or only those active in both. They also hold when we consider only the sub-sample of single-plant firms: confirmation that our theory, which was developed for firms, helps in understanding behavior at plant level too. Finally, the patterns we have found are exhibited by both home-owned and foreign-owned plants. We can conclude that, for this data set, quality- based competence is dominant for firms in differentiated-good sectors, but not for the export sales of firms in non-differentiated-good sectors. Turning to policy, a full consideration of the costs and benefits of different export-promotion strategies is beyond the scope of this paper. Nevertheless, our results are suggestive. Export-promotion policies take a variety of forms, and in the light of our results we can distinguish between those that focus on cost and those that focus on quality. The former type of intervention includes measures to stimulate investment in cost-saving technologies and worker training. The latter type of intervention includes marketing campaigns to stress the advantages of national products, and reductions in the costs of quality certifications (e.g. ISO 9000 or 14000) to improve the producer's image. Without further research, it would be premature to suggest that, in the light of our findings, export-promotion efforts in middle-income countries such as Mexico should focus on helping producers to lower production costs in non-differentiated-good sectors and on improving perceived product quality in differentiated-good sectors. Nevertheless, our findings that Mexican firms compete in foreign markets on either cost or quality, depending on whether they operate in relatively homogeneous or relatively differentiated-good sectors, suggest that a “one- size-fits-all” policy may not be the most effective way of promoting exports. Our findings also have broader implications for the nature of competition in international markets. Our data set shows that within-firm product heterogeneity is not just a rich-country phenomenon, but is also important in at least one middle-income country. Moreover, the evidence we present suggests that only firms in differentiated-product sectors compete in export markets on quality. This has a key implication for understanding how firms compete successfully abroad. While previous studies have shown that all exporters have a productivity premium, our results suggest that those in differentiated-product sectors have a quality premium too, whereas those producing non-differentiated goods behave differently at home and away, competing less on quality and more on price in their export markets."],["The focus of this paper is on news-driven business cycles in small open economies. We make two significant contributions. First, we develop a small open economy model where the presence of financial frictions permits the replication of business cycle co-movements in response to news shocks. Second, we use VAR analysis to identify news shocks using data on four advanced small open economies. We find that expected shocks about the future Total Factor Productivity generate business cycle co-movements in output, hours, consumption and investment. We also find that news shocks are associated with countercyclical current account dynamics. Our findings are robust across a number of alternative identification schemes. --------------------------------------------------------------------------------","Does news about future Total Factor Productivity (TFP) generate business cycles in small open economies? A long tradition in macroeconomics and some recent empirical evidence suggests that news about the future might be an important driver of the business cycle.1 There has been a lot of recent interest in incorporating this idea into modern business cycle models. One of the main challenges emerging from this literature is to develop equilibrium business cycle models that replicate data-congruent macroeconomic co-movement in response to news shocks.2 The emphasis in most of this literature, particularly on the empirical side, is on the effect of news shocks in closed economies. In this paper, we focus on the effect of news shocks in small open economies. We make two contributions. First, we put forward a novel mechanism through which news about future TFP causes business cycles. This mechanism is based on the presence of financial frictions. Specifically, the model incorporates financial frictions à laJermann and Quadrini (2012) into an otherwise canonical small open economy model. The financial friction in this model arises because firms need to arrange a working capital loan prior to production taking place. Access to finance is constrained by the firm's net wealth position. News shocks interact with the financial friction by relaxing the borrowing constraint faced by firms. This allows firms to increase their demand for labour, which raises output and investment in anticipation of future increases in TFP. Greater investment and labour input today creates the expectation of higher dividends in the future, thus raising the share price in anticipation of future TFP. Our second contribution is to identify and analyse the effects of news shock in a set of advanced small open economies. Specifically, we identify the dynamic macroeconomic effects of news shocks in four developed small open economies: Australia, Canada, New Zealand and the United Kingdom. The way news shocks are identified in the data is informed by the theoretical impulse responses of the model. In particular, our identification strategy, based on Beaudry et al. (2011b), imposes model consistent restrictions on the path of TFP, share prices and consumption. For TFP, this implies that news arrives not one, as is the convention in the literature, but two periods in advance. We also impose a restriction that at the end of the news horizon, TFP actually increases for a number of periods. Consistent with our model, we restrict share prices and consumption to rise in response to news. We find consistent evidence that news shocks generate business cycles. As in our theoretical model, a news shock leads to positive co- movement between GDP, hours worked and investment as well as a counter-cyclical trade balance. Our results are robust across a number of alternative identification schemes, including an augmented Barsky and Sims (2011) identification.3 To our knowledge, this is the first account of the effect of news shocks in advanced small open economies. The next section puts our contribution into the context of the literature analysing news shocks. We document the theoretical model and the transmission of the news shocks in Section 3 and 4. Section 5 and 6 present data and our preferred shock identification mechanism. In Section 7 we present our empirical results and perform a number of robustness tests around our baseline identification of news shocks. Further sensitivity analysis is reported in the online appendix to this paper.","News shocks in a standard open economy real business cycle model do not generate news- driven business or Pigou cycles. In this class of model, news about a future TFP improvement creates a positive wealth effect that raises both household consumption and leisure, thus reducing the supply of labour. In the absence of an actual increase in TFP, labour demand remains unchanged and as a result output falls in response to news. Jaimovich and Rebelo (2008, 2009) show that when preferences are such that the wealth effect on labour is small, hours worked do not decline following a news shock. Eliminating the wealth effect on hours is, however, not enough to generate an actual increase in labour demand when an increase in TFP is anticipated but not yet realised. Jaimovich and Rebelo (2009) show that in a closed economy setting, this can be achieved by a combination of investment adjustment costs and variable capital utilisation. The presence of investment adjustment costs causes firms to bring the anticipated increase in future investment, associated with the increase in future TFP, forward into the current period. In a closed economy model, the hump shaped response of investment causes current period Tobin's q to fall, which in turn raises the rate at which firms utilise capital. The rise in capital utilisation raises the marginal product and thus the demand for labour. This is enough to generate new-driven business cycles. For a small open economy, the challenge of generating news-driven business cycles is slightly different. With real interest rates determined abroad, Tobin's q always rises following a news shock. As a result, adding variable capital utilisation does not help generate business cycles in response to news shocks. Jaimovich and Rebelo (2008) show that adding labour adjustment costs that penalise large changes in labour input, will cause firms to bring forward into the current period some of the expected future increase in labour demand.4 Our modelling approach relies not on costly adjustment in the labour market, but on a simple form of financial friction to increase the demand for labour following news about future TFP. This friction introduces a wedge between the marginal product of labour and the real wage. Following Jermann and Quadrini (2012), we assume that because of limited enforcement of financial contracts, firms face an enforcement constraint on working capital loans. Because firms have to borrow the wage bill, the ‘tightness' of the enforcement constraint creates a wedge between the marginal product of labour and the real wage. In our model, good news about future TFP relaxes the enforcement constraint and increases the demand for labour. The intuition behind our results is similar to Pavlov and Weder (2013) who consider a model with counter-cyclical mark-ups. In their model, mark-ups create a wedge between the marginal product of labour and the real wage. A news shock that lowers the mark-up also reduces the labour wedge, raising firms' labour demand. Related to our analysis is Walentin (2012), who examines news shocks in a model with limited enforcement where firms face a collateral constraint when securing external finance. As in our model, the arrival of news relaxes the borrowing constraint and raises share prices, which leads to an accelerator type effect on investment. The key difference between our approach and that of Walentin (2012) is that in our model, the firm faces the enforcement constraint on working capital, whereas in Walentin (2012), the loan is inter-temporal. Financial frictions impact the business cycle via an accelerator type mechanism. This difference between our approaches matters, because only in our case does the financial friction create the aforementioned labour wedge that is helpful in generating a positive co-movement between consumption and hours.5 The literature on news shocks to TFP can be seen as being part of a wider literature on the transmission of total factor productivity shocks. Both news shocks and highly persistent, but contemporaneous, shocks to TFP have an expectations as well as a supply component. In the initial response to a news shock about TFP, the transmission mechanism is dominated by the expectations component of the shock. In a highly persistent shock to TFP, the initial transmission mechanism is affected by both the expectations as well as the supply component. In many dimensions, news and highly persistent shocks may be expected to have similar effects on the economy. In the context of emerging market economies, Aguiar and Gopinath (2007) highlight the importance of highly persistent TFP fluctuations that are akin to shocks to trend growth. These kinds of shocks help explain the somewhat non-standard business cycle fluctuations of emerging market economies. Even for developed economies with ‘standard’ international business cycle characteristics, Corsetti et al. (2008a) show that highly persistent productivity shocks can bring a fairly standard international real business cycle model much closer to the data. In particular, the expectations component of highly persistent TFP shocks creates a wealth effect which helps the model achieve plausible degrees of international risk sharing. Corsetti et al. (2008b) and Corsetti et al. (2014) identify TFP shocks using either long-run and sign restrictions, respectively. In both studies, the US real exchange rate is found to appreciate following a persistent increase in US TFP. Importantly, this real appreciation is not linked to the familiar Balassa-Samuelson mechanism, as both the real exchange rate and the terms of trade are shown to appreciate. In a context of a simple two-good international real business cycle model, this finding is reminiscent of a response to a shock with a significant expectations component. Research by Nam and Wang (2015) lends credence to this view. They use an identification method that divides TFP shocks into a contemporaneous as well as an anticipated, or news, component. When separating the expectations from the supply component of a TFP shock, the authors find that for the US economy, anticipated TFP shocks are associated with a real appreciation, whereas contemporaneous shocks are linked to a depreciation. Whereas the empirical literature on the transmission mechanism of contemporaneous TFP shocks in open economies is well advanced, the empirical literature on in the transmission mechanism of news shocks in open economies, to which our paper is a contribution, is not. There are two notable exceptions. The aforementioned Nam and Wang (2015), who focus exclusively on the US economy and a recent paper by Fratzscher and Straub (2013). The latter authors use a canonical two-country new Keynesian model in which news shocks are used to identify changes in asset prices that are not related to current fundamentals. Related to our work, they analyse asset price shocks in small open economies, including the four in our sample. However, since the focus of their work is on asset price shocks, no effort is made to properly identify news-driven business cycles. Indeed, their baseline model shows a decline in investment following a news shock, and their empirical work does not report the response of investment or employment. In contrast, we identify news shocks as shocks that raise the share price as well as lead to a path of TFP that is consistent with news about future total factor productivity. We use the thus identified shock to check whether news shocks can be a candidate driver of the business cycle and generate business cycle consistent co-movements among macroeconomic aggregates.","We extend the flexible price version of the model presented in Jermann and Quadrini (2012) into a small open economy setting. To turn a closed economy real business cycle model into a small open economy model requires only a few changes to be made to the structure of the model. In an open economy, the savings of households do not have to equal the borrowing by firms. The gap between savings and investment equals the current account balance. Unlike a closed economy, the gross or pre-tax interest rate faced by households and firms is exogenous in a small open economy setting. This rate is determined instead by the world interest rate as well as a small risk premium to ensure a well defined steady state.6 Firms and households produce and consume a homogeneous good. This good is a perfect substitute for output produced in the rest of the world. As a result, the terms of trade defined as the price of imports relative to exports are constant. We make the one-good assumption for two reasons. First, abstracting from terms of trade movements allows us to focus more clearly on the role of financial frictions in the transmission of news shocks. Second, for commodity producers such as Australia, Canada and New Zealand, assuming an exogenous terms of trade is quite realistic, although this is probably not the case for the UK. As in Jermann and Quadrini (2012), we introduce financial frictions into the environment in which domestic firms are operating. The household sector, on the other hand, faces a standard optimisation problem.","In Figs. 1 to 2 we assess the contribution of financial frictions to the generation of positive co-movement between consumption and hours worked in response to news shocks. To do this, we compare our baseline model to a version where the dividend payout cost is set to zero. This version of our model is similar to the model in Jaimovich and Rebelo (2008). The calibration of our model follows Jermann and Quadrini (2012) for the financial frictions part of the model, such that ξ = 0.162 and τ = 0.35. We set κ = 0.146, which corresponds to the value in Jermann and Quadrini (2012). Sensitivity analysis shows that in our baseline calibration with non-separable preferences, we get news-driven business cycles for any positive value of κ. For separable preferences (when γ = 0.999), our model generates news-driven business cycles for any value of κ greater than 0.09. Fig. 1 analyses the response of key macroeconomic aggregates to an increase in TFP that is expected to occur in period t + 2 and announced in period t.8 In our baseline model, hours worked and GDP both increase as soon as the news about future productivity becomes available. Without financial frictions, the agent's preferences over consumption and labour ensure that the wealth effect on hours worked is weak. However, given our value of γ, the wealth effect is small, but not zero, and hence hours worked decline on impact. Because the real interest rate stays largely constant in our small open economy model, share prices rise with the announcement of news. The initial increase in consumption and investment is greater than the response of output, leading to a deterioration of the trade balance and a decline in the economy's net foreign asset position. Because of the way we have closed the model using a small debt-elastic interest rate risk premium, ζ in Eq. (19), the decline in the net foreign asset position raises the real interest rate faced by households and firms and thus also affects the path of consumption and investment. Following a positive news shock about TFP the term Δtφ′(dt) falls, which for a given marginal product of labour raises the real wage. The rise in the real wage causes agents to increase hours worked and thus output to rise. Once TFP increases, the firm's borrowing constraint becomes more binding, Δt rises, due to more output needing to be financed in advance of production. A feature of the Jermann and Quadrini (2012) model is that a tightening of the borrowing constraint causes hours worked to decline. In the next subsection, we analyse why the borrowing constraint is relaxed during the news period, causing hours worked to rise. Intuition ~~~~~~~~~ As dividend payouts decline on impact and then recover, the discount factor falls. In other words, the firm becomes less patient and thus faces a more binding borrowing constraint. As a result, the firm's demand for labour declines. This mechanism explains the observed decline in hours worked once the anticipated increase in TFP materialises in Figs. 1 to 2. When the increase in TFP is anticipated, the firm finds it optimal to start reducing the flow of dividends before TFP actually increases in order to finance increased investment. Thus dividends start to decline as soon as the news about future TFP becomes available. As in the case of an unanticipated increase in TFP, dividend adjustment costs reduce the magnitude of the fall in dividend payouts and smooth the dynamics of dt. The gradual fall in dividends affects the firm's stochastic discount factor via the dividend adjustment cost function. When dividends are expected to fall over time, the firm's discount factor rises, which makes the firm want to hold less debt. Therefore, the firm's desire to gradually reduce dividend payments in order to expand investment and labour input also leads to a de-leveraging of the firm. The combination of a higher expected capital stock and lower inter temporal borrowing thus improves the firm's net asset position, which in turn relaxes the intra-temporal borrowing constraint. The less binding the constraint, the smaller Δt and the greater will be the firm's demand for labour (see Eq. (6)). This mechanism explains the role of financial frictions in generating news- driven business cycles in our model. Importantly, as we show in Fig. 2, this mechanism can be strong enough to off-set the wealth effect on hours worked. Separable preferences ~~~~~~~~~~~~~~~~~~~~~ In this section, we examine whether the mechanism identified above is strong enough to overcome the negative wealth effect on hours that one finds in models with separable preferences over consumption and labour. The utility function put forward by Jaimovich and Rebelo (2008) nests both non-separable (when γ is close to zero) and separable preferences (when γ is close to unity) over consumption and labour. Fig. 2 shows the response to a news shock when γ = 0.99 for an otherwise unchanged calibration. Even with largely separable preferences, GDP and its components display typical business cycle behaviour, with output, consumption and investment and hours all rising on impact. As in the non- separable case, there is a reduction in labour demand once TFP actually increases. This is because as output rises, so does the demand for intra-period finance, which in turn causes the enforcement constraint to tighten.","Next, we estimate a VAR model using data on total factor productivity, stock prices and five macroeconomic aggregates: output, consumption, investment, total hours worked and net trade. Where applicable, our data is normalised by the relevant working age population. For each country our data is from OECD and obtained via HAVER. Our sample period is mainly dictated by data availability and covers 1989Q3 to 2011Q3. Our macroeconomic aggregates and our measure of stock prices are constructed as follows: Output: log(GDP in millions of chain linked domestic currency / working age population); Consumption: log(Private consumption in millions of chain linked domestic currency / working age population); Investment: log(Private non-residential fixed capital formation in millions of chain linked domestic currency / working age population); Total hours: log(Employment in thousands × Hours worked per employee in total economy / working age population); Net trade to GDP: (Export of goods and services in millions of chain linked domestic currency/GDP in millions of chain linked domestic currency - Imports of goods and services in millions of chain linked domestic currency/GDP in millions of chain linked domestic currency); Stock prices: log(Stock prices / Consumer price index). We use the following stock price series: ‘All Ordinaries' for Australia, S&P/TSX Composite index for Canada, NZSX Mid Cap for New Zealand and the FTSE 100 for the UK. A measure of total factor productivity ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our method of constructing quarterly total factor productivity is based on Fernald (2014). Basu et al. (2006) argue that because of sectoral heterogeneity in the marginal product of factors, one should ideally measure TFP at the sectoral level and then aggregate across sectors to derive an aggregate measure of TFP. Furthermore, unobserved variations in factor utilisation must also be accounted for. They proxy the unobserved changes in capital utilisation and labour effort by observed changes in hours worked per capita. Whereas data on capital utilisation, hours worked, output and total labour force are available for our sample of countries at quarterly frequency, we have to construct quarterly data for effort and the capital stock. We follow Basu et al. (2006) and proxy the change in effort by the change in average hours, det = ζdnt. To construct a quarterly series of the capital stock, we use annual capital data from the OECD and transform it to quarterly frequency using information on quarterly investment. Our approach consists of allocating capital into each quarter proportionally to the investment on that particular quarter. Consistently with the calibration of our model we set α to 0.3. The coefficient that links changes in effort to hours worked, ζ is set to 1. We have experimented with values between 0.1 and 10 and found that our results are robust provided ζ is non-zero. Output, yt and total hours nt are defined as above. We construct a quarterly measure of the capital stock using annual OECD data on ‘Productive capital stock’, and quarterly investment. As defined above, investment is taken from the OECD's private non-residential fixed capital formation series. For each quarter, we use the share of total investment over the whole year occurring in this quarter in order to determine this quarter's change in the total capital stock from this year to the next. Therefore, our constructed quarterly capital stock is consistent with annual capital stock taken from the data. HAVER supplies measures of capacity utilisation from a number of sources. For Australia, we use the NAB Business Survey's measure of capacity utilisation, for Canada, we use the ‘Capacity Utilisation - total industry’ measure. For the UK, we use the Harmonized Capacity Utilisation series and for New Zealand, we use capacity utilisation in manufacturing and construction. All series are reported in percentages and divided by 100. Country-specific shocks ~~~~~~~~~~~~~~~~~~~~~~~ It is likely that TFP growth in small open economies has both a domestic as well as a foreign component. To take account of this, we would ideally adjust our TFP series by a measure of global TFP to allow us to isolate country-specific shocks to TFP. There is, however, no obvious measure of world TFP available, hence we take the quarterly TFP series by Fernald (2014), for the United States as a proxy. Stock prices and macroeconomic aggregates ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Stock prices are stock market indexes deflated by CPI for each country. Output and consumption correspond to gross domestic product and private consumption in real prices, respectively. Investment refers to private non-residential fixed capital formation in real prices. Total hours for each country is constructed as a product of employment and hours worked per employee. All macroeconomic aggregates are divided by the working age population and are therefore expressed in per capita terms. Net trade is the the difference between exports and imports of goods and services divided by output.","The number of lags has been selected using information criteria (likelihood ratio test statistic, final prediction error and Akaike's information criterion). All selection criteria suggest that a VAR model with two lags is sufficient to capture the dynamic properties of the macroeconomic data and this is the case for all countries. In order to ensure that our inference is not driven by the selection of a particular lag length we repeat the same analysis using lag choices (VAR(1), VAR(3) and VAR(4)) and the results remain unchanged. For the estimation of the empirical model we rely on Bayesian inference techniques as the large dimension of the observable vector (seven variables) and the small time span of the macroeconomic data set cause Classical (OLS) estimates to be subject of considerable uncertainty. In this case data is combined with prior information about the reduced-form parameter vector in the form of a probability density function. Similar to Beaudry et al. (2011b) we employ a flat Normal-Wishart conjugate prior that leads (after being combined with the likelihood of the model) to a closed-form posterior probability distribution for the VAR parameter vector9. Identification ~~~~~~~~~~~~~~ As in Beaudry et al. (2011b), we identify news TFP shock using a combination of zero type and sign restrictions (Uhlig (2005)). To be precise, the TFP shock anticipated in t + h period is identified by imposing zero restrictions on TFP for periods t, t + 1, …, t + h − 1 and sign restrictions on the responses of a set of variables in the system. Our theoretical model suggests that news shocks are associated with positive co-movement between macroeconomic aggregates and share prices as well as a counter-cyclical current account. Based on this analysis, we impose that stock prices and consumption increase after a positive news about future TFP. Our methodological strategy also enables us to consider, in a VAR framework, news shocks beyond the first period (see, Barsky and Sims, 2011 and Beaudry et al., 2011b). That makes the VAR identified responses more comparable to DSGE ones so the former responses can serve a useful device to either assess the empirical predictions of the structural model about new shocks and/or to calibrate the structural parameter vector in order for the DSGE model to replicate the responses estimated in the data. This closes an important gap in the literature since so far the comparison was achievable only for DSGE model with only one quarter anticipation period (see Barsky and Sims, 2011, Kurmann and Otrok, 2013, Theodoridis and Zanetti, 2013 and Pinter et al., 2013). In summary, the identification puts restrictions on the responses of TFP, consumption and share prices. TFP is restricted to remain at zero for two quarters and positive for the following two quarters. Consumption and share prices are restricted to increase for four quarters following a positive news shock. Testing the identification technique on simulated model data ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In this section, we perform a Monte-Carlo experiment to compare the true impulse responses to news shocks from the model presented in the previous section and those from a VAR using our identification scheme. In order to estimate a VAR with as many variables as we have in the empirical section, we introduce additional structural shocks into our model. In addition to news about TFP, we also consider a standard contemporary TFP shock plus five additional shocks. These are shocks to investment specific technology, financial shocks as in Jermann and Quadrini (2012), government spending shocks and preference shocks and a shock to the world real interest rate. The standard deviation of the two unanticipated technology shocks (TFP and investment specific) is set to 0.75%, the standard deviation of the news shock is set to 0.25% and the standard deviation of all other shocks is set to 0.15%. All shocks are assumed to have an AR(1) coefficient of 0.95. We generate 2000 artificial data sets from the model under our calibration corresponding to Fig. 1. The sample size of each data set is 86, corresponding to the number of observations of the actual data. Using the artificial data sets, we estimate VARs with TFP, output, consumption, investment, total hours worked, net trade and stock prices as observables. The VARs include two lags. We then identify impulse responses using the same identification restriction employed in the empirical section. Fig. 3 displays the theoretical and the range of estimated impulses responses over the Monte-Carlo repetitions. Model based impulse responses are dashed blue and the shaded area represents ± one standard deviation confidence intervals for the estimated impulse responses. The VAR based impulse responses are able to capture the model based dynamics following a news shock. Output and its components, total hours as well as share prices, are all estimated to respond positively while the VAR correctly identifies the negative response of net trade. The VARs somewhat underestimate the magnitude of macroeconomic effects of a news shock, particularly at longer horizons. This bias is probably related to the truncation of the empirical model, as we only use two lags. When we expand the lag order to 4, this bias almost disappears. Nevertheless, the theoretical impulse responses almost always lie within the range of estimated impulse responses. These result suggest that our identification successfully recovers news shocks. As in Barsky and Sims (2011), we argue that our Monte Carlo results suggest that the invertibility or ‘fundamentalness' issue raised by Fernndez-Villaverde et al. (2007) is not a major cause for concern for our work.","In this section, we discuss impulse responses and forecast error variance decompositions obtained from the structural VAR. Fig. 4 shows the estimated impulse responses in four countries for TFP, output, consumption, total hours, investment, net trade and stock prices. The grey shaded area represents 16th and 84th quantiles. As per our identification restrictions, TFP remains constant for two periods and increases after the zero restrictions end, and consumption and stock prices increase on impact. After the identification restriction, the impact of the news shock is significantly positive for all these three variables. The quantitative impact of the news shock is similar for TFP and consumption in all countries. Possibly reflecting their volatility, the impact response of stock prices are one order of magnitude higher than the response of TFP. Now consider the response of other variables which are not subject to any identification restrictions. A first robust finding is that in all countries, output is estimated to increase in response to news. Except for New Zealand, output rises on impact. This is, for example, contrary to Barsky and Sims (2011) who find that, in the US, the output response to a news about future TFP tracks, but does not anticipate the movements in the estimated path of TFP. In our sample of small open economies, positive news about future TFP creates an economic boom on impact and its positive effects on output are estimated to be persistent. Second, total hours and investment are also estimated to persistently increase after a news shock. There is, however, somewhat more heterogeneity in the response of these two variables. In the case of the United Kingdom, in particular, the response of these two variables seem to be building up over time and less front loaded than in the other countries. The impulse responses described so far imply that, for the small open economies we consider, news shock generate a positive co-movement between output, consumption, investment and total hours. Third, net trade is countercyclical in all countries except Australia following a positive news shock, with some differences in the initial responses. Except for Canada, the initial response of net trade is small and is followed with further deteriorations. Comparing Figs. 1 and 4 shows that the estimated impulse responses are qualitatively similar to those of our model. In both cases, there is co-movement between GDP, consumption, investment and hours as well as a counter-cyclical trade balance. Model and data do, however, differ with respect to the persistence of GDP, consumption and hours as well as with respect to the magnitude of the response of share prices. Compared to the data, the model predicts a more persistent response for GDP, consumption and hours and a less volatile one for share prices. We now turn to the relative importance of the identified news shocks in shaping the business cycle dynamics in our set of small open economies. Table 1 reports the share of the news shock in the forecast error variance decomposition for the seven variables in the VAR. The news shock accounts for between 6% to 40% of the 10-quarter ahead forecast error variance of GDP. For consumption, the figures are between 13% and 42%, and for investment between 11% and 44%. The shock also accounts for between 7% and 20% of the 10-quarter ahead forecast error variance of net trade. The large range of these results reflects heterogeneity between countries. In the United Kingdom, the contribution of new shocks to the 10-quarter ahead forecast error variance of GDP is much greater than in Australia, Canada or New Zealand. In New Zealand, news shocks appear to play only a minor role. Whereas in the UK, they account for around 40% of the 10-quarter ahead forecast error variance of GDP, in New Zealand these shocks only contribute around 6%. The role of news shocks in Australian and Canadian GDP is more important than in New Zealand, but in both of these countries the forecast error variance is still only about a two-thirds to half of that of the UK. That country-specific news shocks are somewhat less important in the smaller and more open economies in our sample is in line with the well documented importance of foreign shocks for these economies, see for instance Justiniano and Preston (2010). News shocks and relative prices ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The choice of variables in our baseline VAR is determined by our theoretical model. As such, we abstract from relative price movements and focus on net trade. In Fig. 5, we examine a VAR augmented by two measures of relative prices, country i’s real exchange rate vis-a-vis the United States and the terms of trade measured as the relative price of exports to imports. In both measures, an increase denotes an appreciation. We also replace net trade by real exports and real imports. In the VAR, the response of the real exchange rate is not restricted, while the identification scheme is the same as our baseline formulation. For all countries, the median response the real exchange rate to news shock is an initial real appreciation. This appreciation is, however, relatively short lived and not statistically significant for New Zealand. As the news about TFP is realised and TFP increases, the real exchange rate depreciates as the increased supply of home produced goods depresses its relative price. The qualitative dynamics of the other variables in our VAR remain unchanged. For Canada and the UK, we also observe a significant appreciation of the terms of trade. For our two antipodean commodity exporters, the initial response of the terms of trade is not statistically significant. This suggests that for these small open economies, the terms of trade is exogenous. For Canada and the UK, the real appreciation of the terms of trade suggests that the observed real exchange rate appreciation is not attributable to an increase in the relative price of non-traded goods. News shocks and behaviour of imports and exports ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ So far, our analysis suggests that, except for Australia, news shocks cause a worsening of the net trade position. This is in line with our theoretical model. In a small open economy, the demand for exported goods depends on its relative price and foreign aggregate demand. In Fig. 5, we follow Corsetti et al. (2014) and analyse real exports and real imports separately. For all countries in our sample, a positive news shock is associated with an increase in both real imports and real exports. Except for Australia, the response of imports is greater than the response of exports. Figs. 4 and 5 suggest that a consistent picture emerges regarding macroeconomic dynamics conditional on news shocks in small open economies. Our results indicate that positive news shocks about future TFP are associated with initial increases in output, consumption, investment, total hours, stock prices and are associated with countercyclical net trade dynamics. These results appear to be in contrast with the recent empirical work focusing on the effect of TFP news shock on the US economy.","News about future TFP can be a source of business cycle fluctuations in small open economies. For a set of advanced small open economies, we show that news about future TFP causes positive co-movement between GDP, hours, consumption and investment. News shocks are also associated with counter-cyclical current accounts. This is in contrast with previous studies focusing on the US economy in which news about future productivity are not associated with economic booms. We also find that the contribution of country specific news shock in the forecast error variance decomposition of macroeconomic variables is relatively modest. This is possibly due to larger share of foreign variables in driving business cycle dynamics in these small open economies. In addition to our empirical contribution, we also put forward a theoretical small open economy model that is able to generate business cycles from news shocks to TFP. We introduce financial frictions, akin to those in Jermann and Quadrini (2012) as a mechanism to generate the positive co- movement between hours worked and consumption that is a challenge for canonical small open economy models. Our modelling approach is deliberately parsimonious in order to put forward a particular channel in generating news-driven business cycles."],["The use of renewable energy sources is a major strategy to mitigate climate change. Yet Sinn (2017) argues that excessive electrical storage requirements limit the further expansion of variable wind and solar energy. We question, and alter, strong implicit assumptions of Sinn's approach and find that storage needs are considerably lower, up to two orders of magnitude. First, we move away from corner solutions by allowing for combinations of storage and renewable curtailment. Second, we specify a parsimonious optimization model that explicitly considers an economic efficiency perspective. We conclude that electrical storage is unlikely to limit the transition to renewable energy. --------------------------------------------------------------------------------","In the 2015 Paris Agreement, the world agreed on ambitious targets for reducing greenhouse gas emissions to combat climate change (United Nations, 2015). The use of renewable energy sources is a major strategy for decarbonizing the global economy. As the potentials of hydro, biomass or geothermal energy are limited in many countries, wind power and solar photovoltaics (PV) play an increasing role. For example in Germany, often considered as a frontrunner in the use of variable renewable energy sources, the government plans to expand the share of renewable energy in gross electricity consumption to at least 80% by 2050, compared to 36% in 2017 and only around 3% in the early 1990s (Federal Ministry for Economic Affairs and Energy, 2018). Closing this gap requires a massive further expansion of wind and solar power. Opposed to dispatchable technologies like coal- or natural gas- fired power plants that can produce whenever economically attractive, electricity generation from wind and solar PV plants is variable: it depends on exogenous weather conditions, the time of day, season, and location (Edenhofer et al., 2013; Joskow, 2011). At the same time, maintaining power system stability requires to continuously ensure that supply meets demand. The potential temporal mismatch of supply and demand raises two fundamental questions: how to deal with variable renewable energy at times when there is too much supply, and how to serve demand at times when supply is scarce (cf. also Brown et al., 2018). Evidently, electrical storage can provide a solution, for instance, in the form of batteries or pumped-hydro storage plants, allowing to shift energy over time. In a recent analysis, Sinn (2017) argues that electrical storage requirements may become excessive and could thus impede the further expansion of variable wind and solar power in Germany. Based on historic time series of electricity demand and variable renewable energy supply, he illustrates that without storage a fully renewable electricity supply would imply not using 61% of the possible power generation from wind and solar generators. In contrast, to avoid any “waste” of renewable energy, storage requirements to take up renewable surplus energy1 quickly rise to vast numbers. Under such a strategy, current German storage installations would not allow a share of wind and solar PV in electricity demand greater than 30%.2 And for a fully renewable electricity supply, storage requirements would be more than 400 times as high as the currently installed German pumped-hydro storage capacity, and also much higher than the entire European potential to build such plants (van de Vegte, 2015). These considerations deserve merit as they illustrate essential properties of variable renewable energy sources, which gain relevance throughout the world. Accordingly, economic research on renewables plays an important role in informing policymakers, and Sinn’s analysis can be expected to be widely received both in policy and academic circles.3 Yet the approach is based on strong implicit assumptions, two of which are particularly questionable. First, it only considers two extreme cases in which either all surplus energy is stored or none. In turn, either storage needs are excessive or an excessive share of the available renewable energy is not used. An economically efficient solution is likely to be located in between, i.e., combines some amount of storage and some renewable curtailment. Second, it does not explicitly consider an economic efficiency perspective. Sinn’s approach minimizes the storage energy capacity under the constraint that renewables must satisfy a specified proportion of annual electricity demand. Yet an economically efficient solution would seek to minimize the cost to reach that specified proportion of renewables. Such a solution trades off the costs of investments into storage plants, renewables that may get curtailed at times, and other assets in the power market. We address, and alter, these implicit assumptions and show that their effects are significant. Both results and conclusions change substantially. When moving away from corner solutions, storage needs are up to two orders of magnitude lower in a framework otherwise identical to Sinn (2017). Using a parsimonious optimization model with a more suitable economic objective function that leads to first-best solutions, we also find moderate storage requirements. They are even lower if we consider a future broadening of the electricity sector, that is, an additional and flexible use of renewable electricity in other sectors. Throughout the paper, we provide the economic intuition of what drives storage requirements and use. Variable renewable energy sources are not only variable in supply, they are also nearly free of variable cost. A wind or solar PV plant generates electricity whenever the wind blows or the sun shines without requiring any fuel. Curtailment of renewable energy denotes the operation of a wind or PV plant below its actual temporary generation potential, that is, neither consuming the actually available renewable energy in the moment of generation nor storing it for later use. Analogously, also conventional power plants do not generate electricity at full capacity at all times. The rationale is the following: if electricity demand is satisfied, electrical storage can be used to take up renewable surplus energy. Yet integrating increasing amounts of such surpluses requires disproportionately growing storage capacities that are not valuable at most times (Denholm and Hand, 2011; Schill, 2014). Thus, a corner solution avoiding any curtailment likely leads to inefficiently high storage requirements. Instead, an efficient solution seeks to balance investments into storage, renewables that get curtailed at times, and other capacities to minimize the total cost of providing electricity. The remainder of this paper proceeds as follows: In Section 2, we show that Sinn’s results are outliers compared to the established literature. We then replicate his findings using open data and an open software tool in Section 3. In Section 4, we extend the basic model to target solutions between the two extreme cases. In Section 5, we devise a parsimonious model to endogenously determine optimal storage and renewable capacities as well as renewable curtailment levels. In Section 6, we discuss further relevant factors and flexibility options that influence storage needs. Section 7 concludes that electrical storage requirements are not likely to limit the transition to renewable energy.","Researchers from various fields have addressed the nexus of variable renewable energy and storage. Several review papers highlight different perspectives: the economic and regulatory challenges of integrating variable renewable energy sources (Perez-Arriaga and Batlle, 2012), features of techno-economic models required to generate policy-relevant insights (Pfenninger et al., 2014),4 and the role of long-term storage (Blanco and Faaij, 2018). A synthesis of model-based analyses suggests that electrical storage requirements for renewable energy integration are generally moderate. They may only increase substantially in scenarios approaching a fully renewable energy system (Zerrahn and Schill, 2017). For Germany, Sinn (2017) derives electrical storage capacity needs of 2,100 gigawatt hours (GWh) (5,800 GWh, 16,300 GWh), corresponding to 0.42% (1.15%, 3.23%) of yearly electricity demand, to achieve combined shares of wind and solar power of 50% (68%, 89%). To put these numbers into perspective, we compare Sinn’s results with other studies on future electricity systems with high shares of variable renewables. Also for Germany, Schill and Zerrahn (2018) determine optimal storage requirements in long-run scenarios. For 68% (78%, 88%) variable renewables,5 they arrive at 55 GWh (159 GWh, 436 GWh) storage, corresponding to 0.01% (0.03%, 0.09%) of annual demand. Further results on storage needs for the German energy transition are available among policy studies (Fraunhofer UMSICHT and Fraunhofer IWES, 2014). A particularly influential study concludes that hardly any additional storage investments are necessary in Germany and Europe in the short and medium term (Pape et al., 2014): In 2050 scenarios with European shares of variable renewables around 40% (45%, 55%), additional storage capacity between around 14 and 650 GWh is needed (corresponding to 0.00% to 0.02% of annual demand), largely located outside Germany. A long-term climate policy study commissioned by the German environmental ministry (BMUB) also finds that around 170 GWh of pumped-hydro storage (0.02% to 0.03% of annual demand) suffice to achieve variable renewable shares between 83% and 91% (Repenning et al., 2015). For Europe, Scholz et al. (2017) derive cost-minimal storage capacities corresponding to about 0.08% (0.28%) of yearly demand to achieve 74% (85%) variable renewables in a setting with equal contributions of wind and solar power. Using the same numerical model, Cebulla et al. (2017) derive larger storage needs of around 1% of yearly load at a variable renewables share of 80% in a transmission-constrained European scenario. This number decreases to 0.5% in case of increasing transmission capacity. In a recent long-term scenario commissioned by the German energy ministry (BMWi), storage capacities in the size of 0.01% of yearly demand are enough to achieve a pan-European renewable share of 65% (Federal Ministry for Economic Affairs and Energy, 2017). For the U.S., MacDonald et al. (2016) find that integrating up to 55% variable renewables in 2030 does not require any electrical storage. Instead, pan-U.S. geographical balancing, facilitated by transmission investments, mitigates the variability of wind and solar power. For the Pennsylvania-New Jersey-Maryland market, Budischak et al. (2013) conclude that a large stock of electric vehicle batteries, corresponding to 0.3% of yearly overall demand, would enable pushing the renewables share to 99.9% in 99.9% of all hours. Using stationary batteries would—at higher overall cost—require even less storage capacity. Jacobson et al. (2015) show that a fully renewable (wind and solar power contributing 90%) U.S. energy system covering all end use sectors would be possible with an electrical storage capacity smaller than 0.1% of yearly electricity demand.6 For Texas, papers with different approaches also conclude on moderate storage requirements: capacities corresponding to around 0.02% (0.06%, 0.14%) of annual demand would suffice to integrate a combined share of wind and solar PV of 55% (70%, 80%) with relatively low renewable curtailment (de Sisternes et al., 2016; Denholm and Hand, 2011; Denholm and Mai, 2017). Safaei and Keith (2015) derive optimal storage deployment of 0.10% of annual demand for a 66% wind power share. Based on our review, Fig. 1 plots the shares of variable renewable energy against storage energy requirements, normalized by yearly energy demand, and contrasts them with Sinn’s findings.7 It also includes the results of this paper from Section 5. Two findings stand out: first, storage needs disproportionately grow with higher renewables shares. Therefore, the vertical axis is provided in a logarithmic scale. This is driven by the distribution of surplus energy, which has high peaks in only a few hours of the year and is very small or zero in most other hours. Second, storage requirements found in the literature are considerably lower than those calculated by Sinn (2017)—often by at least an order of magnitude. What drives these large differences? One important factor is renewable curtailment; the data labels in Fig. 1 provide the numbers in percent of the annually available renewable energy. Sinn focuses on corner solutions without renewable curtailment as illustrated by the outer line: the storage has to take up every potential kilowatt hour of surplus energy that could be generated by wind power and solar PV generators, which leads to strongly increasing storage requirements for growing shares of variable renewable energy sources. By contrast, the literature agrees that a complete integration of variable renewable energy is not desirable (see, in particular, Budischak et al., 2013; Schill, 2014; Schill and Zerrahn, 2018; Ueckerdt et al., 2017). A combination of renewable capacity oversizing and some temporary curtailment substitutes storage expansion when imposing economic efficiency criteria, such as finding a technology portfolio for least-cost renewable energy supply. In equilibrium, the marginal effects of adding another unit of storage and adding another unit of renewable generation that gets curtailed at times are equal. A social planner would thus trade off storage against renewable curtailment and other options that can provide flexibility. Accordingly, renewable curtailment is not necessarily inefficient. Beyond curtailment, further options on both the supply and demand sides can provide flexibility for variable renewable energy sources and thus substitute for electrical storage (Lund et al., 2015). These comprise geographical balancing (Fürsch et al., 2013; Haller et al., 2012; MacDonald et al., 2016), demand-side management (Pape et al., 2014; Schill and Zerrahn, 2018), and the flexible use of renewables in other sectors such as heat or mobility (Budischak et al., 2013; Jacobson et al., 2015). To be concise, we largely abstract from such options in our analysis, as in Sinn’s original framework. We further discuss this avenue in Section 6. Our literature review highlights two main insights: first, Sinn’s findings are outliers compared to the consensus of established studies. Second, his extreme findings are driven by not considering relevant economic trade-offs concerning the provision of flexibility, in particular by neglecting potential renewable curtailment.","We first replicate the central results of Sinn’s analysis, using a spreadsheet tool and open-source input data. Following recent discussions on good practice in the field of energy research (Pfenninger, 2017; Pfenninger et al., 2018), we provide our tools and all input parameters under a permissive open-source license in a public repository.8 Focus of our analysis ~~~~~~~~~~~~~~~~~~~~~ In our replication, we focus on the central Section 6 in Sinn (2017). Here, he derives storage requirements to integrate increasing shares of variable renewable energy from wind and solar PV in final electricity demand, ranging between 16.6% and 89%. In Sections 7 and 8, he provides stylized geographical extensions of this approach; while these are illustrative, we stick to the central model and its mechanics from Section 6.9 In Sections 2–4, Sinn suggests transforming variable renewable supply to a perfectly constant output over all hours of a year. Such smoothing results in excessive storage needs. However, there is no economic or technical reason for this kind of smoothing. It seems to be inspired by the notion that renewable generators should mimic the characteristics of conventional power plants. In this case, additional backup capacities (referred to as “double structures”) would be obsolete. However, it is not clear why using existing backup power plants should be less desirable than installing additional electrical storage; the approach is silent about any efficiency or optimality criteria. The lack of practical relevance is illustrated by the fact that resulting storage requirements cannot be empirically observed in countries with high variable renewables shares like Denmark, Ireland or Spain.10 Replication using open data and an open software tool ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We derive input data from the Open Power System Data platform, which collects and provides European electricity market data from official sources (Open Power System Data, 2017). Input parameters comprise hourly time series of realized German electricity demand and availability of onshore wind power and solar PV, defined as capacity factors between zero and one. Electricity demand enters the analysis as inelastic, that is, we do not fit any demand curves. This assumption follows Sinn (2017) and is also standard in much of the literature.11 We calculate capacity factors by relating historic hourly feed-in to installed renewable generation capacity in respective hours. As in Sinn (2017), all input data is taken from the base year 2014. The approach aims at finding the smallest possible storage size to integrate all renewable generation. It is equivalent to Sinn’s objective of scaling the storage such that it is empty in at least one hour. The hourly storage use pattern is shifted up and down until the minimum storage size is found. To avoid free lunch, we require the storage level in the first and last hour of the year to be equal. Using open data and open software tool, we are able to replicate Sinn’s central findings.13 Additionally, we provide results for further shares of variable renewable electricity ranging between 20 and 90%. Fig. 2 shows storage requirements for varying shares of renewable electricity in final demand. Results from our calculations are given in black, Sinn’s results in gray. Storage needs rise sharply if more renewable electricity must be integrated. While current German pumped-hydro storage installations of somewhat below 0.04 terawatt hours (TWh) would suffice to fully integrate almost 30% renewable electricity, even moderate further increases in renewables would substantially drive up storage requirements. For 50% variable renewables, they already amount to 2.1 TWh, that is, they are two orders of magnitude higher. As Sinn does not provide his data and calculations open-source, we cannot trace back the small numerical differences to our findings to a specific reason; presumably, they arise due to slight differences in the input data. (Non-)robustness ~~~~~~~~~~~~~~~~ We address the robustness of findings in a sensitivity analysis using different base years. Both the time series of demand and the availability of renewable energy may change substantially between years. To this end, we repeat the basic analysis using hourly time series of demand and the renewables capacity factor from the base years 2012, 2013, 2015, and 2016. Fig. 3 depicts the results: storage requirements turn out to be highly sensitive to the choice of the base year. Taking data from 2014 yields the highest storage needs up to a variable renewables share of 65%. In contrast, data from 2015 leads to the smallest storage sizes up to 65% renewables. For instance, comparing the 50% renewables case, storage installations are less than a fifth for 2015 data as compared to 2014 data. For a high renewable penetration beyond 65%, data from the base year 2013 leads to the smallest storage. Using 2014 data generally yields comparatively large storage capacities. We explain the drivers of storage requirements below. Intuition: the residual load duration curve ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To gain intuition what drives storage requirements, we use the concept of residual load duration curves (RLDCs). Residual load—also referred to as net load—is defined as hourly demand minus renewable feed-in in the respective hour. It is the remaining load that conventional plants or storage installations must serve.14 A residual load duration curve is a graphical representation of residual load of all 8760 hours of a year, sorted in descending order. The positive integral of the curve corresponds to the energy that must be provided by non-renewable energy, storage generation or imports. The RLDC concept is prominent in the energy economics literature (compare also Ueckerdt et al., 2017). Fig. 4 shows RLDCs for the above analysis in case of 80% renewables, using 2014 as the base year.15 The solid line shows residual load before storage use. To the right, below the horizontal axis, residual load is negative, indicated by area A. In these hours, there is a surplus of renewable generation. To the left, there are hours with positive residual load, areas B and C. In these hours, there is relatively high demand and low renewable energy generation. The dotted line shows the RLDC after storage use.16 The storage takes up energy in hours of excess renewable supply and releases it in hours with excess demand. Graphically, it shifts the surplus energy represented by area A to area B, which equals the size of area A reduced by the storage’s efficiency losses. Area C represents the annual residual load that must be served by other generators, for instance conventional or dispatchable renewable plants. By assumption, it corresponds to 20% of total demand in this case.17 Importantly, nothing forces the storage to shift energy to hours with highest residual loads, that is, greatest scarcity, at the very left-hand side of the RLDC; the operational heuristic prescribes to empty the storage as soon as residual demand is positive again. With the present patterns of demand and renewable feed-in, it is unlikely that an hour with very high residual load follows closely to an hour with renewable surplus generation. The RLDC representation dissolves the temporal sequence of hours during the year.","The residual load duration curve illustrates the two challenges of integrating high shares of variable renewable electricity: (i) on the left-hand side of the curve, there are hours with high demand that variable renewables cannot directly supply; (ii) on the right-hand side, there are hours with renewable surplus generation. Both sides of the RLDC have an impact on storage requirements. Power-oriented renewable curtailment ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The RLDCs suggest that curtailment of renewable surpluses may reduce storage requirements. We first devise a strategy that allows curtailment of all renewable energy surpluses beyond a defined threshold. This threshold can be interpreted as the power capacity of a storage (in megawatt, MW); this is the energy the storage can take up per hour—as opposed to the energy capacity the storage can accommodate in total (in megawatt hours, MWh). This distinction is an essential characteristic of any electric storage technology and missing in Sinn (2017). In case of pumped-hydro storage, the power capacity describes the power of the pumps or of the turbine to generate electricity; the energy capacity indicates the volume (in energy terms) of the storage basin. We extend the basic model by a renewable curtailment threshold, equivalent to the power capacity of the storage. If hourly surplus generation is below the threshold, it is channeled into the storage; all renewable surplus energy beyond the threshold is curtailed. Otherwise, the model is identical to Section 3 and storage use remains myopic. To gain some intuition, Fig. 5 shows RLDCs for 80% renewables in final demand and a curtailment threshold of 44.1 gigawatt (GW), which leads to curtailment of 5% of maximum yearly renewable generation. The threshold is indicated by the horizontal solid gray part of the RLDC after renewable curtailment on the right-hand side. Area D represents all renewable surplus that is curtailed. The storage shifts the remaining surplus energy from area A to area B. The remaining energy demand, area C, is supplied by other, unspecified technologies. This renewable curtailment strategy avoids storing the most excessive surplus events. We iterate through combinations of minimum renewable requirements and maximum allowed renewable curtailment. In doing so, the spreadsheet model endogenously determines the curtailment threshold—or storage power capacity—such that no more renewable energy is curtailed than maximally allowed. Fig. 6 shows the results. It is evident that increasing levels of renewable curtailment lead to lower storage requirements. The decrease is close to linear though somewhat convex. For instance, while a complete integration of 50% variable renewable electricity triggers 2.1 TWh storage energy capacity, allowing curtailment of 5% of the annual renewable generation reduces storage needs to 0.3 TWh. Specifically, we provide solutions that lie between the two extremes “no renewable curtailment” and “no storage”, which Sinn (2017) only considers. The vertical axis of Fig. 6 indicates storage requirements for the corner solution if no renewable curtailment is allowed. The numbers are identical to those that we replicate from Sinn’s approach (compare Fig. 2 and the right panel of his Table 1). The horizontal axis shows renewable curtailment levels for the corner solution if no storage is allowed. Here, we also replicate Sinn’s findings on “efficiency losses” that are given in the left panel of his Table 1. For instance, for 50% renewables, our model returns a renewable curtailment between 6.0 and 6.5%, as indicated by the point where the solid gray line intersects the horizontal axis in Fig. 6. For the same case, Sinn determines an “efficiency” of 93.8%, which corresponds to curtailment of 6.2%.18 Thus, we provide a solution space combining curtailment and storage that lies between the two extreme cases. For other base years, results are qualitatively unchanged; however, they exhibit great variation concerning the level of required storage. To achieve the same share of renewable energy in final demand, the required renewable capacities are necessarily higher when allowing for curtailment, that is, if some of the available energy is not used. However, this increase is moderate, as Fig. 7 shows. For instance, achieving 50% renewable energy in final demand requires 214 GW renewables without curtailment. With 5% curtailment, the necessary renewable capacities are somewhat higher, at 226 GW. Yet renewable curtailment does not increase the necessary backup capacities to supply the remaining residual electricity demand after storage. They are no larger than in the case without renewable curtailment. Inspecting the left-most part of the RLDCs in Figs. 4 and 5, there is also no reason to assume so.19 Energy-oriented renewable curtailment ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ While the power-oriented storage strategy—curtailing renewable surplus whenever it exceeds a defined threshold—seems plausible, it may not be optimal with respect to finding the smallest required storage energy capacity. Given historic input data, it turns out that extended periods of renewable surpluses in contiguous hours determine the maximum energy capacity of the storage, and not single periods with the highest surplus generation. For instance, storing a moderate surplus in ten consecutive hours may require more storage than storing an extreme surplus event in one hour. Therefore, we alternatively implement an energy-oriented renewable curtailment strategy. The storage operational pattern remains myopic and is identical to the above cases; however, renewable curtailment occurs if and only if the storage is fully loaded. Thus, it targets a minimum energy capacity requirement. Again, we iterate through minimum renewable requirements and maximum renewable curtailment constraints to explore the solution space and endogenously determine minimum storage capacities. To provide some intuition, Fig. 8 shows the resulting residual load duration curves for the case of 80% renewables and a maximum curtailment of 5% of the annual renewable energy. Curtailed energy, area D, is identical to the one under the power-oriented renewable curtailment strategy, area D in Fig. 6. However, renewable curtailment is concentrated in hours in which surpluses trigger the highest storage requirements; these are not necessarily the hours with the highest surplus generation. Storage shifts the remaining surplus generation, area A, to hours with positive residual load, area B. Note that renewable curtailment and the storage operational pattern are still myopic and deterministic, that is, they do not require perfect foresight: surplus energy is charged into the storage as long as there are free capacities, and is curtailed otherwise. The stored energy serves residual load as soon as it is positive again. Fig. 9 shows the resulting storage requirements for varying minimum shares of renewable energy. Overall, storage requirements decrease substantially even if only small levels of renewable curtailment are allowed. Again, the intersection points with the axes represent Sinn’s corner solutions: on the vertical axis all renewable energy must be integrated; on the horizontal axis no storage is available. Any combination of both options yields significantly lower storage needs; the decrease is much more convex than under the power- oriented renewable curtailment strategy; already a small curtailment budget triggers a large effect.20 For instance, while a complete integration of 50% variable renewable electricity requires 2.1 TWh storage energy capacity, allowing for 5% renewable curtailment reduces storage needs to 0.019 TWh, or 19 GWh. This is one order of magnitude lower than under the power-oriented curtailment strategy, two orders of magnitude lower than without renewable curtailment, and less than the pumped-hydro power capacity installed in Germany by 2018. Allowing 11% curtailment, 41 GWh of storage, slightly more than installed in Germany by 2018, would suffice to reach 60% variable renewable energy.","The data-driven analysis in Section 4 alters one important implicit assumption of Sinn’s approach: plausible solutions lie between the two corner solutions of no renewable curtailment or no storage. However, the approach is still unlikely to result in an efficient market outcome because of the objective function used. From an economic perspective, finding least-cost solutions is relevant, that is, cost-minimal combinations of conventional and renewable generation, renewable curtailment, and electrical storage.21 As we also include a stylized representation of conventional generators, we explicitly address what Sinn refers to as “double structure buffering.” In this context, both storage energy capacity (in MWh) and storage power capacity (in MW) matter with respect to costs.22 To find optimal solutions, we employ a stylized and parsimonious numerical optimization model.23 We provide the source code and all input data under a permissive open-source license in a public repository.24 The economic optimization approach addresses both challenges of renewable energy integration. First, it delivers an efficient solution to the trade-off how much and when renewable surplus energy to curtail, and how much and when to store. This corresponds to the right-hand side of the residual load duration curve. Second, it determines efficient conventional, renewable, and storage capacities to serve demand at any point in time; this corresponds to the left-hand side of the residual load duration curve. The results of the cost minimization model can be interpreted as long-run equilibria under the assumption of perfect competition and complete information. The model thus mimics a first-best social planner approach. The model ~~~~~~~~~ The numerical model minimizes the total costs of satisfying electricity demand in every hour h of a year. The objective function (1) sums the products of specific investment costs κi and capacity entry N of storage, differentiated by energy Nse and power Nsp, renewables Nr, and conventional capacity Nc. Throughout the model, upper-case Roman letters indicate variables. For conciseness, we consider one stylized technology for variable renewables that aggregates the generation patterns of onshore wind power and solar PV,25 and two stylized conventional technologies c ∈ {base, peak}, parameterized to lignite and natural gas plants. Base generators incur high capacity costs and low variable costs, and vice versa for peak plants. We parameterize storage according to pumped-hydro storage, which features relatively high costs for power capacity κi,sp and relatively low costs for energy capacity κi,se. By focusing on pumped-hydro storage, we follow the narrative in Sinn (2017). Moreover, it is a mature technology and its cost structure renders it a favorable medium-term storage. This complements the variability of renewables well, especially with regard to the diurnal fluctuations of PV.26 Investment costs are annualized using typical lifetimes of power plants. For simplicity, we abstract from the lumpiness of investments. All variables are continuous and positive. Renewable energy does not incur any variable costs. We also do not impose any costs for curtailment of renewables. As the objective function comprises the investment costs of renewable plants, it accounts for the full cost of renewable energy irrespective whether it eventually satisfies demand or is curtailed at times.27 By analogy, also the use of conventional plants below capacity does not receive any penalty in the objective function. The model is a linear program and solved numerically to global optimality. The result is a cost-minimal combination of renewable, conventional, and storage capacity investments as well as their optimal hourly dispatch. Specifically, two strategies can increase the required minimum share of renewables: the use of storage to integrate surpluses, or larger renewable capacities plus curtailment. The model solves this trade-off endogenously. Note that, opposed to the myopic models in Section 4, this approach requires the assumption of perfect foresight to optimally schedule the release of energy from the storage.29 Intuition ~~~~~~~~~ To again provide some intuition before discussing numerical results, Fig. 10 plots the resulting RLDCs for the 80% renewables case. If capacity entry is costly with respect to both storage power and storage energy, the optimal solution combines both channels analyzed in Section 4. Areas D1 and D2 represent the curtailed renewable energy. The kink to the right is driven by power-oriented curtailment, compare Fig. 5. The renewable surplus gets curtailed beyond the threshold of 46 GW, as the shaded area D1 indicates, limited by the optimal storage power capacity. Curtailment of renewable surplus D2 is driven by energy-oriented renewable curtailment, compare Fig. 8, targeted at limiting the storage energy capacity. Area A represents the stored renewable surplus energy, which the storage shifts to hours with positive residual load. The optimal release of the stored renewable energy surpluses economically balances three uses: First, the energy shifted to area B1 reduces generation—and according variable costs—of the base technology, which otherwise operates whenever residual load is positive. Second, the energy shifted to area B2 reduces variable costs of the peak technology. Base capacity is 29.4 GW, indicated by the horizontal part of the dotted line; beyond, the peak plant with higher variable costs additionally generates electricity. Third, the energy shifted to area B3 replaces capacity entry—and respective cost—of the peak technology. If there was no storage, additional generation capacity would have to satisfy the demand exceeding 53.5 GW, indicated by the left-most horizontal part of the dotted line.30 Thus, the cost-minimization model illustrates two economic values of electrical storage beyond avoiding renewable curtailment in a concise way. First, an arbitrage value materializes when stored renewable surplus energy replaces variable costs of other generators, in particular fuel costs. Second, a capacity value materializes when stored renewable surplus energy replaces conventional power plant capacities that otherwise would have to be provided for hours of residual load peaks. Still, only renewable energy can enter the storage by assumption. Opening up the storage for conventional energy would strengthen both economic values. In reality, no reason prohibits such operation.31 Results ~~~~~~~ Fig. 11 summarizes the results of the parsimonious optimization model. Optimal storage capacities, both with respect to energy and power, rise with the share of variable renewable energy. However, overall storage requirements remain moderate. For instance, a storage energy capacity of 35 GWh suffices to achieve a share of 50% renewables in final demand. This is about two orders of magnitude less than in Sinn’s analysis, but slightly more than when targeting minimum storage requirements because the model considers additional values of storage. Yet it is still less than installed in Germany by 2018. Analogous findings prevail for other renewables shares.32 When increasing the targeted renewables share, the optimal storage energy capacity grows much faster than the optimal storage power capacity. For 50% renewables, 35 GWh energy are accompanied by about 6 GW storage power. Dividing energy by power yields an energy-to-power (E/P) ratio of about 6 hours. The E/P ratio is an important metric to characterize a storage technology and reflects its temporal layout: a 6 hours storage is a typical short-to-medium-term storage to compensate diurnal fluctuations, such as of solar PV generation. If it is completely charged, it can generate electricity for 6 hours at maximum power rating.33 For higher renewables shares, the E/P ratio increases to reach about 19 hours for 90% renewables. This highlights the importance of considering both rather inexpensive storage energy and rather expensive storage power separately. Moreover, no need for a true long-term storage arises, that is, storing energy for weeks or months. Optimal endogenous renewable curtailment also grows with higher minimum renewables shares. As such, the economics of renewable electricity provide no reason why curtailment should be avoided. It can be more efficient not to use available renewable energy at times despite costly investment into wind and PV plants. The optimal solution combines conventional plants, storage, and renewables, part of which being curtailed at times. Extension: flexible sector coupling (power-to-x) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In future low-carbon energy systems, renewable electricity supply considered as surplus energy in the above framework is likely to be highly valuable for new uses. Merging electricity, heating, and transport sectors can not only provide flexibility for integrating variable renewables into the power market, but can also contribute to decarbonizing these other sectors (Mathiesen et al., 2015). This concept, often referred to as sector coupling, comprises using renewable electricity, for example, for residential heating (see Bloess et al., 2018, for an overview of the recent literature) or for electric mobility (Richardson, 2013). Moreover, renewable electricity can also be used to produce other energy carriers, such as hydrogen or synthetic gaseous or liquid fuels, by means of electrolysis (Schiebahn et al., 2015). Such sector coupling options are often referred to as power-to-x (or P2X). They can lower electrical storage requirements if the new loads are sufficiently flexible and the additional demand can be shifted to periods in which renewable availability is high. Fig. 12 summarizes the results. For most renewables shares, both optimal storage capacities and renewable curtailment rates are substantially lower in case power-to-x technologies are included. For instance, for 50% renewables, the storage energy capacity drops from 35 GWh to 4 GWh and renewable curtailment from 5% to below 1%. Electrical storage requirements are lower up to a renewables share of 85%. In the 90% renewable case, storage needs are the same as in the case without power-to-x, as the additional demand can be completely satisfied from renewable surplus generation that would otherwise be curtailed. The rationale is the following: in the 50% renewables case, wind and solar capacities rise from 223 GW to 298 GW to supply part of the additional power-to-x demand. Another part of the additional demand is satisfied by renewable electricity previously curtailed or stored. Accordingly, both storage needs and renewable curtailment rates are substantially lower in most cases.35 As we assume power-to-x to be perfectly flexible, the diminishing effects on storage requirements and renewable curtailment constitute an upper bound for less flexible real-world applications. However, they illustrate that flexible sector coupling could substantially drive down electrical storage needs. If additional flexible electricity demand emerges that contributes to decarbonizing other energy sectors, then renewable surplus generation becomes a valuable resource.36","Departing from Sinn (2017), we implement small but relevant changes to move his the setup away from corner solutions. The impact is substantial and reduces storage needs up to two orders of magnitude. At the same time, our analysis remains stylized and tractable. In the following, we highlight further important points that researchers should consider when analyzing storage needs to integrate variable renewable electricity, both against the background of Sinn’s analysis and the large body of academic literature (Brown et al., 2018). First, the definition of efficiency should be clarified when it comes to variable renewable energy sources. Sinn (2017) seems to refer to inefficiency as both curtailment of renewables and storage use as such; the first motivated by avoiding waste, the second by avoiding backup capacities (referred to as “double structure buffering”).37 However, a welfare economic approach should rather target the least-cost provision of electricity for given minimum renewable energy constraints. The result is a combination of conventional and renewable plants, storage, and curtailment of a certain amount of renewable energy. Why the one or the other should be “inefficient” is unclear. In the optimum, the marginal cost of further expanding storage, renewables that are curtailed at times, and conventional capacities is equal. Which shares of renewable energy are optimal when also considering external costs, for instance, arising from climate change, local emissions or land use change, whether the market achieves this solution, or which regulatory measures are required is another issue left for analysis and discussion elsewhere. Second, electrical storage has values beyond arbitrage. It can provide firm capacity and thus reduce the need for backup plants (see Section 5.2), provide balancing reserves and other ancillary services to maintain power system stability, and may also help mitigating grid constraints. Third, other types of energy storage are relevant beyond pumped-hydro. These comprise batteries, which could become much cheaper in the future (Schmidt et al., 2017),38 or power-to-gas storage. Specifically, different storage technologies have different costs for power and energy capacities. While batteries are relatively cheap in power, they are expensive in energy, and vice versa for power-to-gas. Thus, electrical storage technologies have different optimal E/P ratios: batteries are generally suited for short-term storage of a few hours, pumped hydro for around six to ten hours, and power-to- gas for longer periods. An optimal deployment of different storage types can address different types of renewable fluctuations, for example, intra-hourly, diurnal or even seasonal (Safaei and Keith, 2015; Scholz et al., 2017; Zerrahn and Schill, 2017). Such a differentiated storage fleet also tends to be smaller and cheaper. Fourth, scaling up historical feed-in time series of a fixed proportion of wind and solar power tends to over-estimate flexibility requirements. Both market forces and the renewables support scheme in Germany tend to incentivize renewable generation when prices are higher and supply is, accordingly, scarce. Such system-friendly renewables comprise wind turbines that dis-proportionally produce electricity when wind speeds are low, both due to their location and technical layout (May, 2017). A similar argument holds for solar PV panels, which may be directed such that they generate more electricity in morning or afternoon hours. Even more relevant, offshore wind turbines have much smoother generation patterns and higher full-load hours than onshore wind parks.39 The sensitivity toward the base year in Section 3 further illustrates the relevance of appropriate input data choices. Ideally, analyses should be based on bottom-up weather data covering as many years as possible (Staffell and Pfenninger, 2016). Fifth, further electricity market integration across national borders generally yields smoother residual load. When balanced over greater geographical areas, the variability of wind, solar, and demand tends to be evened out (Cebulla et al., 2017; Fürsch et al., 2013; Haller et al., 2012; MacDonald et al., 2016). This results in smoother residual load patterns and lower storage requirements. Finally, a temporally more flexible demand can also substitute electrical storage (Denholm and Hand, 2011; Pape et al., 2014; Schill and Zerrahn, 2018). If a more variable electricity supply triggers more volatile prices, and these prices are passed through to consumers, then the demand side should have increasing incentives to consume more flexibly and profit from arbitrage gains.","The use of renewable energy is a major strategy to mitigate greenhouse gas emissions, reduce fossil fuel imports, and create a sustainable energy system. However, integrating growing shares of variable wind and solar power in electricity markets poses increasing challenges. Electrical storage is an important—albeit not the only—option to address the mismatching time profiles of variable renewable supply and electric demand. In a recent analysis, Sinn (2017) calculates storage needs in a German setting and finds vastly growing electrical storage requirements, already for renewable supply shares only moderately greater than currently the case in Germany. Based on these findings, he suggests that electrical storage may limit the further expansion of variable renewable energy sources. While Sinn’s illustrations deserve merit, the findings are not backed up by the literature. A large body of techno-economic studies conclude on substantially lower storage needs, also for high shares of variable renewables. An important reason for Sinn’s deviating findings is that he only considers corner solutions—either no storage, resulting in vast renewable curtailment, or no curtailment, resulting in excessive storage requirements. We show that addressing these implicit assumptions matters: both results and conclusions change substantially. Our analysis, based on open-source tools and open data, concludes that storage needs are lower by up to two orders of magnitude. Our findings are in line with most of the literature. Cost-efficient solutions optimally combine renewable capacity expansion, renewable curtailment, and electrical storage. We also illustrate that electrical storage needs may decrease further if the electricity sector is broadened to also include flexible additional demand, for example related to heating, mobility or hydrogen production. While we demonstrate that such power-to-x options may substantially change the picture, further and more detailed research on this avenue would be desirable. All things considered, we conclude that electrical storage requirements do not limit the further expansion of variable renewable energy sources."],["We experimentally investigate the relationship between discriminatory behaviour and the perceived social inappropriateness of discrimination. We conjecture that discrimination will be weaker when social norms oppose it. Our results support this prediction. Using a Krupka-Weber social norm elicitation task, we find participants perceive it to be more socially inappropriate to discriminate on the basis of nationality than on the basis of social identities artificially induced using a trivial minimal group technique. Correspondingly, we find that participants discriminate more in the artificial identity setting. Our results suggest norms and the preference to comply with them affect discriminatory decisions and that the social inappropriateness of discrimination moderates discriminatory behaviour. --------------------------------------------------------------------------------","Economic theories seeking to explain discrimination focus on two mechanisms. First, in the presence of incomplete information, profit- or income-maximizing agents use aggregate group characteristics to form statistical beliefs about individual characteristics and then act in accordance with those beliefs by, potentially, treating members of different groups differentially (Arrow, 1972). Second, individuals are assumed to derive direct utility from favouring certain groups relative to others, i.e. they are assumed to have a ‘taste for discrimination’ (Becker, 1957). Such tastes explain why discrimination is observed even in settings where asymmetric or incomplete information is not an issue (e.g. Chen and Li, 2009; Abbink and Harris, 2012). The focus of our paper is on this second form of discrimination, taste-based discrimination, and in particular on the psychological foundations of the tastes or preferences for discrimination, which have received remarkably little attention in the literature. Specifically, in this paper we use experimental methods to investigate whether tastes for discrimination are systematically associated with social norms, i.e. collectively recognised rules of behaviour that define which actions are viewed as socially appropriate within a specific social group.1 As we discuss further below, there may be a host of factors that shape the tastes for discrimination, including direct altruism towards members of one's own social group. The key contribution of our paper is to provide evidence that one important taste-shaping factor is a norm-based mechanism that regulates the extent to which actions that favour one's own group relative to others are regarded as permissible and appropriate. Uncovering this normative component is an important step towards understanding how patterns of taste- based discrimination are shaped. If social norms moderate the taste for discrimination, the incidence of discriminatory behaviour should positively correlate with beliefs about the appropriateness of discrimination. Similar correlations have been found in relation to other types of economic behaviour. Following Krupka and Weber (2013), lab and lab-in-the- field experiments have shown that in a variety of economic contexts people are more likely to take an action the more socially appropriate they perceive it to be (e.g. Burks and Krupka, 2012 – corporate ethics; Gächter et al., 2013 – gift-exchange; Krupka et al., 2016 – informal contract enforcement; Banerjee, 2016 – bribery). There is also evidence from econometric research (e.g. Buonanno et al., 2009) and natural field experiments (e.g. Allcott, 2011) suggesting norms drive behaviour outside the lab. Thus, in driving behaviour, social norms may effectively substitute for laws (e.g. Huang and Wu, 1994), or may complement them (e.g. Sunstein, 1990; Kübler, 2001; Lazzarini et al., 2004; Posner, 2009; Benabou and Tirole, 2011). However, a correlation between individuals' beliefs about the appropriateness of discrimination and the prevalence of discriminatory behaviour is a challenge to document empirically using naturally occurring data, not least of all because of the difficulties associated with accurately measuring such beliefs.2 Occasionally, attitudinal surveys include questions that can be interpreted as eliciting respondents' perceptions of the appropriateness of discrimination. For instance, the 2002 wave of the Scottish Social Attitudes Survey asked respondents whether they believed that ‘sometimes there is good reason for people to be prejudiced against certain groups’. One can interpret responses to this question as a proxy for the perceived social appropriateness of discrimination. Using this interpretation, we calculated the percentage of residents in each local authority area of Scotland who agreed with the statement. For each area, Fig. 1 plots this variable against the number of racist incidents,3 per 100 non-white residents,4 reported to the police in the financial year 2003–4 (Scottish Executive Statistical Bulletin, 2007). A correlation coefficient of 0.27 between the two variables suggests a positive relationship between the social appropriateness of racial discrimination and the incidence of racially discriminatory behaviour, which is consistent with the notion that norms moderate the taste of discrimination. The acceptability of prejudice-based humour has sometimes been used as a proxy for the normative appropriateness of discrimination (see, e.g., Crandall et al., 2002). Fig. 2 plots, over the period 2004 to 2014, the frequency of Google searches in the US for ‘N***** jokes’ (we apply the censorship for this paper; the original search term was uncensored5), as a proportion of all Google searches in the US (Google Trends, 2016). Searching for racist jokes about black people can be treated as evidence that the searcher perceives discrimination against black people to be socially appropriate. Fig. 2 also plots, on an annual basis over the same period, the number of incidents in the US involving hate crimes motivated by an anti-black bias that were reported to the FBI, per every 100 people living in areas where the hate crimes are reported (United States Department of Justice, 2015).6 Both the frequency of anti- black joke searches and the rate of anti-black hate crime incidents declined considerably over the period. This is suggestive of a positive relationship in the US between the change over time in the social appropriateness of discrimination against black people and the change over time in discriminatory behaviour against black people. In spite of these examples, the paucity of useful naturally occurring data with which to investigate the empirical relevance of norms for discriminatory behaviour advances the case for using experimental methods to address the question. Our paper does this, with an empirical strategy relying on four main elements. First, we use standard experimental techniques to prime participants to think about particular dimensions of their identities. The priming aims to trigger a process of social identification by encouraging subjects to identify with half of the participants in their experimental session and not with the other half. Second, in the decision-making phase of the experiment we ask subjects to distribute a given amount of money between two potential recipients, one an individual sharing their primed identity (‘in-group’), the other an individual not sharing their primed identity (‘out-group’). This simple allocation task allows us to measure discrimination as the extent to which individuals are willing to favour members of their own social group at the expense of the out-group. Third, crucially, we exogenously vary the dimension of identity that is primed. We do this across two treatments that we designed to vary the perceived appropriateness of discriminating in favour of the in-group and against the out-group, while holding other aspects of the decision-making context constant.7 Under one treatment, social identities are based on nationality; we form groups in the laboratory based on whether participants are British or Chinese. Under the other treatment, social identities are entirely artificial; groups are formed according to the colour of ball that each participant draws blindly from a bag. We expect the norms that mandate how a decision- maker should treat in-groups and out-groups in our experiment to differ across the two treatments. Specifically, we expect discrimination against out-group and in favour of in- group members to be perceived as less appropriate when identity groups are formed on the basis of nationality than when they are artificially formed on the basis of the colour of balls randomly picked. Indeed, when identity groups are artificially formed, participants have no directly relevant social norm to which to refer for guidance about the social appropriateness of discrimination. If this is the case, our exogenous manipulation varies the strength of the norm relating to discrimination across our treatments and, if discrimination is systematically shaped by norms, we thus expect discrimination to be stronger between the artificial groups. Fourth, as well as measuring discrimination, we directly measure the perceived social appropriateness of discrimination in each treatment. We do this by employing the ‘norm-elicitation’ task introduced by Krupka and Weber (2013), in which participants are described the allocator game and are asked to evaluate the social appropriateness of each and every possible action available to the allocator. We use this norm-elicitation task to construct an incentivized measure of the extent to which participants' perceptions of the appropriateness of discrimination vary across our two treatments and to examine the extent to which these differences in perceived appropriateness translate into differences in discriminatory behaviour in the allocation task. Our results show that, in both treatments, discriminatory actions are viewed as socially inappropriate. However, as expected, discrimination is perceived to be significantly less appropriate in the nationality treatment compared to the artificial identity treatment. The results of the decision task correlate with these differences in perceived appropriateness: while few participants discriminate in either treatment, discrimination is significantly stronger between artificial groups than between nationality groups. These results are consistent with the notion that the perceived social appropriateness of discrimination varies according to the way identity groups are defined, and this corresponds with individuals' revealed preferences for discrimination. That discrimination can be observed along a trivial, artificially-induced dimension of identity highlights the strength of the human inclination to discriminate against out-group members, and the ease with which in-group bias can be triggered (Ashburn-Nardo et al., 2001). That we observe weaker discrimination when identity is based upon the more meaningful characteristic of nationality, and that such discrimination is perceived to be more socially inappropriate, suggests that the extent to which human society has been effective in curbing the inclination to discriminate is owing to the development of shared norms proscribing this behaviour. Our study's main contribution is in linking discrimination to social norms and social identity theory. In this sense, our study is closely related to the paper by Chang et al. (2017), who investigate the effect of priming US citizens' political identities on redistributive behaviour. They show that individuals' primed political identities (Democrat or Republican) determine their perceptions of the social appropriateness of redistribution, and that this explains differences in redistributive behaviour between Democrats and Republicans. Like Chang et al.'s, our experiment shows that both individuals' distributive decisions and their perceptions of the social appropriateness of such decisions are sensitive to the dimension of identity that is salient in a given context. However, while the normative prescriptions upon which Chang et al. focus relate to the social identities of the decision-makers alone, we focus on the social identities of both the decision-makers and other individuals affected by the decision-makers' behaviour, and on how those social identities relate one to another. Thus, unlike Chang et al., in our experiment both the priming and the distributive decisions have an intergroup component which allows us to investigate the relationship between social identities, social norms, and discriminatory behaviour. Our paper is also related to work on the associations between social identity and norm enforcement.8 Bernhard et al. (2006) and Goette et al. (2006), for instance, use third-party punishment games to study whether the willingness to enforce norms of sharing and cooperation depends on the social identities of the norm violator and of the victim of the norm violation and on how those identities relate to that of the norm enforcer. Both papers find that social identity systematically affects the patterns of norm enforcement: enforcers are generally more willing to mete out punishment against violators when the victim of the norm violation is an in-group rather than an out-group member. Also related is Harris et al. (2014), who study whether in-group favouritism is proscribed by social norms by observing the extent to which individuals are willing to incur costs to punish it. They find that in-group favouritism goes largely unpunished when the punisher belongs to the same identity group as the norm violator or when she belongs to a neutral group. In-group favouritism is instead frequently punished when the punisher belongs to a different identity group. Harris et al. conclude that in-group favouritism is not always considered a violation of social norms, as this depends on the identities of the agents involved in the interaction. While these studies strongly suggest an association between discrimination and social norms and identities, none of them has directly measured the norms that underlie the observed patterns of behaviour. Moreover, none of these studies has investigated whether variations in primed social identity trigger differences in norms that, in turn, predict variations in discrimination. Thus, our study fills an important gap in this literature, as we are the first to provide direct evidence not only that discrimination co-varies with social norms, but also that these norms vary across particular dimensions of an individual's identity. The rest of the paper is set out as follows: Section 2 sketches a simple theoretical model of identity and norm-compliance that we use to motivate and inform our empirical strategy. Section 3 outlines our experimental design; Section 4 presents our results; Section 5 concludes and discusses our findings.","We assume that the decision-maker's utility can be broken into three components. The first component, Vi(a), describes individual i's utility over material payoffs, which in turn depends upon his or her own actions and the actions of others. Note that this accommodates standard self-regarding preferences, where the individual only cares about his or her own material payoff, as well as various forms of outcome-based other-regarding preferences, where individual i's utility also depends on others' material payoffs (e.g. Fehr and Schmidt, 1999; Bolton and Ockenfels, 2000). Importantly, this component of utility does not depend on the identities of the decision-maker or the others. In contrast, we assume that the second and third components of utility depend on social identities. These capture the decision-maker's willingness to treat others differently depending on how those others' identities compare to his or her own identity. There are several psychological mechanisms that form the basis for these components. Social identity theory (Tajfel and Turner, 1979), for example, posits that discrimination helps individuals satisfy their need for positive self-esteem since it confers a relatively high status on the in-group at the expense of the out-group. Subjective uncertainty reduction theory (Hogg, 2000) takes a different approach: individuals strive to reduce uncertainty about their attitudes, beliefs, and perceptions. Self-categorization and identification with groups that provide normative prescriptions for behaviour can reduce this uncertainty and lead to differential treatment of in-groups and out-groups. We view these psychological mechanisms as distal motivations for the ‘taste for discrimination’ that has been discussed in the economics literature. In our model we operationalise these mechanisms using two distinct components of utility. The component Si(a| I) captures what has traditionally been thought of as the taste for discrimination, i.e. utility derived by individual i from i's and others' material payoffs that is conditional on how the social identities of individual i and the others relate, one to another. In our model, this component of utility can be thought of as primal – as the direct utility that is or would be derived from favouring the in-group in the absence of any self-moderation. As in the utility functions proposed by McLeish and Oxoby (2007), Chen and Li (2009), and Chen and Chen (2011), this component of utility is not conditional on which specific dimension of identity is salient in the decision-making environment. Rather, we simply assume that i places a higher weight on the material payoffs of those players who are in-group, i.e., have the same social identity as him- or herself, as compared to the payoffs of players who are out-group, i.e., have a different identity.9 This can accommodate simple forms of favouritism towards the in-group, such as in-group altruism, as well as more complex forms of identity-contingent other-regarding preferences, as in the models by Chen and Li (2009) and Chen and Chen (2011). The third component of utility in our model, γiN(ai| I, d), captures the decision-maker's preference to self-moderate his or her primal inclination to favour the in-group with reference to what is or is not socially appropriate. Specifically, we assume that the individual derives utility from complying with normative prescriptions, captured by the function N(.), which defines the social appropriateness of each action ai available to individual i. These normative prescriptions depend on social identities. They may, for example, prescribe different behaviours towards in-group and out-group others. In addition, these prescriptions depend on the dimension of social identity, d, that is salient given the decision-making context. So, the same action may be viewed as more or less socially appropriate depending not only on how the identities of the decision-maker and others compare, but also on what dimension of identity it is appropriate or meaningful to compare given the context. In some contexts, this third term mitigates the second. For example, an individual might have a primal desire to direct mildly insulting comments at members of an ethnic group other than his or her own, but refrains from doing so because it is socially inappropriate. In other contexts, this third term may build on the second. For example, an individual might have a primal desire to direct mildly insulting comments at the supporters of a soccer team other than the one he or she supports and is further motivated to do so because, especially on match days, such behaviour is socially appropriate. Finally, γi is an individual-specific parameter defining the importance that individual i attaches to complying with social norms. In our experiment, subjects face a simple allocation task (described in detail in the next section), where they have to divide an amount of money between two other participants. In all treatments of the experiment, we keep constant the set of material payoffs available to players and the mapping from actions into payoffs. Thus, the first component Vi(a) of the utility function above is held constant across treatments for any given set of actions a. Moreover, in all treatments subjects are asked to divide the money between a participant who belongs to the same identity group as themselves and a participant who belongs to a different identity group. Thus, the second component Si(a| I) of the utility function is also kept constant across treatments. Our treatments vary the dimension of identity d that is made salient to the decision-makers and, hence, the process by which the relevant identity groups are defined in the experiment. As we describe in detail in the next section, in one treatment identity groups are formed on the basis of a random event, while in the other treatment identity groups are based on a meaningful personal characteristic. An implication of this treatment manipulation is that the normative prescriptions, N(ai| I, d), that regulate the third component of the utility function described above, may differ across treatments. Specifically, the same action ai available to the decision-maker may be evaluated differently depending on how identity groups are formed. Our experiment empirically explores the effect of varying the salient dimension of identity on the normative prescriptions relating to discriminatory behaviour and the role of these normative prescriptions in predicting such behaviour. Note that our model does not specify ex-ante the underlying determinants of the perceptions of appropriate behaviour captured by the function N(.), or how these will vary across treatments. Instead, we follow Krupka and Weber (2013) and Chang et al. (2017) and employ a norm-elicitation technique to quantify, in an incentive-compatible way, the function N(.) in each treatment.10 This allows us to assess empirically the extent to which normative prescriptions do indeed differ across treatments; and thereby examine the extent to which differences between treatments in the level of discrimination in the allocation task are predicted by differences in the perception of its appropriateness. Measuring discrimination – the allocator game ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In the allocator game, one participant was endowed with £16 and asked to allocate it between two passive players, one belonging to his or her own identity group and the other belonging to a different identity group.11 The decision-maker could not keep any of the money for him- or herself but knew he or she would receive a payment, between £6 and £10, which the computer would randomly pick at the end of the experiment.12 Allocators could split the money any way they liked between the other two players, as long as each amount was a multiple of two. Thus, the allocator had to choose one of nine possible allocations of money between the two passive players, ranging from (£16; £0) to (£0; £16). In order to maximize sample sizes, we elicited decisions using a role randomisation method: all participants were asked to make a decision in the allocator role knowing that their actual role would be determined at random at the end of the experiment (participants had a one- third chance of being assigned the allocator role and a two-thirds chance of being assigned a passive player role). Role assignment was implemented at the end of experiment, once everyone had submitted an allocation decision. Decisions were made anonymously and the only information allocators had about their recipients was the identity group that each of them belonged to. We chose the allocator game as our discrimination-eliciting device for the following reasons. First, given our focus on taste-based discrimination, we wanted a decision-making task within which statistical discrimination had no relevance; in the allocator game the decision-maker's material payoff does not depend on what any other player does, so statistical beliefs about other players are irrelevant.13 Second, to maximize our chances of discerning treatment differences, we wanted a task that reliably produces discriminatory behaviour and, in a meta-analysis, Lane (2016) found the allocator game to be the experimental task that yielded the strongest discrimination. Finally, in the allocator game it is obvious to participants what the experiment is about and any observed discrimination is interpretable as conscious rather than subconscious. Thus, the game is an ideal subject for a norm-elicitation task; it is much simpler to assess the social appropriateness of conscious behaviour than of subconscious behaviour. Measuring the social appropriateness of discrimination – the Krupka-Weber norm-elicitation task ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We elicited the social appropriateness of discrimination in the allocator game using an adaptation of the task design pioneered by Krupka and Weber (2013). Participants were described the allocator game, were presented with a table listing the nine possible actions an allocator could take, and were asked to evaluate the social appropriateness of each by selecting one option on a four-point scale: ‘Very socially inappropriate’, ‘Somewhat socially inappropriate’, ‘Somewhat socially appropriate’ or ‘Very socially appropriate.’ To ensure that the relevant perceptions of appropriateness are measured, the evaluators should be, to the greatest extent possible, in the mind-set of the person making the decision they are evaluating. In our experiment, in contrast to the original Krupka and Weber method, participants in the norm-elicitation task were the same as those playing the allocator game. This allows us to look at within-individual correlations between norms and actions. To facilitate an investigation into whether this had implications for either the behavioural or normative data, we varied which task came first (participants were unaware of the content of the second task until they had completed the first).14 All participants were assigned to identity groups before their first task, so those taking the norm-elicitation task first had had their identities primed in exactly the same way as the allocator game participants whose behaviour they were evaluating. Each individual in the norm-elicitation task only evaluated the appropriateness of actions made by allocators of the same identity group. The evaluation of actions was incentivised. Participants were told that, at the end of the experiment, one of the nine actions they had evaluated would be randomly selected, and each participant's evaluation of the action would be compared to that of another randomly selected participant. If a participant's evaluation matched that of the person they were compared with, that participant would earn £8; otherwise they would earn nothing. These incentives transform the task into a coordination game, where participants are incentivised to match other participants' evaluations of appropriateness. Krupka and Weber (2013) argue that this gives participants an incentive to reveal their perception of what is commonly regarded as appropriate or inappropriate behaviour in the decision situation, rather than their own personal evaluation of the actions they are asked to consider. This is important because social norms are collectively recognised rules of behaviour, rather than personal opinions about behaviours (e.g. Elster, 1989; Ostrom, 2000). Moreover, because we wanted to incentivise participants to coordinate on identity-specific social norms (i.e. the social norms that were recognised by those belonging to a specific identity group), participants were told that the person whose evaluation theirs would be compared to would be a member of their own identity group. Participants were told: ‘By socially appropriate, we mean behaviour that you think most participants [of your group] would agree is the “correct” thing to do. Another way to think about what we mean is that if [the allocator] were to select a socially inappropriate action, then another participant [of your group] might be angry at [the allocator].’ Treatments ~~~~~~~~~~ Our treatments, labelled Nationality and Artificial, differed in the way identity groups were formed. In Nationality participants in the experiment were segregated into identity groups based on nationality (previous economics studies taking this approach include Netzer and Sutter, 2009; Guillen and Ji, 2011; Goerg et al., 2016). In Artificial participants were split into ‘minimal groups’, using a variant of the technique first introduced by Tajfel et al. (1971), wherein social identities are artificially instilled in participants during the experiment. For both treatments we recruited British and Chinese students at the UK campus of the University of Nottingham, a British institution which hosts a large number of students from China.15 In the Nationality treatment, upon arrival, the British were seated on one side of the lab and the Chinese on the other. At every computer terminal on the British (Chinese) side was placed a sign reading ‘YOU ARE ON THE BRITISH (CHINESE) SIDE OF THE ROOM. ALL PARTICIPANTS ON THIS SIDE OF THE ROOM ARE BRITISH (CHINESE)’ (see Supplementary Online Materials B). In the instructions at the beginning of the experiment, it was again made explicitly clear that the lab and the participants had been divided based on nationality. In the Artificial treatment, upon arrival, participants blindly drew a ball from a bag. In each session the bag initially contained equal numbers of green and yellow balls, and participants continued to draw from it until the bag was empty, thus ensuring an equal split of green and yellow balls drawn. Those with green balls were then seated on one side of the lab, and those with yellow on the other. Consistent with the Nationality treatment, signs were placed at each terminal, reading ‘YOU ARE ON THE GREEN (YELLOW) SIDE OF THE ROOM. ALL PARTICIPANTS ON THIS SIDE OF THE ROOM DREW A GREEN (YELLOW) BALL’, and it was again made explicit at the beginning of the instructions that the lab and the participants had been divided on the basis of ball colour. As in the Nationality treatment, we invited an equal mix of British and Chinese students to the Artificial sessions. This ensures comparability between the two treatments.16 We conjectured that the normative prescriptions regulating the third component of utility in Eq. (1) would differ across the two treatments. Specifically, we conjectured that favouring the in-group at the expense of the out-group would be viewed as less appropriate in the Nationality compared to the Artificial treatment. This specific conjecture was derived from the following assumptions and observations. First, we considered the Nationality treatment. In Britain, a liberal society with a long history of in-migration, it seemed reasonable to assume: first, the existence of a norm proscribing discrimination against people from nations other than one's own; and second, that both the British and the Chinese participants in our experiment recognised such a norm and its relevance under the Nationality treatment. Under these assumptions, the third component of utility in Eq. (1) is discrimination prohibiting. Second, we considered the Artificial treatment. In this case, there was no directly relevant social norm to which participants could refer. However, we conjectured that participants could have referred to one or more apparently partially relevant social prescriptions or norms.17 So for some, the specifics of the Artificial treatment could have brought to mind dimensions of identity such as nationality or ethnicity. However, it seemed reasonable to assume that for others it would have been more likely to invoke dimensions of identity such as sports fandom and team game-playing, across which discrimination is condoned. Thus, we conjectured that, under Artificial, to the extent that any normative prescription or prescriptions were regulating the third component of utility in Eq. (1), on average, they would be less discrimination prohibiting than those at work under Nationality. Finally, we noted that this conjecture is consistent with the fact that previous experiments priming national identity (e.g. Goerg et al., 2016; Netzer and Sutter, 2009; Willinger et al., 2003) have often not found significant discrimination, while experiments involving minimal group identity (e.g., Ahmed, 2007; Chen and Li, 2009; Hargreaves Heap and Zizzo, 2009) do so more frequently.18 Indeed, according to a recent meta-analysis by Lane (2016), on average, discrimination is significantly weaker in the former compared to the latter type of experiment. Procedure ~~~~~~~~~ All participants participated in both the allocator game and the norm-elicitation task, as well as completing a post-experimental questionnaire. In each session, everyone received payment either for the allocator game or for the norm-elicitation task, as determined by a coin toss at the end of the experiment. Participants also received a £4 show-up fee. The order in which the tasks were performed was randomised between sessions, so that we could check for ordering effects. We do not find such effects (see Supplementary Online Materials C for the analysis), which is consistent with the findings of Erkut et al. (2015) and D'Adda et al. (2016). Therefore, in the analysis below we pool across ordering conditions. All sessions had 24 participants – twelve belonging to each group – and were conducted in March or April 2015, using z-Tree (Fischbacher, 2007). We conducted ten sessions, with 120 participants participating in each treatment.19 Treatment differences – social norms ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We look first at the social appropriateness of discrimination in each treatment, as measured by the norm-elicitation task. Fig. 3 plots the mean appropriateness ratings assigned to each allocation in the Nationality and Artificial treatments. Following the approach of Krupka and Weber (2013), we assign evenly-spaced values of −1 for the rating ‘very socially inappropriate’, −0.33 for the rating ‘somewhat socially inappropriate’, 0.33 for the rating ‘somewhat socially appropriate’ and 1 for the rating ‘very socially appropriate.’ The table at the bottom of the figure displays the distribution of evaluations for each allocation in each treatment, and presents the results of randomisation tests on the treatment differences in mean ratings. Our results are corrected for the fact that we are performing multiple tests; applying the Benjamini- Hochberg False Discovery Rate method (Benjamini and Hochberg, 1995), we sort our p-values in ascending rank and multiply each by the number of separate tests being performed (in our case nine, one for each possible allocation) before dividing each by its rank – thus greater adjustments are made to smaller p-values.20 In each treatment the mean and modal evaluations follow the same general pattern. Participants tend to regard extreme discrimination against recipients belonging to either identity group to be very socially inappropriate, while the equal split is generally regarded as very socially appropriate. There is a lack of strong consensus on allocations mildly favouring members of one group or the other. This pattern is consistent with a social norm of equality. However, in both treatments the perceived social appropriateness decays faster as allocations move away from equality towards favouring the out-group member than when they move towards favouring the in-group member, indicating that social norms against discrimination are stronger when the victim is a member of one's own identity group.21 By design, any treatment differences in the ratings assigned to a given allocation can only be driven by contextual differences in the perceived appropriateness of discrimination. We observe subtle but significant treatment differences. Whereas 95% of participants in the Nationality treatment perceive the equal split to be very appropriate, the equivalent figure is only 84.2% in the Artificial treatment; mean ratings for the equal split are significantly higher in the Nationality treatment. Furthermore, as the allocations move away from the equal split towards favouring the in-group, the appropriateness ratings decline at a faster rate in the Nationality treatment than in the Artificial treatment. For the extreme (16,0) split, 92.5% of participants in the Nationality treatment opt for ‘very inappropriate’, while only 80.8% do so in the Artificial treatment. And while only 5% of participants rate the (16,0) allocation as socially appropriate in the Nationality treatment, 18% do so in the Artificial treatment. In fact, Fig. 3 shows that, for any in-group-favouring allocation, there are more participants in the Artificial than Nationality treatment who find discrimination to be socially appropriate.22 As a consequence, all in-group-favouring allocations are, on average, perceived to be more appropriate in the Artificial treatment, and the differences are statistically significant at the 5% level or better in three out of the four possible cases (the exception being the allocation 14, 2 for which the difference is significant at the 10% level). Moreover, the differences in perception of appropriateness of discrimination only pertain to in-group favouritism and not to any form of discrimination; Fig. 3 shows that, while out-group-favouring allocations are, on average, perceived to be slightly more appropriate in the Artificial treatment, only for the (6,10) allocation is the difference significant, and then only at the 10% level.23 Treatment differences – discrimination ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Fig. 4 presents the distribution of decisions made in the allocator game in each treatment. In the Nationality treatment, 83.3% of participants choose to allocate the money evenly between the in-group member and the out-group member. Only 69.2% of the participants in the Artificial treatment make this choice. The remainder of participants in each treatment discriminate against out-group members; no participant in either treatment allocates more money to the out-group member than the in-group member. 12.5% of participants in the Artificial treatment allocate all the money to the in-group member, while only 4.2% do so in the Nationality treatment. In the Nationality treatment, participants allocate an average of £8.67 to the in-group member and £7.33 to the out- group member, resulting in a mean difference of £1.33. In the Artificial treatment, participants allocate an average of £9.52 to the in-group member and £6.48 to the out- group member, resulting in a mean difference of £3.03. A randomisation test indicates that the mean difference in the Artificial treatment is significantly higher than that in the Nationality treatment (p = 0.007). This is consistent with the conjecture that discrimination is stronger in the treatment where it is perceived to be more socially appropriate. It suggests that norm-compliance moderates discriminatory behaviour.24 In Table 1, an OLS regression confirms that the treatment effect on discrimination is robust to the inclusion of various controls – such as age, gender, nationality and the extent to which participants understand the tasks.25,26 Econometric analysis of individual perceptions of social appropriateness and behaviour ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ So far we have analysed the link between behaviour and norms at the group level, by showing that there is more discrimination in the treatment where it is perceived as less socially inappropriate. We now exploit the within-subject nature of our experiment to extend the analysis to the individual level. Specifically, we investigate whether a model that incorporates a preference for norm compliance is better able to explain the behavioural regularities in our experiment than a model that does not incorporate such a factor. Following the theoretical framework introduced in Section 2, we assume that the utility that allocators derive from choosing allocation x depends on three components, defined respectively on material payoffs, identity-contingent preferences over material payoffs, and normative prescriptions. We assume that the first component depends on the squared difference between the material payoffs of the two passive players implied by allocation x. The parameter v captures the weight that the allocator places on the material payoff component of the utility function, regardless of the identities of the passive players: allocations that implement unequal payoffs carry the same weight to utility, regardless of whether the inequality favours the in-group or out-group. Thus, the parameter v simply captures (identity-blind) preferences associated with payoff inequality. In contrast, the parameters s and γ capture the weight that the allocator places on the components of utility that are contingent on the identities of the passive players. The parameter s captures simple in-group altruism: the allocator places an extra weight on the material payoff of the passive player who belongs to the same group as him- or herself (and zero weight on the payoff of the out-group). Finally, the parameter γ captures the weight that allocators place on (identity-related) normative prescriptions. Our objective, here, is to show that a norm-augmented model is better able to capture the data patterns observed in the experiment than a model which contains only the first two components of utility captured in Eq. (2) above. Thus, in Table 2 we report the output of two fixed-effects conditional logit models, each estimated using all of the allocation decisions and, in Model (B), all of the social appropriateness evaluations generated under either the Nationality or the Artificial treatment. In the first model we impose the restriction γ = 0 to the utility function in Eq. (2) and, thus, estimate a choice model where the decision-maker is purely concerned with payoff inequality and simple in-group altruism. In the second model this restriction is removed and utility is allowed to depend on payoff inequality, simple in-group altruism, and the individual's normative evaluation of the action under consideration. The significant negative estimates of v in both models indicate that actions which yield larger payoff inequalities are less likely to be chosen. The estimate of s is positive and significant in both models, indicating that allocations which favour the in-group are more likely to be chosen. Finally, the significant positive estimate of γ in model (B) indicates that an individual is more likely to choose actions he or she perceives to be more socially appropriate. The significant estimate of γ in a model that also includes the v and s parameters indicates that the normative component of the utility function can explain variation in choice behaviour that cannot be entirely captured by (identity-blind) inequality considerations combined with simple in-group altruism. This also explains why the Bayesian Information Criterion is significantly lower for model (B) than (A) (p < 0.001 on a likelihood-ratio test) indicating that the norm- augmented model fits the data significantly better than the model without norms. The reason why the norm-augmented model performs better is made clear in Fig. 5, in which the aggregate action choice rates predicted by each of the models are graphed next to the actual choice rates (as displayed in Fig. 4). The left-hand panel of Fig. 5 presents the choice rates predicted by model (A). The right-hand panel presents the choice rates predicted by model (B). For ease of comparison, actual choice rates are reproduced in both panels. In each panel, the predicted choice rates (striped bars) and actual choice rates (shaded bars) of the Nationality (Artificial) treatment are shown in dark (light) grey. Model (A), in which participants care only about inequality and in-group altruism, captures some important aspects of the choice data. In particular, the model predicts that deviations from equality are asymmetric across the choice space. That is, the probability of choosing an unequal and in-group-favouring allocation is predicted to be higher than that of choosing an allocation which creates the same payoff inequality but favours the out-group. This is what we observe in the actual choice data, as no-one chooses out-group- favouring allocations, while 24% of participants choose an in-group-favouring allocation. However, Model (A) fails to capture a second key feature of the choice data, the difference in allocations across treatments. In contrast, the norm-augmented model (B) predicts a lower probability of choosing the equal split allocation and higher probabilities of choosing in-group-favouring allocations in the Artificial compared to the Nationality treatment. This is in line with what we observed in the experiment. Moreover, although the model still assigns positive probabilities to out-group-favouring allocations, they are markedly lower than those predicted by Model (A). Before concluding, we need to investigate whether identity-contingent preferences over material payoffs (captured by the term s[∥Ii=Ijπj(ax) + ∥Ii=Ikπk(ax)] in Eq. (2)) depend on the dimension of social identity that is salient in the decision-making situation. We would expect to observe such a dependence if, for example, in-group altruism varies depending on how strongly individuals identify with the group and this, in turn, depends on the dimension of identity along which groups are defined. Here, as in Model (B) in Table 2, v is highly significant and negative and both γ and s are highly significant and positive, while the newly added σ is statistically insignificant.28 Thus, we cannot reject the null hypothesis that direct utility derived from in-group altruism is independent of the dimension of social identity that is salient in the decision-making situation, and retain Eq. (2) as our preferred specification of utility.","We show that discrimination is perceived to be socially inappropriate. However, the extent of this perceived inappropriateness depends on the identities upon which discrimination is based: when the identities are defined with reference to a brief, random event, discrimination in favour of the in-group is viewed as less inappropriate than when the identities are based on nationality. Furthermore, we show that discrimination in the allocator game is stronger in the setting where it is perceived to be less inappropriate, and that, at the individual level, perceived inappropriateness predicts actual behaviour. Our findings are supportive of a theoretical framework within which taste-based discrimination is partly driven by normative considerations about the appropriateness of discriminatory behaviour. We offer direct evidence that, across choice contexts that are otherwise identical, differences in the way identity groups are defined translate into differences in the perceived normative prescriptions, and corresponding differences in behaviour towards in-group and out-group members. These findings are in line with models of social identity that have emphasised the role of social norms, such as Akerlof and Kranton (2000, 2005). Consistent with longstanding results from the minimal group literature, our study shows how remarkably easy it is to trigger discrimination between groups whose identities are based on artificial, trivial characteristics. That we find weaker discrimination on the basis of more meaningful identity characteristics such as nationalities, and that discrimination is perceived to be more socially inappropriate in that setting, suggests that shared norms opposing discrimination help moderate this most natural of human inclinations. We remain agnostic as to exactly why these norms opposing discrimination are more strongly triggered by national identity than by minimal group identity. National groups are different from minimal groups in various ways, so there are multiple possible explanations. We conjecture that a norm proscribing discrimination on the basis of nationality is likely to have developed over time in a liberal society, with a long history of in-migration, such as Britain, perhaps enhanced by sensitivities about the negative historical effects of racism and xenophobia. For these reasons, grouping participants by nationality seems likely to invoke stronger norms against discrimination than if we grouped them by other types of natural identity, such as university affiliation (as implemented, for instance, in Ockenfels and Werner, 2014). In the case of minimal groups, while their members have no directly relevant norms to refer to, we argue that they are relatively likely to invoke types of identity such as sports fandom and team game-playing, across which discrimination is considered harmless. It is also possible that some participants under minimal group identity perceive that the experimenter intends for them to discriminate, and that this experimental demand leads them to perceive a social norm in favour of discrimination. These are matters for interesting future research. While further investigation is needed, our findings are consistent with and strongly suggest that shared norms opposing discrimination do, as one would expect, help moderate discrimination. One likely implication of this would be that if society allows such prohibitive social norms to be eroded by whatever means, discrimination will increase. This would be consistent with the recent co-emergence of a backlash in various western countries against ‘political correctness’, spurred on by leaders promoting nationalism and identity politics, and the apparent rise in hostility towards immigrants and ethnic minorities in these countries."],["Intellectual property accounts for a growing share of firms' assets. It is more mobile than other forms of capital, and could be used by firms to shift income offshore and to reduce their corporate income tax liability. We consider how influential corporate income taxes are in determining where firms choose to legally own intellectual property. We estimate a mixed (or random coefficients) logit model that incorporates important observed and unobserved heterogeneity in firms' location choices. We obtain estimates of the full set of location specific tax elasticities and conduct ex ante analysis of how the location of ownership of intellectual property will respond to changes in tax policy. We find that recent reforms that give preferential tax treatment to income arising from patents are likely to have significant effects on the location of ownership of new intellectual property, and could lead to substantial reductions in tax revenue. © 2014. --------------------------------------------------------------------------------","The growing importance of intellectual property as a factor in production,1 and concern that it is easier for firms to shift income from this source than it is from others, presents challenges for tax design. Firms can and do position their intellectual property with a view to reducing tax liabilities. However, despite these concerns, firms do not by and large locate the legal ownership of intellectual property in the lowest tax countries, and corporate income taxes still raise considerable amounts of revenue in most developed countries. In this paper we address the question of how influential corporate income taxes are in determining where firms choose to legally register ownership of an important form of intangible assets, patents. Our contribution is to extend the empirical literature on public policy and firm location choice by introducing new methods to this area of public economics. We estimate a mixed (or random coefficients) logit model that incorporates both observed and unobserved heterogeneity in firms' location choices (see inter alia, Berry et al. (1995, 2004), Nevo (2001) and Train (2003)). A key strength of this approach is that it allows us to compute own and cross tax elasticities across locations that reflect patterns of correlation in observed choices in the data, and therefore to capture more realistic substitution patterns than standard logit models. Our estimates allow us to conduct ex ante analysis of how the location of ownership of intellectual property will respond to changes in policy. We use our estimates to simulate responses to recent policy reforms that provide preferential tax treatment to income arising from patents. We find that these reforms are likely to have significant effects on the location of ownership of new intellectual property, and could lead to substantial reductions in tax revenue. Our estimates could be used to simulate a wide range of other counterfactual situations. We use comprehensive panel data on all patent applications made to the European Patent Office (EPO) by a large number of innovative European firms over 1985–2005. A patent is a legal document that grants a firm the exclusive rights to use or licence a novel technology for a specified period of time. A firm can register legal ownership of a patent in a subsidiary that is located in a country different to the firm's headquarters, different to the location where the underlying technology was created and different to the location where the intellectual property will be applied. Lipsey (2010) notes that, in multinational firms, intangible assets “have no clear geographical location, but only a nominal location determined by the parent company's tax or legal strategies.” For example, Fig. 1 shows the share of patent applications made by UK parent firms where the legal ownership is registered outside of the UK and in a separate place to where the underlying innovative activity occurred. This share has increased six-fold over the past two decades. The largest proportion has gone to countries that have a lower tax rate than the UK, but the amount going to countries with a higher tax rate has also increased. We model the impact of tax on where firms choose to locate the legal ownership of patents. Tax could influence this decision because the legal ownership of the patent will be one of the determinants of where the income derived from the patent is taxed. The profits earned from the exploitation of intellectual property will be the result of a number of activities, including the research and development (R&D) investment undertaken to create the new idea, the financing of this investment and the subsequent commercialisation. When these activities take place in multiple countries, as is often the case for multinational firms, the returns must be allocated to individual jurisdictions for tax purposes. Firms have an incentive to arrange their activities in such a way that, all else equal, profits accrue in the country in which they would pay the lowest tax. There are a number of strategies that can be used to achieve this. Such strategies commonly require that the income earned from exploiting intellectual property accrues outside of the country in which the underlying R&D took place. One way to achieve this is through contract R&D. For example, a subsidiary in a relatively low tax country may finance (and bear the risk for) R&D activities that are contracted to a related subsidiary in a higher tax country (possibly with the benefit of R&D tax incentives and access to high skills levels). The contract will specify the payment to be made for the R&D activities (commonly equal to the costs incurred plus an arm's length mark-up). Returns above this payment, either from using the technology directly or from licensing it, will accrue to the subsidiary that bore the financial risk. There is a tax advantage to this strategy if the true value of the R&D activities is less than the price paid for the contract R&D. A similar result may be achieved through the use of a cost sharing agreement that specifies how subsidiaries will share the costs, risks and returns associated with an R&D project. Such agreements may be designed such that the right to exploit and capture the returns from a technology accrues to a subsidiary in a low tax country. The strategies available to a firm depend on how the firm is organised and on the precise tax rules they are subject to (Finnerty et al. (2007)). Tax rules limit a firm's ability to manipulate where income arises for tax purposes. Shifting income typically requires that payments made to compensate the company that conducts the R&D, or royalties made for the use of a technology, are at preferential prices. There are transfer pricing rules that aim to enforce the principle that the prices of intra-firm transactions are set as if they had occurred between unrelated parties — this is the arm's length principle. However, these transactions often do not have market counterparts, which means that firms may have opportunities to set the prices of related transactions in such a way as to reduce tax liability.2 Tax rules, including those that dictate how a firm can allocate the returns to innovative activities, differ across European countries and are different to those faced by US multinationals. For example, countries differ on the acceptable methods used to calculate payments for contracted R&D services, and where there are cost sharing agreements, countries differ in the requirements over whether all subsidiaries involved in the agreement need be engaged in R&D (in contrast to the US, not all European countries allow holding companies in low tax locations to be part of cost sharing agreements). The corporate tax rate is likely to be an important determinant of the location in which a firm chooses to hold legal ownership of intellectual property. However, it is unlikely to be the only factor; we would not expect all intellectual property to be legally registered in the lowest tax countries. Indeed, legal ownership of patents is rarely in the set of small countries that are often considered to be tax havens. The patents that are legally owned in such countries accounted for fewer than 0.5% of all patent applications made to the European Patent Office over the period 2001–2005, and many of those are unrelated to European firms.3 This could be due, at least in part, to the operation of Controlled Foreign Company (CFC) regimes, which effectively seeks to tax income at the higher home country tax rate if it is deemed to be located in a low tax country for tax purposes. More generally, there may be characteristics of a location over and above its corporate tax rate that firms value. For example, the strength of intellectual property rights protection and market size might play a role, and, all else equal, firms may be more likely to co-locate ownership of intellectual property with associated real innovative activity due to externalities from co-location. There is likely to be a large degree of heterogeneity in how responsive firms are to tax when deciding where to locate the legal ownership of their intellectual property; a number of papers have emphasised the importance of incorporating heterogeneity in firms' decisions (Melitz (2003), Bernard et al. (2007a, 2007b), Krautheim and Schmidt- Eisenlohr (2011)). This heterogeneity will arise for a number of reasons, some of which relate to observable factors, and others that relate to factors unobserved by the econometrician. For example, firms are likely to be more sensitive to tax when choosing the location in which to legally own patents with a relatively high expected value (Becker and Fuest (2007), Bohm et al. (2012)). Firms are also likely to be differentially responsive to tax due to differences in their organisational structures. Their existing network of subsidiaries, the proficiency of their tax department and the tax strategies they are able to employ for managing income from intellectual property will play a role. Firms with headquarters in different countries might respond differently if countries differ in the stringency of their tax rules and in the effectiveness with which they are applied. Firms operating in some markets or using certain technologies might respond differently, because, for example, transfer pricing rules may be easier to circumvent for firms operating in markets where a high share of transactions are intra-firm meaning it is difficult for tax authorities to accurately assess what is a fair market price. Both firm size and industry have been highlighted as important in the context of firm decision making over how to organise offshore activities (Graham and Tucker (2006) and Desai et al. (2006)). Indeed, the value of a patent, the relative attractiveness of a location and a firm's strategies and organisational structures are likely to vary across industries and, within industries, across firms. Our work relates to several papers in the literature. Most closely related, Dischinger and Riedel (2011) and Karkinsky and Riedel (2012) estimate the relationship between corporate tax and, respectively, the quantity of intangible assets and the number of patent applications made by subsidiaries located in each of a number of European countries. Also related is Ernst and Spengel (2011) who estimate the impact of R&D tax incentives and corporate tax on patenting. In common with these papers, we are interested in the relationship between corporate tax and where firms choose to locate intellectual property. We extend this literature by estimating a choice model that allows us to compute the full set of own tax and cross tax elasticities and which allows us to carry out ex ante analysis of how location decisions will respond to potential policy changes. Our work is also related to Cohen (2012), which uses a discrete choice framework to study how the design of US state tax rules influence US firms' decisions over in which state to incorporate. There is a considerable literature in the Hall and Jorgenson (1967) tradition that considers the impact of taxes on production activity and on the location of R&D. Hines (1996, 1999) and Devereux (2006) provide surveys of the empirical literature. This literature finds that, despite the many factors that will influence a firm's location decision, tax exerts a significant effect on location choices. Hines and Jaffe (2001) show that tax affects the location of firms' innovative activities within US multinational groups. Most relevant for our analysis, previous work has highlighted the role that intangible assets play in allowing firms to organise their activities with a view to reducing their tax burden (Altshuler and Grubert (2006)). Empirical studies provide indirect evidence of tax avoidance by, for example showing that firms have relatively high profitability in low tax countries (Grubert and Mutti (1991), Hines and Rice (1994)) and that the share of royalty payments associated with low tax countries is higher than expected (Grubert and Mutti (2009)). Grubert (2003) formalises how intangible assets can be used to shift income and finds that about half of the income shifted from high-tax to low-tax countries by US manufacturing firms can be accounted for by income from R&D linked intangibles. The structure of the paper is as follows. In Section 2 we outline a model of a firm's decision over where to locate the legal ownership of a patent. In Section 3 we describe the data we use to estimate the model. Section 4 presents the estimated coefficients and the tax elasticities between locations. An example of how the model can be used to conduct policy simulations is given in Section 5, where we consider the impact of recent reforms that reduce the tax rate for income derived from patents. A final section summarises and concludes.","When a firm generates a new idea, it expects to earn a stream of income on the application of that idea in the future. Ideas will vary both in their expected values and in the number of patents they give rise to (some will lead to one patent and some will lead to many). A firm faces the decision over where to initially locate the legal ownership of each patent. It will make this decision based, in part, on the rate of tax that it expects to face on income generated by the use of the patent in the future. Unobserved attributes of ideas are likely to be crucial, and potentially could generate correlations in patent location decisions. The firm will also take account of other characteristics of locations that it may value, for example, whether the real innovative activity associated with that intellectual property is also located there, the potential size of the market (if it also expects to commercialise the idea in that location), intellectual property rights' protection, technological condition, and many other location specific factors, at least some of which are likely to be unobserved by the econometrician. The importance of these location characteristic are likely to vary across ideas. For example, high value ideas may be more tax sensitive and the importance of intellectual property protection may differ across industries. We develop a tractable empirical model that captures these determinants of location choice. Firm payoffs ~~~~~~~~~~~~ We specify a model in which a parent firm decides where to locate the legal ownership of each of its patents. Firms, indexed f = 1,…,F, realise ideas, indexed i = 1,…,I. Ideas are assumed to arise exogenously over time, indexed t. Each idea can yield a single patent or a group of related patents; patents are indexed p = 1,…,P. We model the country, indexed j = 1,…,J, in which the parent firm decides to locate the legal ownership of each patent, allowing for correlation in decisions between related patents (those that are part of the same idea). We consider all patents taken out by a parent firm that are technologically related in a quarter as part of the same idea; the precise definition of an idea is given in Section 1. For each patent, the parent firm chooses the location that yields the highest payoff. The payoff the parent firm gets from choosing a location depends on the tax rate it expects to face, τfjt, the quality of the idea, qi, whether any research activity that gave rise to the idea is located there, aijt, the strength of the country's intellectual property rights protection, the size of the local market (measured as GDP), and the level of technological innovativeness (measured as total annual business R&D expenditure as a share of GDP), captured in the vector xjt. Crucially, we also allow location choice to depend on unobserved characteristics of both the idea and the location. We allow the impact of all observed and unobserved factors to vary across medium and large firms and for technologies in different industries; the subscript r = 1,…,R denotes the industry–firm size category an idea belongs to. The tax rate τfjt varies across firms, because the tax system in a firm's residence jurisdiction may interact with the rules of the countries in which it is considering locating ownership of a patent through the operation of Controlled Foreign Company (CFC) rules. We use the tax rate dated t, making the assumption that when a firm chooses the location of a patent it expects that the current tax regime will apply in the future. The tax parameter, αi, is at the idea level. We allow it to vary with an observed measure of idea quality, qi. Patents that are part of a high quality idea are likely to have a higher expected value, and thus their location may be more sensitive to tax. We also allow the idea level tax parameter to include a random term ηi. This captures all components of ideas that determine the responsiveness of location choice to tax and are observed by the parent firm but not by the econometrician. For instance, the quality variable is likely to be an imperfect measure of the expected value of the idea. There could be other factors that are correlated with the idea's expected values, unobserved by us, but available to firms, which will be captured by ηi. Similarly, we model the parameter on real innovative activity, βi, as an idea level random coefficient; the impact of real innovative activity on the payoff function varies across patents with the random term νi. Firms may value locating the legal ownership of intellectual property in the same country as it was created, and the strength of this motive is likely to vary across ideas. A central assumption of the standard multinomial logit model is that the stochastic error term associated with the payoff from a particular option (in our case the decision to locate ownership of a patent in a particular location) has an iid type I extreme value distribution. This rules out correlation in latent payoffs. This assumption leads to a closed form solution for the location choice probabilities, which is empirically convenient. However, it is restrictive, leading the multinomial logit model to imply restrictive substitution patterns. In particular, the lack of correlation in payoffs endows the model with the independence of irrelevant alternatives (IIA) property. We include in xjt a number of time varying location characteristics that firms are likely to value when choosing where to locate the legal ownership of intellectual property. However, there are likely to be other location characteristics that firms value that we do not observe. To capture these we include location–industry–firm size fixed effects, ξrj. These will control for location specific costs, such as the legal costs associated with setting up a subsidiary, or location specific benefits, such as government provided public goods, the relevance of which might differ for firms of different size or industry. Note that there may be many other patent, idea or firm specific factors that do not vary across location but that do influence the costs or benefits of a location. However, because these do not vary across location they will not enter the location choice decision, and therefore are not explicitly entered into Eq. (1); including them would lead to an observationally equivalent choice model, because they would drop out when payoff comparisons are made across locations. Identification ~~~~~~~~~~~~~~ Our primary interest lies in pinning down the ceteris paribus impact of a change in the corporate tax rate set by any one country on the shares of patent applications made by subsidiaries in both that and in alternative locations. To do this we must consistently estimate the parameters of the payoff function outlined in Eqs. (1)–(4), and in particular the parameters governing the marginal impact of tax on the payoff associated with selecting a location, which are modelled as random coefficients. Train (2003) shows that the random coefficient model has a duel interpretation as an error component model. Under this representation, the mean of the random coefficient can be interpreted as a fixed coefficient — it is pinned down by variation in location choices in response to variation in taxes faced by firms, conditional on other observables included in the model. The standard deviation of the random coefficient is interpreted as a component in the error term — it is pinned down by correlation in payoffs across locations (both in a given choice set and across location decisions within a given idea). Berry and Haile (2010) have established that in random utility multinomial logit models the distribution of unobserved preference parameters is non-parametrically identified given sufficiently rich micro data. However, non-parametric estimation of this model is computationally burdensome, and we therefore follow most papers in the literature by assuming payoffs are linear with independent additive shocks and that the distribution of unobserved parameters is normally distributed. These assumptions give us a convenient approximation. The standard identification concerns still apply here; to consistently estimate the parameters we require that the additive shock (ϵpjt) and the idea specific random terms (ηi and νi) are independent from each other and from the other explanatory variables. Specifically, if there are factors which influence location choice, and that are not captured by the observed and unobserved controls that we include, this would lead to inconsistent estimates. To mitigate this concern, we include a number of controls in the model. We include location–industry–firm size fixed effects; these control for all country characteristics that affect a firm's payoff and that do not vary through time, but that potentially do vary across firms in different industries and different parent firm sizes. We include time varying (non-tax) location characteristics. These include a measure of the presence of real innovative activity associated with the intellectual property. This controls for the fact that some firms, for reasons other than the tax rate they will face in a particular location, may wish to co-locate legal ownership of intellectual property with real innovative activity. If decisions over the location of real innovative activity are influenced by corporate tax rates, then failure to control for this would result in an inconsistent estimate of the impact of tax on patent location choice. We also control for the strength of intellectual property rights protection in the location, market size and technological innovativeness — all factors that vary over time and location, and may be expected to impact on intellectual property location choices. Identification of the tax coefficients also relies on the presence of informative variation in taxes in the data; specifically we need to observe variation in the set of taxes in potential locations across patent choice situations, conditional on all other factors that influence location choice. Crucially, it is necessary that there is variation in differences in taxes between locations across choice sets. So, for instance, if the only source of tax variation was that all tax rates changed simultaneously by the same amount, the marginal effect of tax on the payoffs would not be identified. Such variation arises in our framework for two reasons. First, there is variation over time in statutory tax rates; as outlined below in Section 2, over the period for which we have data (1985–2005) there has been a general pattern of declining statutory tax rates. The size of this decline has varied across countries, and the changes have occurred at different times, meaning that tax reforms have given rise to variation in the set of tax rates across locations. Second, in addition to this time series variation, CFC regimes lead to another source of variation; two parent firms taking a decision at the same point in time, but resident in different countries, can face a different set of tax rates due to cross country differences in CFC rules. The variation in location choices, conditional on other factors, in response to this variation in tax rates pins down the impact of tax on location choice.","To estimate the model we need information on where firms have chosen to locate the legal ownership of their patents, the corporate tax regime and other conditioning variables. Patents data ~~~~~~~~~~~~ We use data on patent applications filed at the European Patent Office (EPO) by the European and US subsidiaries of parent firms located in fourteen European countries. We exclude from our analysis firms that patent infrequently. The number of patent applications by location of the subsidiary that filed the patent application is shown in Table 3.1. Our data include 1083 parent firms that collectively have 4,823 patenting subsidiaries, which file 379,849 patent applications over the period 1985–2005. These account for a 70% of all corporate applications filed at the EPO by firms parented in these fourteen European countries during this period. Each patent application lists the firm that files the application (the applicant), this is the legal owner of the patent.4 We identify the parent firm using information from company accounts (from Amadeus), company websites, business directories and other sources (see Abramovsky et al. (2008) for details). We use ownership information at a fixed point in time (2004), and we do not observe changes in ownership after an application has been filed. For each patent application this gives us a mapping between the location of the parent firm and the location of the subsidiary that legally owns the patent. We also observe the location of the inventors (individuals) that created the technology underlying the patent application.5 There are often inventors located in multiple countries, and often in different countries to that of the applicant.6 The location of both the applicant and the inventors are distinct from the patent office to which the firm is applying for protection. For patent applications filed at the EPO, each application also designates the individual countries in which final patent protection will be sought; it is these individual countries, not the EPO, that grants patent protection. We use Thomson's Derwent database to classify patent applications based on the technology embodied in the patent and the markets in which the technology is used. We use three broad industry groups — Chemical, Electrical and Engineering. A patent application can be relevant for more than one industry group if it has applicability in more than one of these industries. Where this is the case we include the patent application in each of the industry sub-samples, and when we calculate the market level elasticities we weight each patent application so that the sum of the weights equals one (so if a patent application is in two industries it will get a weight of 0.5 in each). Table 3.1 columns (2)–(4) shows the industry split of patent applications, 32.3% are in Chemical, 36.8% in Electrical and 30.8% in Engineering, but this varies across countries. We restrict our analysis to firms that are above the 20th percentile in terms of the number of patent applications per firm in their industry. We distinguish large firms as those with a level of patenting above the 80th percentile in their industry. 78.8% of patent applications in our sample are held by large firms. Firms often take out a number of related patent applications at the same time; we allow for correlation in these decisions. We group together related patent applications that can be considered to be part of the same idea. We identify patent applications as part of the same idea if they are made by the same parent firm, are filed in the same quarter (i.e. three month period), are classified in the same industry and share a network of common inventors. The number of patent applications in an idea varies: on average an idea contains one patent application; ideas containing more than one application account for 26% of all patent applications. The importance of ideas in our empirical strategy is that we allow correlation across patent applications at the level of the idea; the decisions over where to locate ownership of these related patent applications is unlikely to be independent, and the inclusion of random coefficients at this level allows for them to be correlated. Patent quality There is a large literature that highlights the skewness of patent value and quality (Pakes (1985, 1986), Blundell et al. (1999), Lanjouw and Schankerman (2004) and Hall et al. (2007)). Firms file patent applications for a variety of reasons; some are filed to protect valuable new ideas, others are filed strategically to provide option values or to block competitors. Patent value also varies because some ideas are commercially more valuable than others. We identify high quality patent applications as those that are part of a triadic patent family, i.e. a related patent application has been filed at each of the EPO, the US Patent and Trademark office and the Japan Patent Office. The OECD uses triadic patent families “… to improve the international comparability and quality of patent-based indicators … patents included in the triadic family are typically of higher economic value: patentees only take on the additional costs and delays of extending the protection of their invention to other countries if they deem it worthwhile.” (OECD, 2012). We expect triadic patent applications to be of a higher value since there is a cost to filing patent applications at each of these patent offices, and the main incentive to do this is if firms expect the technology to have a wide application. Each idea (group of patent applications) is classified as high quality if over half of the associated patent applications are triadic. As seen in Table 3.1 (column (7)) on average 36.5% of patent applications are classified as being part of a high quality idea. Patent ownership and income from patents We model the impact of tax on where firms choose to locate the legal ownership of patents. In the introduction we discuss the reasons that we might expect tax to affect a firm's decision of where to hold legal ownership of intellectual property. The extent to which firms have arranged their activities in such a way that income can reasonably be deemed to be attributable to the subsidiary that legally owns the intellectual property will differ; firms will differ in how aggressively they seek to manage their tax liabilities. For some firms, the choice of where to earn income may be a choice between those countries in which real innovative activities already takes place; others may employ strategies that allow them to earn income in a separate country. There are many factors that affect the costs and benefits of choosing a particular location. For tax havens these costs may be particularly high: CFC rules are more likely to bind; the transfer of profits to locations where there is little real activity will be more difficult; tax havens are likely to be less attractive locations along non-tax dimensions such as intellectual property rights protection. We would not expect legal ownership of all patents to be located in such countries. However, it is possible that some firms are particularly aggressive in their tax planning and organise their activities in such a way that income is earned in a location that is not where legal ownership is located. We do not observe income flows, so we do not explicitly model this behaviour; to the extent that it makes the decision over the location of legal ownership less related to tax we would be less likely to find an impact of tax. In our model we aim to capture this variation in behaviour across firms, and within firms across ideas, by the inclusion of observed and unobserved heterogeneity. An additional complicating factor is that a firm might file a patent application from one subsidiary, but later transfer ownership of that patent to another related firm. However, firms have an incentive to consider tax when making the initial location decision, because in many situations there are tax costs to transferring the ownership of intangible assets. For there to be a tax benefit to the sale or transfer of an asset it must be the case that this can happen at a value below the true market value. The transfer of intellectual property will be subject to transfer pricing rules, which will act to limit how much value can be shifted to a low tax country. In addition, many European countries operate exit taxes that attempt to levy tax on the net present value of the expected revenue stream on an intangible asset when it is moved out of the country. Such tax provisions reduce (if not remove) any tax advantages to re- locating to a lower tax jurisdiction. If firms do intend, with some positive probability, to re-locate the ownership of a patent in the future, and if transfer pricing rules and exit taxes do not act to perfectly off-set any tax advantages of doing so, this would reduce the importance of corporate tax in the initial location decision. This is an additional reason that it is important that we allow heterogeneity in the importance of corporate tax across intellectual property. The place where we need to make more restrictive assumptions about the relationship between legal ownership and income from intellectual property is when we carry out the ex ante analysis of the Patent Box tax reforms and calculate the revenue implications of these reforms. In order to do this we need to assume that the relationship between legal ownership and income is not changed by the policy reform. Taxes ~~~~~ We measure the impact of tax on payoffs using the statutory tax rate. We assume that returns from intellectual property are expected to be sufficiently high that deductions such as capital allowances are relatively unimportant, so that the effective tax rate faced by the firm is approximately the statutory tax rate (see Devereux and Griffith (2003), where Fig. 1 shows that the marginal effective tax rate asymptotes to the statutory tax rate as profitability increases). Our identification strategy relies on variation over time and across countries in the tax rate. Table 3.1 (columns (8)–(11)) summarise the variation in corporate tax rates. In general, main statutory tax rates fell in the two decades up to 2005, but with the timings of changes differing across countries. The Scandinavian countries – Denmark, Finland, Norway and Sweden – reduced tax rates significantly around 1990. Italy enacted a reduction of over 10 percentage points in 1998, as did Germany in 2001. France and the UK have enacted a series of gradual reductions. There can be additional tax levied in the parent firm's home country as a result of Controlled Foreign Company (CFC) rules, which aim to prevent firms locating income in lower tax countries in order to avoid taxation in their home country. CFC rules set out criteria for identifying subsidiaries that are located in a country deemed to be ‘low-tax’ and earning a significant amount of ‘passive income’ (income that is not associated with real activity). When a CFC regime is in place in a parent firm's country of residence, and a subsidiary is located in a country that is deemed a ‘low tax’ location (as judged against parent firm country specific thresholds), then we set the tax variable, τfjt, equal to the parent firm country's statutory rate. A description of the country pairs for which this is the case is given in Table 3.2. There is variation in whether a parent firm country operates a CFC regime (some regimes are introduced during the period for which we have data) and in the applicant countries that are deemed low tax (which differ over time when statutory rates change). This definition of whether CFC rules bind effectively assumes that the income received from a patent is deemed to be passive income, and that the share of passive income is sufficient to trigger the CFC rules. This is clearly an approximation. However, if we look across all location options that firms in our data face and that are deemed low tax by CFC regimes, then it is rarely the case that the parent firm has both inventors and holds legal ownership of a patent application in the same location. The results we present below are robust to the alternative assumption that patent applications with ownership located in countries where associated real innovative activity is also located would be treated as active income, so that CFC rules do not bind. Descriptive statistics ~~~~~~~~~~~~~~~~~~~~~~ The variables included in the model are defined and summarised in Table 3.3. The top panel contains the observed location attributes we include. These comprise the tax rate that the parent firm would face if it earned income from the application of intellectual property in the location, a measure of the presence of real innovative activity in a location defined as an indicator of whether at least one of the inventors associated with the patent applications that form the idea are located in that country and country-time varying observable characteristics. The latter includes a measure of intellectual property rights protection. This is based on a measure developed by Ginarte and Park (1997) and Park (2008). The countries we consider all have advanced systems of property rights and therefore rank relatively highly on the protection of intellectual property. We define a country as having a strong intellectual property regime if it scores above the median of countries in our sample. Other country-time varying variables include market size, as measured by Gross Domestic Product (GDP) and the technological innovativeness of a country, proxied by business R&D investment in the country as a share of GDP. We allow the valuations firms place on location characteristics to vary across patent applications. A summary of observable patent (or idea) characteristics is given in the bottom panel of Table 3.3. In estimation we allow all coefficients to vary with the industry the patent application belongs to and the size of the associated parent firm. This allows the model to capture, for example, that large firms are more likely to have organisational structures that assist the location of intellectual property for tax purposes. The tax rate is interacted with a measure of the idea quality, reflecting the possibility that firms' location choices may be more responsive to tax when they expect intellectual property to earn higher returns.","Table 4.1 shows the estimated coefficients of the choice model outlined in Section 2. The model is estimated using simulated maximum likelihood (see Train (2003)). We allow all coefficients to vary across industry and firms size, indicated by the different columns. We include a full set of location–industry–firm size fixed effects (not reported in Table 4.1, but available upon request). The top row of Table 4.1 shows that the mean marginal impact of tax on the payoff from placing legal ownership of a patent in a location is negative and statistically significant across all industries and parent firm size groups. The second row shows that in both the electrical and engineering industries the payoff for high quality patents is more sensitive to taxes. This is true both for large and medium firms. In the chemical industry the payoff for a high quality patent is estimated to be marginally less responsive to tax than for lower quality patents for large firms, with there being no statistically significant difference between the high and low quality patents for medium firms. Row three shows that there is a substantial degree of unobserved heterogeneity in the importance of tax on location choice across ideas, the standard deviations of the random coefficients on tax are both large and statistically significant across all industries and size categories. The fourth row shows that, ceteris paribus, having real innovative activity in a location is associated with a higher payoff from placing legal ownership of a patent in that location across all industries and size categories; the fifth row shows that there is a significant amount of variation in the importance of this characteristic across ideas. Together the large and statistically significant standard deviations on the random coefficients on tax and real innovative activity (in all industry-firm size groups) indicates the presence of important correlations in payoffs, both across locations for a given patent, and across patents in a given idea. These correlations will generate patterns of substitution that will depart from the more restrictive patterns implied by a standard multinomial logit model. The remaining three rows of Table 4.1 describe the impact of having strong intellectual property protection, and the marginal impacts of market size and technological innovativeness, on the payoff function. For five of the six industry-firm size groups, a location having strong intellectual property protection is, all else equal, associated with firms obtaining higher payoffs from locating legal ownership of their patents there (the exception is medium electric firms, for which the strong intellectual property rights dummy is negative). Larger market size is associated with statistically significantly larger payoffs for five of the size industry-firm size groups, and a higher degree of technological innovativeness is associated with statistically significantly larger payoffs for four of the size industry-firm size groups. Table 4.2 shows the matrix of own and cross tax elasticities implied by the choice model. It contains the elasticities of the share of patents located in each of 14 European countries with respect to the rate of corporate tax set in each of these countries and in the US. These are calculated as described in Section 2. We report the matrix of elasticities using tax rates and the distribution of patent applications for the most recent year in our data, 2005. Each cell shows the elasticity of the share of patents located in the country indicated in column 1 with respect to the tax rate set by the country in row 1. The emboldened diagonal shows the own tax elasticities. For all locations, except Luxembourg, the own tax elasticities are less than one in magnitude. There is a limited literature on the elasticity of the location of corporate income with respect to tax. De Mooij and Ederveen (2008) report that empirical studies considering the effect of differences in statutory tax rates on various measures of profitability (with a view to indirectly capturing the effects on profit shifting) tend to find a semi-elasticity of around − 1.2. As in this paper, Karkinsky and Riedel (2012) consider the link between corporate tax rates and patent applications. They estimate a semi-elasticity that, depending on the functional form of their model, implies that a 1 percentage point increase in the rate of corporate tax translates into a 3.5%–3.8% fall in patent applications from that location. Direct comparison with our results is made difficult by the fact that our model allows tax effects to vary across all locations. We find that the share of patents held in Luxembourg is most sensitive to tax (the Luxembourg semi-elasticity is 3.9%) and least sensitive for Germany (the German semi- elasticity is 0.5%).7 Theoretically, we might expect smaller countries to have relatively high own tax elasticities, as a change in their tax rate will not affect the market rate of return, making the cost of capital more responsive to tax changes (see Wilson (1999)). This may be one of the reasons that such countries are more likely to compete for corporate income using low rates; a change in the rate leads to a larger change in the relatively small tax base (see, for example, Bucovetsky and Haufler (2007)). The own tax elasticities in Table 4.2 show some evidence of this; they are higher for the Benelux countries than for France and Germany. The importance of allowing for observed heterogeneity and correlation in locations' payoffs can be seen by looking at the cross tax elasticities. In a multinomial logit with no observed or unobserved heterogeneity all cross tax elasticities in a column would be the same — a reduction in the tax rate in location A would lead to patent applications switching from other locations in proportion to their original shares. This implausibly restrictive pattern of substitution is not imposed in our more flexible model, meaning that elasticities vary substantially within a column. In particular, our model allows the data to capture the fact that firms are more likely to choose to switch between locations with similar characteristics (whether that be because the firm has inventors located in several locations, or because locations have similar tax rates).","One of the advantages of estimating the model outlined above is that it captures patterns of substitution across locations, and it therefore allows us to simulate counterfactual policy situations. We illustrate this by considering a recent set of policy reforms. A number of European countries have introduced polices that offer substantially reduced rates of corporation tax on the income derived from patents, and in some cases other forms of intellectual property (these are often called Patent Boxes). Firms are able to declare that some portion of their profits are derived from either the use or licence of patents, and these profits are taxed at a lower rate. Patent Box rules differ across countries, for example, in terms of how eligible income is measured, how the rules that apply when calculating how much income can be allocated to patents, and how the related expenses are treated.8 None of the countries require that the R&D underlying the intellectual property took place in that country, as this is not permissible under European law. We use the most recent year of our data (2005) to simulate the impact of the two sets of policies. First we consider the introduction of Patent Boxes in the Benelux countries, and second the later introduction in the UK. We simulate the impact of these policies on the share of new patents for which legal ownership is placed in each of these countries using the choice model presented above. For illustrative purposes, we assume that the total level of patenting activity by European firms is not affected by the policy reforms. We also consider the impact of these policy reforms on tax revenue; this requires the further assumption that the relationship between where tax is levied and the location of legal ownership is not altered by the policy reform. The policies are summarised in Table 5.1. In 2007 Belgium introduced a Patent Box that reduced the tax rate on income derived from patents from 34% to 6.8%, and the Netherlands introduced a Patent Box that reduced the rate from 31.5% to 10%. In 2008 Luxembourg reduced the rate from 30.4% to 5.9%. The UK government introduced a Patent Box at the rate of 10% in 2013; the main rate of corporate tax in the UK was 30% in 2005, but had fallen to 24% by 2013. We simulate the impact of the reduction from 30% to the Patent Box rate.9 Table 5.2 sets out the results of these simulations for the four locations that introduced Patent Boxes.10 A note of caution in interpreting these results is that the lowest tax rate we observe in the data is 10% in Ireland, whereas two of the Patent Box rates are below this level, and so are outside the observed range of taxes in our data. We carry out the simulation on the full set of patent applications, shown in the top panel. It may be the case that many patents do not earn much income, and so in the bottom panel we carry out the simulation using only the high quality patents, under the assumption that these are the patents that are expected to earn the highest income. The estimates suggest that the location of these patents were on average more sensitive to tax. The first column shows the actual share of patent applications in each location in 2005 (prior to the introduction of Patent Boxes). The second column shows the predicted share of patent applications in each location after the introduction of Benelux Patent Boxes. The standard error of these predicted shares are shown in parenthesis. The third column expresses the % change from column 1 to column 2. The introduction of Patent Boxes in the Benelux countries leads to a large and statistically significant increase in the share of new patents whose legal ownership is located in Belgium and Netherlands. The increase in Luxembourg is proportionally large, but is not statistically significant. There is no change in the share in the UK. There is a decline in the share of patent applications located in other non-Patent Box locations (not shown). The fourth column shows the predicted shares after the introduction of the UK Patent Box (in addition to the Benelux Patent Boxes). The fifth column shows the % changes from column 1 to column 4. The UK Patent Box leads to a reduction in the share of new patent applications made by subsidiaries located in the Benelux countries, but for Belgium and the Netherlands they still have a statistically significantly higher share than prior to the introduction of any Patent Boxes. The share of new patent applications made by subsidiaries located in the UK increases by a statistically and economically significant amount. The results with high quality patents are similar. In columns (6)–(8) of Table 5.2 we consider the impact on tax revenue from income derived from patents. These combine two effects. The reduction in the statutory tax rate will reduce revenue, but the increase in the share of income from patents will increase it. We demonstrate the impact on tax revenue by computing the product of the statutory tax rate in each country and the share of patent applications. We index this to 100 before the introduction of any Patent Boxes. In the upper panel of the table we assume that all patents are equally valuable, and that the relationship between legal ownership and taxable income is not affected by the reform. All countries experience a decline in revenue. Although the countries that introduce Patent Boxes attract more new patents, the increased share is not sufficient to outweigh the effect of the lower tax rate. With all four Patent Box policies in place, revenues are less than half of their previous levels in these countries. Ernst et al. (2013) provide evidence that lower rates of tax on patent income attract particularly innovative projects with high earning potential. In the lower panel we consider the effect on revenue when we consider only high quality patents. The picture is here is similar; the introduction of Patent Boxes results in a substantial reduction in revenues.","The literature has emphasised the downward pressure on corporate income tax rates that arises from factor mobility. There is also a large literature that discusses the strategies firms use to shift income for tax purposes and to circumvent anti-avoidance rules, and that highlights an important role for intangible assets. However, we know relatively little about the extent to which the location of intangible assets responds to tax. The evidence there is on the impact of tax on the location of capital more generally has tended to suffer from the imposition of restrictive a priori assumptions placed on the underlying model of firm behaviour. From a policy perspective it is clearly important to understand how responsive firms are to corporate income taxes when they make location decisions. In this paper, we estimate a model of firms' decisions over where to locate the legal ownership of their patents. We find that corporate tax rates are an important determinant of location choice. We extend the current literature on the determinants of firm location choice by estimating a flexible choice model, which accounts for both observed and unobserved heterogeneity in behaviour. We are able to generate own and cross tax elasticities across locations that capture complex patterns of substitution in the data. The model can be used to conduct ex ante analysis of policy changes. We find that this heterogeneity is important for explaining location choices. Our model also shows that other factors influence where firms choose to hold legal ownership of patents. For instance, firms are more likely to locate patent ownership in countries where they have associated real innovative activity. This may reflect co-location externalities, or the influence of tax rules which seek to limit the extent to which income and real innovative activity can be geographically separated. Firms also value other non-tax location characteristics. Such factors, along with tax rules like the operation of CFC regimes that limit the tax advantages of locating patent ownership in low tax jurisdictions, help explain why we do not see firms choosing to hold all legal ownership of patents in the lowest tax locations. We use the model to consider the impact of the recent introduction of preferential tax regimes for income from patents. These Patent Boxes are likely to attract patent income, but our estimates suggest that they will also lead to substantial falls in tax revenues. Of course some of this revenue loss might be offset by gains from attracting activities that yielded positive externalities; these would need to be taken into account in a calculation of the welfare impact of the policy. It is also possible that the tax reforms will affect firms' decisions over whether to apply for a patent on a new technology or whether to rely on secrecy. We do not have information that would allow us to directly estimate this margin, but this would be an interesting avenue for future research. The introduction of Patent Boxes by several European countries in a relatively short space of time has given rise to concerns that countries are engaging in tax competition for patent income. In future work we intend to build on the framework developed here to consider whether governments are engaged in a strategic game to attract income from intellectual policy that ultimately will continue to exert downward pressure on corporate taxes."],["We evaluate the impact of a policing experiment that depenalized the possession of small quantities of cannabis in the London borough of Lambeth, on hospital admissions related to illicit drug use. To do so, we exploit administrative records on individual hospital admissions classified by ICD-10 diagnosis codes. These records allow the construction of a quarterly panel data set for London boroughs running from 1997 to 2009 to estimate the short and long run impacts of the depenalization policy unilaterally introduced in Lambeth between 2001 and 2002. We find that the depenalization of cannabis had significant longer term impacts on hospital admissions related to the use of hard drugs, raising hospital admission rates for men by between 40 and 100% of their pre-policy baseline levels. The impacts are concentrated among men in younger age cohorts. The dynamic impacts across cohorts vary in profile with some cohorts experiencing hospitalization rates remaining above pre-intervention levels three to four years after the depenalization policy is introduced. We combine these estimated impacts on hospitalization rates with estimates on how the policy impacted the severity of hospital admissions to provide a lower bound estimate of the public health cost of the depenalization policy. © 2014. --------------------------------------------------------------------------------","Illicit drug use generates substantial economic costs including those related to crime, ill-health, and diminished labor productivity. In 2002, the Office for National Drug Control Policy estimated that illicit drugs cost the US economy $181 billion (ONDCS, 2004). For the UK, Gordon et al. (2006) estimated the cost of drug-related crime and health service use to be £15.4 billion in 2003/4. It is these social costs, coupled with the risks posed to drug users themselves, that have led governments throughout the world to try and regulate illicit drug markets. All such policies aim to curb both drug use and its negative consequences, but there is ongoing debate among policy-makers as to relative weight that should be given to policies related to prevention, enforcement, and treatment (Grossman et al., 2002). The current trend in policy circles is to suggest regimes built solely around strong enforcement and punitive punishment might be both costly and ineffective. For example, after forty-years of the US ‘war on drugs’, the Obama administration has adopted a strategy that focuses more on prevention and treatment, and less on incarceration (ONDCS, 2011), although the two primary enforcement and policy agencies of the Drug Enforcement Agency and the Office for National Drug Control Policy remain more focused on traditional supply-side approaches. Other countries such as the Netherlands, Australia and Portugal, have long adopted more liberal approaches that have depenalized or decriminalized the possession of some illicit drugs, most commonly cannabis, with many countries in Latin America currently debating similar moves.1 While such policies might well help free up resources from the criminal justice system and stop large numbers of individuals being criminalized (Adda et al., 2013), these more liberalized policies also carry their own risks. If such policies signal the health and legal risks from consumption have been reduced, then this should reduce prices (Becker and Murphy, 1988). This can potentially increase the number of users as well as increasing use among existing users, all of which could have deleterious consequences for user's health. The use of certain drugs might also provide a causal ‘gateway’ to more harmful and addictive substances (van Ours, 2003; Melberg et al., 2010). This paper considers the impact of a localized policing experiment that reduced the enforcement of punishments against the use of one illicit drug-cannabis-on a major cost associated with the consumption of illegal drugs: the use of health services by consumers of illicit drugs. The policing experiment we study took place unilaterally in the London Borough of Lambeth and ran from July 2001 to July 2002, during which time all other London boroughs had no change in policing policy towards cannabis or any other illicit drug. The experiment-known as the Lambeth Cannabis Warning Scheme (LCWS)-meant that the possession of small quantities of cannabis was temporarily depenalized, so that this was no longer a prosecutable offense.2 We evaluate the short and long run consequences of this policy on healthcare usage as measured by detailed and comprehensive administrative records on drug- related admissions to all London hospitals. Such hospital admissions represent 60% of drug-related healthcare costs (Gordon et al., 2006). To do so we use a difference-in- difference research design that compares pre- and post-policy changes in hospitalization rates between Lambeth and other London boroughs. Our analysis aims to shed light on the broad question of whether policing strategies towards the market for cannabis impact upon public health, through changes in the use of illicit drugs and subsequent health of drug users. Our primary data comes from a novel source that has not been much used by economists: the Inpatient Hospital Episode Statistics (HES). These administrative records document every admission to a public hospital in England, with detailed ICD-10 codes for classifying the primary and secondary causes of each individual hospital admission.3 This is the most comprehensive health related data available for England, in which it is possible to track the admissions history of the same individual over time. We aggregate the individual HES records to construct a panel data set of hospital admissions rates by London borough and quarter. We do so for various cohorts defined along the lines of gender, age at the time of the implementation of the depenalization policy, and previous hospital admission history. As such these administrative records allow us to provide detailed evidence on the aggregate impact of the depenalization policy on hospitalization rates, and to provide novel evidence on how these health impacts vary across cohorts. To reiterate, these administrative records cover the most serious health events. Patients with less serious conditions receive treatment elsewhere, including outpatient appointments, accident and emergency departments, or primary care services. If such health events are also impacted by drug policing strategies, our estimates based solely on inpatient records provide a strict lower bound impact of the depenalization of cannabis on public health. The balanced panel data we construct covers all 32 London boroughs between April 1997 and December 2009. This data series starts four years before the initiation of the depenalization policy in the borough of Lambeth, allowing us to estimate policy impacts accounting for underlying trends in hospital admissions. The series runs to seven years after the policy ended, allowing us to assess the long term impacts of a short-lived formal change in policing strategy related to cannabis. Given the detailed ICD-10 codes available for each admission, the administrative records allow us to specifically measure admission rates for drug-related hospitalizations for each type of illicit drug: although the depenalization policy would most likely impact cannabis consumption more directly than other illicit drugs, this has to be weighed against the fact that hospitalizations related to cannabis usage are extremely rare and so policy impacts are statistically difficult to measure along this margin. Our main outcome variable therefore focuses on hospital admissions related to hard drugs, known as ‘Class-A’ drugs in England. This includes all hospital admissions where the principal diagnosis relates to cocaine, crack, crystal-meth, heroin, LSD, MDMA or methadone.4 The administrative records also contain information on the length of hospital stays (in days) associated with each patient admission, and we use this to explore whether the depenalization policy impacted the severity of hospital admissions (not just their incidence), where the primary diagnosis relates to hospitalizations for Class-A drug use. Ultimately, we then combine the estimated policy impacts on hospitalization rates and the severity of hospital admissions for Class-A drug use, to provide a conservative estimate of the public health costs of the depenalization policy that arises solely through the increased demand on hospital bed services. We present four main results. First, relative to other London boroughs, the depenalization policy had significant long term impacts on hospital admissions in Lambeth related to the use of Class-A drugs, with the impacts being concentrated among men. Exploring the heterogeneous impacts across male cohorts, we find the direct impacts on Lambeth residents to be larger among cohorts that were younger at the start of the policy. The magnitudes of the impacts are large: the increases in hospitalization rates correspond to rises of between 40 and 100% of their pre-policy baseline levels in Lambeth, for those aged 15–24 and aged 25–34 on the eve of the policy. To underpin the credibility of the difference-in- difference research design, we also probe the data to: (i) check for pre-existing divergent trends in hospitalization rates between Lambeth and other London boroughs; (ii) evaluate the robustness of the results to alternative control boroughs to compare Lambeth to; (iii) examine whether differential changes over time in health care provision between Lambeth and other locations, or other policies impacting hospitalizations for Class-A drug use, could confound the results, and; (iv) shed light on whether individuals changed borough of residence in response to the policy. Second, the dynamic impacts across cohorts vary in profile with some cohorts experiencing hospitalization rates remaining above pre- intervention levels three to four years after the depenalization of cannabis was first introduced. Third, we explore the impacts of the policy on hospitalizations related to alcohol use among Lambeth residents. There is a body of work examining the relationship between cannabis and alcohol use: this has generated mixed results with some research finding evidence of the two being complements (Pacula, 1998; Williams et al., 2004), and other studies suggesting that the two are substitutes (DiNardo and Lemieux, 2001; Crost and Guerrero, 2012). We add to this debate using a novel policy experiment and administrative data. Our results suggest that for the youngest age cohort, if depenalization causes the price of cannabis to fall, then alcohol and cannabis might well be substitutes. However for older age cohorts, we find no evidence that the policy leads to increased admissions related to alcohol use, or the combined use of alcohol and Class-A drugs. Finally, the severity of hospital admissions, as measured by the length of stay in hospital, significantly increases for admissions related to Class-A drug use. We then combine this impact with our baseline estimated impacts on hospitalization rates by age cohort, to calculate the annual cost of the policy. We find the increased hospitalization rates and length of stays conditional on admission to be around £80,000 per annum, and this more than offsets the downward time trend in hospital bed-day costs that exists in the rest of London in the post-policy period. Taken together, our four classes of results suggest that policing strategies towards the market for cannabis have significant, nuanced and long lasting impacts on public health. Our analysis contributes to understanding the relationship between drug policies and public health, an area that has received relatively little attention despite the sizable social costs involved. This partly relates to well known difficulties in evaluating policies in illicit drug markets: multiple policies are often simultaneously targeted towards high supply locations; even when unilateral policy experiments or changes occur they often fail to cause abrupt or quantitatively large demand or supply shocks, and data is rarely detailed enough to pin down interventions in specific drug markets on other drug-related outcomes (DiNardo, 1993; Caulkins, 2000; Chu, 2012). Our analysis, that combines a focused policy and administrative records, makes some progress on these fronts. To place our analysis into a wider context, it is useful to compare our findings with two earlier prominent studies linking illicit drug enforcement policies and health outcomes: Model (1993) uses data from the mid-1970s to estimate the impact on hospital emergency room admissions of cannabis decriminalization, across 12 US states. She finds that policy changes led to an increase in cannabis-related admissions and a decrease in the number of mentions of other drug related emergency room admissions, suggesting a net substitution towards cannabis. Our administrative records also allow us to also check for such broad patterns of substitution or complementarity between illicit drugs. Our results suggest that the depenalization of cannabis led to longer term increases in the use of Class-A drugs, as measured by hospital inpatient admissions rather than emergency room admissions as in Model (1993).5 More recent evidence comes from Dobkin and Nicosia (2009), who assess the impact of an intervention that disrupted the supply of methamphetamine in the US by targeting precursors to methamphetamine. They document how this led to a sharp price increase and decline in quality for methamphetamine. Hospital admissions mentioning methamphetamine fell by 50% during the intervention, while admissions into drug treatment fell by 35%. Dobkin and Nicosia (2009) find no evidence that users substituted away from methamphetamine towards other drugs. Finally, Dobkin and Nicosia (2009) find that the policy of disrupting methamphetamine supply was effective only for a relatively short period: the price of methamphetamine returned to its pre- intervention level within four months and within 18 months hospital admissions rates had returned to their baseline levels. In contrast, the cannabis depenalization policy we document has an impact on hospitalization rates that, for many cohorts, lasts for up to four years after the policy was initiated and despite the fact that the policy itself was only formally in place for one year.6 The paper is organized as follows. Section 2 describes the LCWS and the existing evidence on its impact on crime. Section 3 details our administrative data, discusses the plausibility of a link between policing-induced changes in the cannabis market and the consumption of Class-A drugs, and describes our empirical method. Section 4 presents our baseline results which estimate the impact of the LCWS by cohort and the associated robustness checks to underpin the credibility of the research design. Section 5 presents extended results related to dynamic effects, spillovers in alcohol-related admissions, and the severity of admissions. Section 6 estimates the public health costs of the policy. Section 7 discusses the broader policy implications of our findings, and the potential for opening up a research agenda on the relationship between police behavior and public health.","The Lambeth Cannabis Warning Scheme (LCWS) was unilaterally introduced into the London borough of Lambeth on 4th July 2001 by the borough's police force. The scheme was initially launched as a pilot intended to last six months, and represented a change in policing policy towards the market for cannabis. Under the scheme, those found in possession of small quantities of cannabis for their personal use in Lambeth: (i) had their drugs confiscated; (ii) were given a warning rather than being arrested.7 The main reason behind the policy change was to reduce the number of individuals being criminalized for consuming cannabis, and to free up police time and resources to deal with more serious crimes, including those related to hard drugs or ‘Class-A drugs’ (Dark and Fuller, 2002; Adda et al., 2013). The underlying motivations for the policy, as well as the way in which it was implemented and the targeted outcomes, were very similar to the way in which cannabis depenalization policies have often been implemented throughout the world. In keeping with other experiences of depenalization, the primary motivation behind the policy was to free up police time and resources to tackle other crimes, and there was little or no discussion of the depenalization policy's potential impact on public health. To this extent our results can be informative of the existence of links between police drugs policy and public health in settings outside of the specific London context we study.8 Anecdotal evidence suggests that local support for the scheme began to decline once the policy was announced to have been extended beyond the initial six-month pilot. Media reports cited that local opposition arose due to concerns that children were at risk from the scheme, and that the depenalization policy had increased drug tourism into Lambeth. The LCWS formally ended on 31st July 2002. Post-policy, Lambeth's cannabis policing strategy did not return identically to what it had been pre-policy, partly because of disagreements between the police and local politicians over the policy's true impact. Rather, it adjusted to be a firmer version of what had occurred during the pilot so that police officers in Lambeth continued to issue warnings but would now also have the discretion to arrest where the offense was aggravated.9 Hence our measured long run impacts of the depenalization policy capture the total effects arising from: (i) the long run impact of the introduction of the depenalization policy between June 2001 and July 2002; and (ii) any permanent differences in policing towards cannabis between the pre- policy and post-policy periods. The impact of the LCWS depenalization policy on patterns of crime in Lambeth and other boroughs is extensively studied in Adda et al. (2013). For the purposes of the current study on the relationship between drug-policing and public health, three key results on the impact of the localized depenalization policy on crime need to be borne in mind: its impact on the market for cannabis in Lambeth, its impact on the market for Class-A drugs, and drug tourism induced into Lambeth from other parts of London due to interlinkages in illicit drug markets across boroughs. First, Adda et al. (2013) find that the LCWS led to a significant and permanent rise in cannabis related criminal offenses in Lambeth. Using finely disaggregated data by type of drug offense, they find that both the demand and supply of cannabis are likely to have risen significantly in Lambeth after the introduction of the depenalization policy, and that this impact persists into the long run, well after the LCWS policy officially ended. This result is important for the current study because it suggests that the depenalization policy caused an abrupt, quantitatively large and permanent shock to the cannabis market, leading to the equilibrium market size to have likely increased by around 60% in the longer term, as proxied by the number of criminal offenses for cannabis possession.10 Second, this expansion will consequently affect the equilibrium market size for Class-A drugs if the markets are related in some way, either because of economies of scale in supplying both drug markets, or because on the demand side preferences are such that cannabis and Class-A drugs are complements/substitutes. Along these lines, Adda et al. (2013) report that the longer term effect of the LCWS was to lead to a significant increase in offenses related to the possession of Class-A drugs: offense rates for the possession of such substances rose by 12% in Lambeth in the post-policy period relative to the rest of London. However, there is little evidence that the police reallocated their efforts towards crimes relating to Class-A drugs: Adda et al. (2013) report no change in police effectiveness against Class-A drug crime in Lambeth based on two out of four such measures (arrest and clear-up rates). Rather, Adda et al. (2013) document that the policy appears to have allowed the police to reallocate effort towards non-drug crime. The fact that the LCWS policy did not lead to a major reallocation of police resources towards crime related to Class-A drugs suggests that in the current study, any link between the depenalization policy and hospitalizations for Class-A diagnoses most likely stems from the interlinkages between the demand sides of the markets for cannabis and Class-A drugs. Given the addictive nature of Class-A drugs, potential lags between cannabis use and the use of Class-A drugs later in life, and potential lags in seeking out and receiving treatment (Fergusson and Horwood 2000; Patton et al., 2002; Arseneault et al., 2004), we might also reasonably expect any impact of the LCWS on hospital admissions related to Class-A drug use to last well into the post-policy period. We therefore later consider how the effects of the LCWS on drug-related hospital admissions evolve over time. Third, Adda et al. (2013) document how the LCWS likely induced drug tourism into Lambeth. Such changes in the location where individuals decided to purchase cannabis stems from the fact that local markets for illicit drugs are inherently interlinked across London boroughs. To explore this further in terms of health outcomes, we later investigate whether there is any evidence of individuals permanently changing their actual borough of residence to Lambeth, after the LCWS is introduced. Standard consumer theory provides clear set of predictions on how such depenalization policies can impact the use of cannabis and other illicit drugs. Most existing studies assume that such policies cause significant reductions in the price of cannabis (Thies and Register, 1993; Grossman and Chaloupka, 1998; Williams et al., 2004). This will, all else equal, increase the demand for cannabis in part because of greater demands from existing users and also because of an impact on the extensive margin so that new individuals choose to start consuming cannabis at the lower price. This will have a positive impact on the consumption of Class-A drugs if cannabis and Class-A drugs are contemporaneous complements in user preferences. It will also increase the demand for Class-A drugs over time if the use of cannabis serves either as a gateway to the use of other harder illicit drugs, or there is state dependence so that cannabis users have particular characteristics that also lead them to subsequently misuse Class-A drugs. Of course if cannabis and Class-A drugs are substitutes, then the increased demand for cannabis resulting from the depenalization of cannabis possession should reduce Class-A drug use and related hospitalizations. Such cross price impacts might also exist between cannabis and alcohol (Pacula, 1998; DiNardo and Lemieux, 2001). Hence we later also examine how hospitalization rates for diagnoses primarily related to alcohol use respond to the depenalization policy. Administrative records on hospital admissions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Data on hospital admissions are drawn from the Inpatient Hospital Episode Statistics (HES). These provide an administrative record of every inpatient health episode, defined as a single period of care under one consultant in an English National Health Service hospital.11 These administrative records are the most comprehensive data source on health service usage for England. Inpatients include all those admitted to hospital with the intention of an overnight stay, plus day case procedures when the patient is formally admitted to a hospital bed. As such, these records cover the most serious health events. Patients with less serious conditions receive treatment elsewhere, including outpatient appointments, accident and emergency departments, or primary care services. If such health events are also impacted by the depenalization policing strategy, our estimates based solely on inpatient records provide a strict lower bound impact of the policy on public health. For each patient-episode event in the administrative records, the data record the date of admission, total duration in hospital, and ICD-10 diagnoses codes in order of importance. Background patient information covers their age, gender, and their zip code of residence at the time of admission.12 As discussed in more detail later, some specialist services required to treat diagnoses involving the use of Class-A drugs, such as those relating to mental health, are concentrated in a small subset of facilities that are dispersed across London. Each of these specialist facilities would be expected to treat patients from across London. Hence the geographic information we use to understand the impact of the localized LCWS policy relates to the patient's borough of residence, not the borough in which they are hospitalized. This helps ameliorate concerns that any changes in drug related hospitalization rates are driven by changes in the provision of specific drug-related services through specialized hospitals in London (that serve individual residents in multiple boroughs). In Section 4.2 we provide evidence ruling out potentially confounding changes on the supply side of medical care for heavy users of Class-A drugs. Hence, any documented change in hospital admissions for Class-A drug related diagnoses in Lambeth following the introduction of the LCWS might then operate through two mechanisms: (i) a change in behavior of those residents in Lambeth prior to the policy; and (ii) a change in the composition of Lambeth residents, with the policy potentially inducing a net inflow of people into the borough with a higher propensity for Class-A drug use. In Section 4.2 we use our data to examine the relative importance of these channels: we find little evidence of systematic changes of residence in response to the policy, implying that most of the impacts are driven by changes in behavior among those already residing in Lambeth pre-policy. The administrative records also allow us to create panels based on prior histories of patient admissions because the HES records have unique patient identifiers that allow the same patient to be tracked over episodes between 1997 and 2009. We focus on histories of admissions related to the use of either drugs (Class-A drugs, cannabis, or other illicit drug), or alcohol, and create panels by borough-quarter-age cohort-gender, for those with and without pre-policy histories of admissions related to drugs or alcohol. Among those with no pre-policy admissions, we calculate admission rates as per Eq. (1), where by construction of this admission rate is zero before the policy. For this group, we effectively estimate whether the depenalization policy differentially impacted hospital admission rates between Lambeth and other non-neighboring boroughs in the period after the policy is first initiated. For those with pre-policy admission rates (an obviously far smaller group of individuals than those without admission histories), we change the numerator in the admission rate to reflect the relevant ‘at risk’ population: hence Popby is replaced by the number of distinct individuals admitted for diagnoses related to illicit drugs or alcohol while residing in borough b in the pre-policy period between April 1997 and June 2001 (which given the small number of individuals with histories of such hospitalizations, is not measured in thousands). The depenalization policy likely lowers prices for cannabis in Lambeth, all else equal. Depenalization might then impact hospitalizations for Class-A diagnoses differently across cohorts based on their prior histories of illicit drug use. Among those with no prior history of hospitalization for drug or alcohol use, the reduced price of cannabis induced by the policy might lead to greater consumption of Class-A drugs if they are complements to cannabis, or, for example, cannabis acts as a gateway to such substances. To be clear, among these cohorts we pick up the combined impacts among those that were previously using illicit drugs (and potentially other substances) but not so heavily so as to induce hospitalizations, as well as those that begin to use cannabis and Class-A drugs for the first time as a result of the reduced price of cannabis. The administrative data utilized does not allow us to separate out the policy impacts stemming from each type of individual. Among the cohorts with histories of hospitalization for drug or alcohol use even before the LCWS is initiated, there are likely to be long term and heavy users of illicit substances. Such individuals' consumption of Class-A drugs might reasonably be more habitual and so less sensitive to changes in the price of cannabis, so that this cohort might be less impacted by the depenalization policy, all else equal. Cannabis and Class-A drug use ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our primary interest is to understand how changes in police enforcement strategies towards the cannabis market-as embodied in the LCWS policy impact public health through changes in hospitalization rates related to illicit drug use. Of course the policy would most directly affect the consumption of cannabis, but changes in inpatient hospital admissions related to cannabis use are statistically hard to detect given the rarity of such events, as documented in detail below. It is therefore instructive to first compare rates of drug related hospital admissions from the HES administrative records, to rates of self-reported drug use from household surveys the most reliable of which is the British Crime Survey (BCS). Estimates from the BCS in 2002/3 indicate that cannabis was by far the most popular illicit drug, with 16% of 16–24 year-olds and 9% of 25–34 year-olds reporting to have used cannabis in the month prior to the survey. The corresponding figures for Class-A drug use are just 4% and 2% respectively (Condon and Smith, 2003). The HES records show that there are seven times as many inpatient hospital admissions for Class-A drugs than for cannabis. This reinforces the notion that cannabis related policing policies such as the LCWS, may not lead to a rise in statistically detectable cannabis related to hospital admissions even if there is a substantial increase in cannabis usage caused by the policy. What is important for our analysis is that a body of evidence suggests the cannabis and Class-A drug markets are linked: while little is known about such potential linkages on the supply side, on the demand side this might be because cannabis users are more likely to consume Class-A drugs, both contemporaneously and in the future. There are of course multiple explanations for this positive correlation between admissions for cannabis and subsequent usage of Class-A drugs. One explanation is state dependence so that cannabis users have particular characteristics that also lead them to subsequently misuse Class-A drugs, a channel shown to be of first order importance using data from the NLSY97 by Deza (2011). Alternatively, the use of cannabis might act as a causal gateway to the use of harder drugs, has been suggested by Beenstock and Rahav (2002); van Ours (2003); Bretteville- Jensen et al. (2008) and Melberg et al. (2010). Clearly the empirical debate on the relative importance of state dependence and gateway impacts is far from settled. For our study what is important is that some correlation between the market sizes for cannabis and other illicit drugs exists, be it either because of state dependence or gateway effects. To show the relatedness between these markets as recorded in the hospital admission records we exploit, we present descriptive evidence from the HES to suggest how cannabis consumption today might correlate to Class-A drug use in the future. To do so we exploit the individual identifiers in the administrative records, allowing us to track the same person over time. We then calculate the probability, conditional on an admission in 1997 or 1998, of being readmitted to hospital at least once between 2000 and 2004. Four groups of admission are considered: (i) “cannabis admissions”, who were admitted for cannabis; (ii) “Class-A admissions”, who were admitted for the use of a harder drug; (iii) “alcohol admissions”, who were admitted for alcohol-related diagnoses; (iv) “all other admissions”, who were admitted for any other cause and serve as a benchmark for the persistence of ill- health over these time periods. Table 1 shows the mean and standard deviation for each probability of readmission, conditional on prior admissions.15 Two points are of note. First, there is substantial persistence in hospital admissions for the same risky behavior, as shown on the leading diagonal in Columns 1–3. Persistence is particularly high for Class-A drugs and alcohol, where 26 and 23% of individuals respectively, were readmitted for the ill-effects of the same risky behavior over the two time periods. Reading across the last row of Table 1 on subsequent readmission to hospital from 2000 to 2004 for any diagnosis unrelated to drugs or alcohol, we see that this readmission probability is between 15 and 28% conditional on having been previously admitted in 1997–8 for some risky behavior related to illicit drug or alcohol use. Second, although admissions for any form of risky behavior in 2000–4 is best predicted by admission for the same behavior in 1997–8, we note that for those admitted for Class-A drugs in 2000–4, 5.4% will have been admitted for cannabis related diagnoses in 1997–8. This is significantly higher than having been previously admitted for alcohol-related diagnoses (2.2%) over the same period. This highlights the particularly robust correlation between cannabis use at a given moment in time, and future hospital admissions for Class-A related drug use. In this paper our focus is on establishing whether a change in police enforcement in the cannabis market — as embodied in the LCWS — has a causal impact on hospital admissions for Class-A drugs. The evidence presented in Table 1 and the existing evidence documenting a linkage between cannabis consumption on the subsequent use of other illicit substances, suggests that as long as the policy affects the usage of cannabis consumption in some way, this is likely to have a knockon effect on the usage of Class-A drugs in the long run. It is these longer term effects on public health that we now focus on. Empirical method ~~~~~~~~~~~~~~~~ β0 captures London-wide cohort trends (excluding Lambeth's neighbors) in hospitalization rates occurring at the same time as the LCWS was in operation in Lambeth. β2 captures longer term London-wide cohort trends in hospitalization rates for the age cohort after the depenalization policy in Lambeth officially ends. This coefficient mostly picks up the natural time profile of any change in hospitalizations as the cohort ages say because of varying usages of illicit substances, or changes in susceptibility to the same levels of usage. These coefficients also partly pick up any impacts on hospitalization rates related to diagnosis-d for London and nationwide policies, including the nationwide depenalization of cannabis possession that occurred from January 2004 through to January 2009.16 The parameters of interest are estimated using a standard difference-in-difference research design: β1 and β3 capture differential changes in hospital admission rates for a given age cohort, in Lambeth during and after the depenalization policy period, relative to other London boroughs excluding Lambeth's neighbors. Our research design identifies whether: (i) hospitalization rates in Lambeth significantly diverge away from London-wide cohort trends during and after the depenalization policy is in place; and (ii) these divergences coincide with the depenalization policy's operation in Lambeth. In Xbqy we control for two sets of borough-specific time varying characteristics. The first contains the shares of the population under 5 and over 75 (by borough and year), who place the heaviest burden on health services. Second, Xbqy includes controls for admission rates, by borough-quarter- cohort, for conditions that should be unaffected by the LCWS, in particular malignant neoplasms, diseases of the eye and ear, diseases of the circulatory system, diseases of the respiratory system, and diseases of the digestive system. These capture contemporaneous changes in healthcare provision or levels of illness in the population that could affect drug-related admissions. The admission rates for these diagnoses are all constructed from the HES administrative records. The fixed effects capture remaining permanent differences in admissions by borough (λb) and quarter (λq). Observations are weighed by borough shares of the London-wide population. Defining t as quarters since April 1997: t = [4 × (y − 1997)] + q, we assume a Prais–Winsten borough specific AR(1) error structure, ubqy = ubt = ρbubt − 1 + ebt, where ebt is a classical error term. ubqy is borough specific heteroskedastic, and contemporaneously correlated across boroughs.17 As with any difference-in-difference research design, the coefficients of interest measure causal impacts only under some identifying assumptions. First, we have to assume common trends in hospitalization rates between Lambeth and the rest of London. We later present evidence to establish whether there is any evidence of such convergent/divergent trends in the pre-policy period, and we also estimate our baseline specifications allowing for borough specific linear time trends. Second, we require that there is no ‘Ashenfelter dip’, which might otherwise indicate that the policy was introduced in response to divergent/convergent hospitalization rates between Lambeth and the rest of London. The descriptive time series evidence presented below helps ameliorate this concern. Third, we require there to be no confounding changes on the supply side of medical care impacting hospital admissions for Class-A drug use, nor any other confounding policies impacting such outcomes. We later provide descriptive evidence to show how the availability of health care in Lambeth for such diagnoses changed over time. We also show the impacts of the LCWS on Class-A admissions in Lambeth prior to the introduction of the nationwide depenalization policy in 2004. Fourth, we require that there is no differential change in the underlying populations who are resident in Lambeth and the rest of London that might drive divergences in hospitalization rates for Class-A admissions. Given our hospitalization rates are based on borough of residence and not borough of treatment, we later discuss, given the available evidence, the plausibility of individuals with differing propensities for drug use changing their borough of residence in response to the policy. Hospitalization counts Table 2 shows the raw count data (Totdbqy) for the average number of hospital admissions for diagnosis d, that occur in borough b in quarter q in year y, covering diagnoses related to the use of illicit substances such as Class-A drugs and cannabis, as well as for alcohol (in each case we show the sum of primary and secondary diagnoses). We break down admission numbers for Lambeth and the rest of London, averaging over the pre-policy and post-policy periods. Given that hospitalization rates for such diagnoses are higher for men than women, Table 2 presents the data for three male age cohorts, where age is defined on the eve of the introduction of the LCWS policy. Three points are of note. First, admission rates for Class-A related diagnoses are low in absolute numbers in the pre-policy period for all age cohorts. These low levels of baseline counts imply that large percentage increases can be generated by a small change in the absolute number of admissions related to the use of Class-A drugs. Second, for the younger two age cohorts, admission numbers for Class-A diagnoses rise dramatically post-policy. In each case, the absolute increase between the post- and pre-policy periods is larger in Lambeth than the rest of London average (despite Lambeth having higher admission counts than other London boroughs pre-policy for all age cohorts). For the oldest cohort, those aged 35–44 on the eve of the LCWS policy, the count data suggest a slight fall in Class-A admissions in Lambeth but a rise in the average for the rest of London. These broad descriptive patterns in absolute counts will be replicated later in the formal analysis when Eq. (2) is estimated for admission rates. The third point of note from Table 2 on counts relates to diagnoses for cannabis or alcohol. We see that for each male age cohort, admission counts for cannabis related diagnosis are considerably rarer than for Class-A related diagnoses, and this remains true post-policy. As argued above, using these administrative records on hospital admissions, it is therefore considerably harder to statistically detect any significant impact of the LCWS on cannabis use through hospitalizations for cannabis. In contrast, we see that alcohol-related hospital admissions are the most frequent for all age cohorts: pre-policy, there are around four times as many such admissions in London on average than for Class-A related diagnoses. Given the body of existing evidence on potential interlinkages between the use of cannabis, Class-A drugs and alcohol, we later examine whether the depenalization policy had any impact on hospital admissions involving alcohol- related diagnoses. Unconditional impacts of hospitalization rates The core outcome considered in the empirical analysis is hospital admissions rates as defined in Eq. (1). Fig. 1A shows the time series for hospital admission rates in Lambeth against the rest of London averages, for each male age cohort. Each figure is centered on the time of policy change in Lambeth: the dashed red lines indicate the start and end points of the official period of operation of the LCWS policy. Each time series is averaged annually. In order to line up with the policy period, each year starts from Q3 of that year and averages to Q2 in the following year (so for example the value for 1997 is the average over 1997Q3–1998Q2). Although the time series for Lambeth is quite volatile, three points emerge from the comparison with other London boroughs: (i) there is no systematic difference in pre-trends between Lambeth and the rest of London at least for the two older age cohorts, nor is there any evidence of an ‘Ashenfelter dip’ in admission rates in Lambeth just prior to the introduction of the LCWS; (ii) there are divergences in admission rates in Lambeth relative to the rest of London for each age cohort; and (iii) London wide time series in hospital admission rates appear rather flat and not trending upwards or downwards, certainly for the two older age cohorts. Fig. 1B repeats the figures comparing Lambeth only to other boroughs with a similarly (high) incidence of Class-A drug related hospital admission pre-policy. The same broad patterns can be seen in the three time series for each male age cohort in Lambeth against this control group.18 Table 3 then provides descriptive evidence on the unconditional long term effects of the depenalization policy on Class-A related hospital admission rates, with each row showing hospital admission rates (Admitdbqy) as defined in Eq. (1). We again first focus on male cohorts of various ages on the eve of the LCWS policy. Columns 1 and 2 present mean hospital admission rates related to Class-A drug usage in Lambeth during the pre-policy and post-policy periods respectively; Columns 3 and 4 give the corresponding statistics for the average borough in the rest of London (excluding Lambeth's neighboring boroughs). Pre-policy, Lambeth had substantially higher rates of admissions than the London average. Indeed, in ranking boroughs by their per- policy hospital admission rates related to Class-A drugs, Lambeth has the third highest for men and the second highest for women. However, as suggested in Fig. 1 and shown more formally later, there is no evidence of diverging or converging trends in Class-A related hospital admission rates between Lambeth and the London average in the pre-policy period from 1997 to 2001. In Lambeth, admission rates in the pre-policy period are lowest for the youngest cohort, reflecting the overall pattern of drug admissions by age. Comparing Columns 1 and 2 re-iterates the basic pattern of potential health impacts of the depenalization policy, that was previously shown in the raw counts data in Table 2: hospital admission rates in Lambeth rise over time for the 15–24 and 25–34 cohorts, but fall slightly for the oldest cohort. In contrast for the rest of London admission rates rise only for the youngest cohort and are stable or declining for the older two age cohorts. Columns 5 and 6 then present difference-in-difference estimates of how Class-A drug admission rates relate to the LCWS policy. Column 5 shows that unconditional on any other factor, admission rates for both the 15–24 and 25–34 cohorts significantly rise in Lambeth relative to the London borough average, after the introduction of the policy to depenalize the possession of cannabis. The relative increases in admission rates of .054 and .079 per thousand population for the youngest two age cohorts are statistically significant at the 5% level: the increases correspond to a 146% rise relative to the pre-policy level for the 15–24 cohort, and a 44% increase above the baseline level for the cohort aged 25–34 on the eve of the policy. The effect for the oldest cohort is not statistically significantly different from zero. Column 6 then shows this basic pattern of difference-in-differences to remain in magnitude and significance once borough and quarter year fixed effects are controlled for. These results suggest that among younger male age cohorts, the policy of depenalizing the possession of cannabis is associated with significantly higher hospitalization rates in Lambeth for Class-A drug use in the longer term. Table A1 shows the corresponding results for female age cohorts: we find no significant impacts on Class-A related hospitalization for any female age cohort. The rate of admissions for such diagnoses among women is generally lower than among men and this might be one reason that it is harder to statistically detect any impact at conventional significance levels. At the same time, the fact that there are very different trends in hospitalizations for Class-A drugs across genders within Lambeth, suggests that the earlier results for men are not merely picking up other changes in hospital behavior or how diagnoses are recorded within Lambeth, that might otherwise have been expected to impact men and women equally. Table A2 shows the corresponding descriptive evidence for hospital admissions related to cannabis use for male cohorts. Cannabis hospital admission rates are generally lower than for Class-A drugs, especially among older age cohorts, despite much higher levels of cannabis usage as suggested by survey data. The difference-in-difference results suggest the LCWS had no significant impact on hospital admissions for cannabis: the point estimates for the youngest male cohorts are positive but not precisely estimated, and a similar set of findings is obtained when examining the impact of the depenalization policy on hospitalizations for cannabis related diagnoses among female cohorts (not shown). To relate these findings to the literature, recall that Model (1993) find that the de facto decriminalization of cannabis in twelve US states from the mid-1970s significantly increased cannabis-related emergency room admissions. Chu (2012) similarly finds that the passage of US state laws that allow individuals to use cannabis for medical purposes leads to a significant increase in referred treatments to rehabilitation centers. Our evidence from London suggests that if a similar effect occurs from the depenalization of cannabis possession, it does not then feed through to significantly higher rates of hospitalization that involve extreme consequences on health leading to overnight hospital stays, which is what our inpatient administrative data measures. For the bulk of our core analysis, we therefore continue to focus on Class-A drug-related hospital admissions among men. The impact of the LCWS by cohort ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 4 presents estimates of the full baseline specification (2), where we consider the impact of the LCWS on Class-A drug related hospital admission rates for three male age cohorts in Columns 1 to 3. These findings represent our core results: they show that the addition of time varying borough controls (Xbqy) produces estimates very similar to the unconditional estimates shown in Table 2. The first row shows that in the longer term post-policy period, there are statistically significant rises in admission rates of .038 and .075 for the youngest two cohorts in Lambeth, relative to other non-neighboring London boroughs. In line with the earlier descriptive evidence, no policy impact is found on the oldest age cohort, who were aged 34–44 on the eve of the LCWS's introduction in Lambeth. Comparing these increases in admission rates to the mean admission rate in Lambeth pre- policy as reported at the foot of Table 4, the percentage increases conditional on other factors are 103% for the youngest age cohort, and 42% for those men aged 25–34 on the eve of the policy, which are slightly smaller than the unconditional percentages reported in relation to Table 3. The second row of Table 4 shows that in the short-run, during the 13 months in which the LCWS was actually in operation, there are no statistically significant effects on hospitalization rates for any cohort. Hence, as might be expected, any impact of the cannabis depenalization policy on hospitalization rates for Class-A drug use takes time to work through (in line with the descriptive evidence in Table 1). Table 4 also shows the estimates of β0 and β2. These highlight that for London as a whole (excluding Lambeth's neighbors), there are no significant long-term time cohort trends in admission rates during and after the policy period for the older two cohorts. For the youngest cohort aged 15–24 (column 1), hospital admission rates for Class-A drug related admissions are naturally rising over time as the cohort ages, but the results overall show that hospitalization rates in Lambeth are significantly diverging away from this London-wide average in the post-policy period, all else equal. Taken together, our results suggest that the depenalization of cannabis led to longer term increases in the use of Class-A drugs and subsequent hospitalizations related to Class-A drug use among the two youngest aged cohorts on the eve of the LCWS policy. If depenalization led to a decline in the equilibrium price of cannabis in Lambeth, as is often argued to be an unambiguous effect of such policies (Kilmer et al., 2010), then this result suggests that cannabis and Class-A drugs have a negative cross-price elasticity, so that the two types of illicit drug are contemporaneous complements, or the use of cannabis leads through some mechanism to the later use of harder illicit drugs.20 This would be in line with some other studies that have estimated the cross-price elasticity between cannabis and a specific Class-A drug: cocaine, either using decriminalization as a proxy for a price reduction (Thies and Register, 1993; Grossman and Chaloupka, 1998), or using actual price information (Williams et al., 2004). An obvious concern with these results is that they might in part be confounded by natural time trends, by age cohort, in hospitalization rates for Class-A drugs, that are not fully being captured in the policy and post-policy dummies. To address this, we repeat the analysis but augment (2) with controls for borough specific linear time trends. Columns 4 to 6 in Table 4 present the results for each male age cohort when time trends are conditioned on. We find that for the two older male age cohorts, hospitalization rates are significantly higher in Lambeth relative to the rest of London comparing the post- and pre-policy periods. Hence policy impacts remain even once linear within borough time trends are controlled for, although we note the descriptive evidence in Fig. 1 does not provide compelling evidence that such time trends should necessarily be controlled for. In summary the evidence suggests that there are significant impacts of the police policy of depenalizing cannabis on public health, as measured in hospitalization rates for Class-A related drug use. These impacts are quantitatively large, apply to more than one male age cohort, and are observed well after the policy depenalizing the possession of cannabis is officially ended. To be clear, these results cannot be interpreted as suggesting that there are some individuals that start taking Class-A drugs as a result of the depenalization of cannabis. All we can infer is that there are individuals, who prior to the policy might either have not been consuming illicit drugs at all, or were consuming them in quantities that did not lead to hospitalization, who are then in the longer term post-policy, significantly impacted by the depenalization policy so as to require hospitalization for diagnoses related to Class-A drug use. Robustness checks ~~~~~~~~~~~~~~~~~ We now present evidence to underpin the credibility of the difference-in-difference research design. These relate to probing the data to: (i) check for pre-existing divergent trends in hospitalization rates between Lambeth and other London boroughs; (ii) evaluate the robustness of the results to alternative control boroughs to compare Lambeth to; (iii) examine whether differential changes over time in health care provision between Lambeth and other locations, or other policies impacting hospitalizations for Class-A drug use, could confound the results, and; (iv) shed light on whether individuals changed borough of residence in response to the policy. Pre-trends The research design implicitly assumes that in the absence of the depenalization policy, there would have been no natural divergence/convergence in admission rates between Lambeth and the rest of London. The previous set of specifications that allowed for borough specific time trends already partly addressed this concern. A second way to address the issue is to exploit the four years of panel data prior to the introduction of the depenalization policy, from 1997 Q2 until 2001 Q2, using this period to test whether there is any evidence of a divergence in trends in hospitalization rates between Lambeth and the rest of London pre-policy. To do so, we estimate a specification analogous to Eq. (2) in the pre-policy sample but allow for only one split of the sample, midway through the pre-policy period. We then test whether there are divergent trends across Lambeth and the rest of London in admission rates between the first and second halves of the pre-policy period. As Table A5 shows (and consistent with the descriptive evidence in Fig. 1), for all male age cohorts this pre-policy sample split dummy interaction is not significantly different from zero suggesting that hospitalization rates in Lambeth are not diverging from London in the years prior to the depenalization policy. As discussed in Section 2, this is very much in line with the policy discussion around the underlying motivation for the policy, that emphasized the policy enabling the police to reallocate their effort towards non-cannabis crime, and which hardly mentioned the potential impacts on public health. Hence the data supports the assertion that the depenalization policy was not introduced specifically into Lambeth because of worsening public health related to drug- related hospital admissions. Nor is there any evidence of reversion to the mean in hospitalization rates with Lambeth converging back towards London-wide averages. In short, any form of ‘Ashenfelter dip’ does not appear to be confounding the estimated parameters, as was also suggested by the descriptive evidence in Fig. 1. Control boroughs We now examine the robustness of the findings to comparing Lambeth to other subsets of boroughs, rather then all boroughs London-wide (excluding only the immediate neighbors of Lambeth). To begin with, we follow on from the descriptive evidence in Fig. 1B and compare Lambeth to a more limited set of nine other London boroughs with similarly high levels of hospital admission rates for Class-A drugs pre-policy. As shown in Columns 1–3 of Table A6, with this restricted comparison sample of boroughs to Lambeth, there remains a significant impact of the policy among males aged 25–34 on the eve of the policy, an effect significant at the 1% level. The point estimate of the impact (.104) is actually larger than the corresponding coefficient in the baseline specification (.075), as reported in column 2 of Table 4. The remaining columns in Table A6 then show this result to be robust to alternative modifications to the control group of boroughs included: (i) boroughs with the very highest pre-policy admissions related to Class-A drugs (Columns 4–6, restricting the sample to four boroughs); (ii) including neighbors to Lambeth in the control group where spillovers in hospital admissions might have been greatest (Columns 7–9, expanding the sample to 32 boroughs). Along similar lines, the remaining columns in Table A6 consider restricting the control group of boroughs to those in which there exists: (i) a mental health trust headquarters (there are six such boroughs including Lambeth); (ii) a teaching hospital (there are eight such boroughs including Lambeth).21 The intuition for these comparisons is that residents of such boroughs might have access to especially high levels of quality in hospital care, or similar degrees of specialization in dealing with mental health disorders associated with the use of illicit drugs as in Lambeth where one mental health trust is headquartered. In line with the baseline results, we see that there were significant impacts on hospitalization rates post-policy in Lambeth for the youngest two age male cohorts in both these restricted samples. Taken together, these comparisons suggest that our baseline results are not driven solely by differences in health care between Lambeth and other London boroughs. Supply side changes and other confounders To provide further evidence on whether changes on the supply side of health care could be driving the difference-in-difference estimates, we utilize information in the HES administrative records both on the borough of residence of the individual, and the borough of treatment for each hospital episode. We then construct the time series for the percentage of men treated for a Class-A related diagnoses in a Lambeth health facility (by age cohort), that actually reside in Lambeth. To be clear, nearly all London boroughs have a hospital in them (with a handful of boroughs containing two). However, the specialist services required to treat diagnoses involving the use of Class-A drugs, such as those relating to mental health, are more concentrated in a small subset of facilities that are more dispersed across London. Each of these specialist facilities would be expected to treat patients from across London. If such services expanded in Lambeth during the post-policy period and were especially targeted towards Lambeth residents, then we would expect to observe, over time, a greater share of individuals treated in Lambeth to also reside in Lambeth. Fig. 2 presents the relevant time series evidence on this, again split by male age cohorts where ages are defined on the eve of the policy. Within each cohort, we find no evidence of changes in this percentage over time: for all male cohorts, between 30 and 40% of admissions into hospital are from individuals that are residents of Lambeth, and this does not change much over time. This evidence suggests that even if medical capacity were expanding in Lambeth, it did not lead to a differential treatment of Lambeth residents versus non-residents for the diagnoses related to Class-A drug usage we focus on in our main analysis.22 There are other potential confounding factors to consider. First, if the LCWS policy allowed the policy to reallocate their effort towards crime involving Class-A drugs, then the impacts we have documented would not solely be occurring through any demand side linkage between the use of cannabis and Class-A drugs (whether it arises from unobserved heterogeneity or state dependence via a gateway effect). However as discussed in Section 2, the body of evidence presented in Adda et al. (2013) suggests that the LCWS policy did not lead to a reallocation of police resources towards crime related to Class-A drug crime (rather, the police used the policy to reallocate their effort towards non-drug crime). Hence, in the current study, any link between the depenalization policy and hospitalizations for Class-A diagnoses most likely stems from the interlinkages between the demand sides of the markets for cannabis and Class-A drugs. A second potential confounding factor is that between January 2004 and January 2009 cannabis was declassified from a Class-B drug to a Class-C drug throughout the UK. This declassification effectively decriminalized the possession of small quantities of cannabis for personal use, mirroring the LCWS policy experiment in many ways.23 Such a nationwide policy would obviously only bias the difference-in-difference estimates that we focus on if its impact differed between Lambeth and other London boroughs. To show the policy impacts that we have documented between Lambeth and other London boroughs likely stem from the localized depenalization policy that only operated in Lambeth, we re-estimate our baseline specification (2) using only data running up to 2003 Q4, so up to the point where the nationwide policy change occurred. The result in Columns 1–3 of Table A7 show that for two out of three male age cohorts, there are significant impacts on hospitalization rates for Class-A related diagnoses in Lambeth, even over this restricted post-policy period before any changes in nationwide policy take hold. Residential mobility Throughout the analysis we have used the borough of residence at the time of admission to build hospitalization rates across cohorts. The documented increase in hospital admissions for Class-A drug related diagnoses in Lambeth following the introduction of the LCWS might then operate through two mechanisms: (i) a change in behavior of those residents in Lambeth prior to the policy; and (ii) a change in the composition of Lambeth residents, with the policy inducing a net inflow of people into the borough with a higher propensity for Class-A drug use. Undoubtedly, the geographical distances between London boroughs are small and travel costs are low relative to the fixed costs of permanently changing residence. Similarly the nationwide depenalization policy in place between 2004 and 2009 would further have weakened incentives for individuals to relocate residence with Lambeth in response to the LCWS policy. However, if drug users perceive the depenalization of cannabis in Lambeth as signaling a wider weakening of police enforcement against all illicit drugs, there might be longer term benefits to relocating to the borough. Given the importance of assuming the underlying populations, and hence propensity for drug use, in Lambeth and the rest of London to remain unchanged over time in the difference-in-difference design, we now try to use the administrative records to shed some light on the extent to which drug users relocate their residence into Lambeth from other parts of London as a result of the depenalization policy.24 The HES data contain information on borough of residence for each individual admission to hospital, with individual identifiers allowing us to link patients across episodes and time. The major limitation of using hospital administrative records to shed light on changes in borough of residence in response to the policy, is that for those that are admitted only once during the study period, the data does not allow us to identify whether they have changed residence over time prior to the admission, or will do so subsequent to the admission. These individuals, that form the bulk of hospital admissions and that are included in the main analysis, cannot be included in the analysis below examining migration patterns. While this obviously limits our ability to shed light on the potential net migration into Lambeth of drug users in response to the depenalization policy, we know of no data set representative at the London borough level, that would match both changes in residence over time with individual hospital admissions or health outcomes over time. We therefore proceed by documenting changes in borough of residence for those that have at least two admissions into hospital between 1997 and 2007. To get a sense of the sample selection this induces, we note that in the pre-policy period, 326,683 men are admitted into hospital for any diagnosis, of which 10.6% are re-admitted (at least once) somewhere in London during the one-year period in which the LCWS policy is in place, and 25.3% are re-admitted (at least once) anytime in the post- policy period. Among those 1746 individuals admitted for Class-A drug related diagnosis in the pre-period, only 14.7% are observed being re-admitted for any diagnosis during the policy period, and 28.2% are observed being re-admitted for any diagnosis during the post-policy period. If individuals are induced to migrate to Lambeth in response to the depenalization policy, they might do so at some point during its actual period of operation between June 2001 and July 2002. To check for this, we first focus on those 1630 individuals that are admitted to hospital for any diagnosis in Lambeth during the policy period, and that are observed having at least one prior hospital admission somewhere in London pre- policy. Of these 1630 individuals, 1.7% are admitted for Class-A related diagnosis in Lambeth during the policy period. These are perhaps the most likely individuals to have moved to Lambeth in specific response to the depenalization policy. However we note that among this group, almost all their earlier pre-policy admissions (for any diagnosis) occur in Lambeth, so that there is no strong evidence of these individuals having recently moved to Lambeth during the policy period. While these results focus on those admitted for Class-A drug related diagnosis in Lambeth during the policy period, it might well be the case that drug users that migrate into Lambeth because of the policy are first admitted for some other diagnosis. Hence, we next focus on the 98.3% of hospital admissions in Lambeth during the policy period for any diagnosis unrelated to Class-A drug usage. Among these individuals, nearly all of them are observed with all their earlier admissions in Lambeth; only 10.3% have their last prior admission in some other borough, indicating that they changed their borough of residence at some point between their last admission and the end of the policy period. Taken together, these two pieces of evidence show that among those men with at least two hospital admissions since 1997, there is very limited evidence of there being significant changes of residence into Lambeth during the formal policy period between June 2001 and July 2002. Our next set of results examines longer term patterns of changes in borough residence. Given the fixed costs of changing residence and that in the post-policy period policy enforcement in Lambeth remained somewhat different than other boroughs, it might be reasonable to assume that a net inflow of drug users into Lambeth simply takes some time to occur. To check for this we examine whether inflows into Lambeth from other London boroughs change between two four-year windows: the first four year window occurs entirely pre-policy from April 1997 to April 2001, and the second four year window occurs entirely post-policy from April 2003 to April 2007. In each window we check whether among those admitted to hospital at least twice in the four-year window, and, where at least one admission relates to a diagnosis indicating Class-A drug use, whether changes in borough of residence between the first and last admission vary over time. In the first four-year pre-policy window from 1997 to 2001 we observe: (i) of those that have their first admission outside of Lambeth, 1.4% are observed with a later admission in Lambeth; (ii) of those that have their first admission in Lambeth, 16% are observed with a later admission outside of Lambeth. Doing the same for the later four-year window from 2003 to 2007 to see if this pattern of migration is altered in the longer term, we find that: (i) of those that have their first admission outside of Lambeth, 3.0% are observed with a later admission in Lambeth; and (ii) of those that have their first admission in Lambeth, 30% are observed with a later admission outside of Lambeth. Hence there is evidence of more frequent changes of residence among this subsample post- policy, but that this increase occurs both into Lambeth and from Lambeth: the inflow into Lambeth from other boroughs in the post-policy window relative to the pre-policy window increases (3.0% relative to 1.8%), but this is offset by the percentage increase in outflows from Lambeth to other boroughs among such individuals (30% relative to 16%).25 Overall this suggests is that, among those with multiple hospital admissions, there is increased mobility of residents across boroughs over time, but there is no strong evidence of systematically increased inflows into Lambeth over the second four year window relative to the first.","We now consider four margins of policy impact in more detail: the dynamic responses within age cohorts over time, the heterogeneous impacts within age cohorts by previous admission history, spillover impacts onto hospital admissions for alcohol-related diagnoses, and the severity of hospital admissions. Establishing the existence and magnitude of each effect is important to feed into any assessment of the overall social costs of this localized change in drug enforcement policy related to the market for cannabis. The dynamics of the response ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ On longer term dynamics, Fig. 3 shows that for each cohort there is an inverse-U shaped pattern of dynamic responses across time in the post-policy period. For each cohort the depenalization policy has no significant impact on hospitalization rates during the policy period, and estimated impacts increase thereafter for some time before starting to decline. In line with the evidence in Table 4, the magnitude of the impacts are largest for those in the younger two cohorts aged 15–24 and 25–34 on the eve of the policy. For these age cohorts: (i) the impacts on hospitalizations related to Class-A drug use take a year or two to emerge after the policy is first initiated; and (ii) the post-policy impacts are the highest three to four years into the post-policy period, where the peak impacts correspond to a near doubling of admission rates relative to the pre-policy period for each cohort. Although the pattern of coefficients for the oldest cohort also follows an inverse-U shape, the sign of the point estimates is quite different by the final period considered: 4–6 years into the post-policy period, there is a significant and negative impact on hospitalization rates. This might in part be driven by a different link between the consumption of cannabis and Class-A for those in this age cohort. In comparison to the literature linking policies regulating the market for illicit drugs and public health, all of these dynamic responses are of significant duration. For example, Dobkin and Nicosia (2009) study the impact of a government program designed to reduce the supply of methamphetamine on hospitalization rates (by targeting precursors to methamphetamine), as well as other outcomes. This policy is sometimes claimed to have been the DEA's greatest success in disrupting the supply of an illicit drug in the US and indeed Dobkin and Nicosia (2009) find that the policy had significant impacts on public health. However, they document that these effects were short lived: within 18 months admission rates had returned to pre-intervention levels. In contrast, the depenalization policy we document has an impact on hospitalization rates that lasts at least 3–4 years post-policy for two of the three male cohorts even though the policy itself is only formally in place for a year. Admission histories ~~~~~~~~~~~~~~~~~~~ We next examine how the long run policy impacts are heterogeneous within the same age cohort. To do so, we exploit the full richness of the administrative records to consider differing impacts by individual histories of hospital admission for drug and alcohol- related diagnoses during the pre-policy period from April 1997 to June 2001. This allows us to shed light on whether those with a prior record of heavy substance abuse, respond differentially to the depenalization of cannabis than does the rest of the population. Relative to the existing literature linking drug enforcement policies and health, exploiting this aspect of the data allows us to present novel evidence on the characteristics of the marginal individuals most impacted by a policy of depenalizing cannabis. Examining heterogeneous impacts along this margin is informative because previous heavy users of illicit drugs might be engaged in habitual behaviors so there is less scope for further increases in hospitalization rates for Class-A related diagnoses. Hence the post-policy impacts are measured relative to the period in which the LCWS policy is actually in place. Columns 4–6 in Table 5 consider policy impacts within each age cohort among those that have a prior history of at least one hospitalization for drug or alcohol-related diagnoses. The results suggest that in the longer term such cohorts are either not affected by the depenalization policy, or for the oldest age cohort, their admission rates significantly decline in Lambeth in the long term.26 Such long term users, at least among the two younger cohorts, might be more habituated in their behavior and less price sensitive to any change in the price of cannabis induced by the depenalization policy. If so, this result would be consistent with the evidence based on NLSY97 data in Deza (2011) who uses a dynamic discrete choice model to document that the gateway effect from cannabis to hard drug use is weaker among older age cohorts. An obvious concern with these results is that they might in part be confounded by natural time trends in hospitalizations for Class-A drugs. These time trends might also differ across age groups and by hospital admission histories. To address this, we repeat the analysis but augment Eqs. (4) and (2) with controls for borough specific linear time trends. Table A8 presents the results, again broken down for cohorts based on age and prior admission histories.27 The inclusion of borough specific linear time trends serves to reinforce the earlier conclusions among those without a prior history of admissions (Columns 1–3, Table A8). Among those with a history of admissions, we continue to find no impact among the two youngest age cohorts, although among the oldest cohort the policy now has a positive and significant impact on hospitalization rates.28 Alcohol ~~~~~~~ There is an established body of empirical work examining the relationship between cannabis and alcohol use: this has generated mixed results with some research finding evidence of the two being complements (Pacula, 1998; Farrelly et al., 1999; Williams et al., 2004), and other studies suggesting that the two are substitutes (DiNardo and Lemieux, 2001; Crost and Guerrero, 2012; Anderson et al., 2013), or that there is no statistically significant relationship between the two (Crost and Rees, 2013; Yörük and Yörük, 2013). Many of these studies have identified these impacts among young people, sometimes exploiting minimum legal drinking age that should create discontinuities in alcohol consumption for those aged around 21. We provide a novel contribution to this debate by examining the effect of depenalization on extreme forms of alcohol usage, leading to hospitalizations. We do so for all three male age cohorts. As documented in the raw counts data in Table 2, such admissions occur with far higher frequency for men in all age cohorts, than admissions for either Class-A drugs or cannabis. Hence any positive or negative impacts on admissions for alcohol can have dramatic implications for the monetary health costs of the policy. Throughout, we measure admission rates for alcohol-related diagnoses analogously to those used as our dependent variable in the baseline specifications, Eq. (1). To begin with, we focus on admissions for alcohol-related diagnoses where the primary diagnosis refers to alcohol. We exclude any admission that additionally refers to the use of Class-A substances as the secondary cause of admission. The results in Columns 1 to 3 of Table 6 show that there was a significant reduction in alcohol-related admissions among the youngest cohort in Lambeth relative to the rest of London, but there were no impacts on such alcohol-related admissions for older cohorts. The result suggests that for the youngest age cohort, if depenalization causes the price of cannabis to fall, then alcohol and cannabis might well be substitutes. The next set of specifications probe further to examine the evidence of whether and how the policy impacts the combined use of alcohol and Class-A drugs: here we define admission rates where the primary diagnosis is again for alcohol-related diagnosis, but the secondary diagnosis refers to the use of Class-A drugs. We find no evidence that the policy causes such combined admissions to change in the longer term (and this occurs against a backdrop of London-wide increases in such combined diagnosis admissions). Again, if the depenalization of cannabis caused increased cannabis and Class-A drug use in the longer term, this last set of results supports the assertion that such substances are not being used together, at least among those most prone to extreme abuse of such substances. Severity of hospital admissions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A final dimension along which to consider policy impacts relates to the severity of hospitalizations, as measured by the number of days the individual stays in hospital for, conditional upon admittance. This margin is of policy relevance because it maps directly into the resultant healthcare costs associated with the depenalization of cannabis, as calculated in the next section. We therefore first document how the length of individual hospital episodes for diagnoses related to Class-A drug use changes differentially between Lambeth and other London boroughs post-policy relative to the pre-policy period. To do so, we estimate a specification analogous to Eq. (2) but where the dependent variable is the individual length of hospital stay in days and the sample is confined to episodes where the primary diagnosis relates to Class-A drugs. We focus on the first episode for any hospital stay (that is the same as the entire hospital stay for 93% of observations), and to avoid the results being driven by outliers, we drop observations where the length of the stay is recorded to be longer than 100 days (that excludes a further 2% of all stays).29 As the outcome variable now relates to individual outcomes (rather than borough- quarter-year aggregates), we cluster standard errors by borough to capture unobservables determining the length of hospital stays that are assumed correlated across residents of the same borough. Columns 1 to 3 in Table 7 present the results, again split by age cohort. The data suggests that in the longer term post-policy, across all three age cohorts, the length of stay for Class-A drug related admissions significantly increases in Lambeth relative to the London average. For example, among the 15–24 age cohort, hospital stays increased by 3.7 days, and this is relative to a baseline pre-policy hospital stay length of 7.2 days, an increase of 49%. The proportionate changes for the other age cohorts are 29% for the 25–34 age cohort and 20% for the oldest age cohort. Hence, the proportionate changes in length of hospital stay are greater for the age cohorts that were younger at the time the depenalization policy was introduced. This emphasizes that quite apart from the impacts of the depenalization of cannabis on hospitalization rates for Class-A diagnosis that has been the focus of our analysis so far, the policy also has impacts on the severity of those admissions for Class-A drug use. Both margins are relevant for thinking through the public health costs of the policy as detailed in the next subsection. We note further that the coefficients in the third row of Table 7 show that in other London boroughs there are negative time trends in the duration of such individual hospitalizations conditional on all other controls in Eq. (2). Hence the findings for Lambeth post-policy do not appear to be driven by some systematic lengthening of hospital stays for such diagnosis that might be occurring more generally across London.","We use published the National Health Service estimates of the cost per hospital bed-day.30 This cost is comprised largely of hospital ward costs (nursing, therapies, basic diagnostics and overheads), hence there is actually little variation by diagnosis: the average cost per additional bed-day across all adult inpatient diagnoses categories is £240 (Department of Health, 2012), but the upper end of the hospital bed-day costs, relating to those for adult acute (inpatient) mental health stays, are only slightly higher at £295 (PSSRU, 2011). We therefore use a figure between these estimates, of £250 per hospital bed-day, as quoted by the NHS Institute (2012). This likely represents a lower bound of the true cost of a hospital bed-day because it does not include any specific treatment costs or the additional costs from any associated stay in intensive care.31 We then take our estimates from Tables 2 and 7 for each age cohort to calculate each component of Eq. (7). For the youngest age cohort of those 15–24 on the eve of the depenalization policy, ∆Hbc = 27.2 bed-days per quarter; of this total change, 24.1 bed- days operate through the first channel of increased hospitalization rates, and 4.1 bed- days through the second channel of longer hospital stays conditional on admission. Among those aged 25–34 on the eve of the policy, ∆Hbc = 52.5 bed-days per quarter, where the first channel corresponds to an increase of 42.6 bed-days, and the second component generates an increase of 9.9 bed-days. Finally, for the oldest cohort of 35–44 year olds, ∆Hbc = − 26.9 bed-days per quarter, where the point estimate in the third row of Table 2 implies a decrease in hospitalization rates post-policy of 30.7 bed-days (although this point estimate was not statistically different from zero), and this is only partially offset by the increase through the second channel of 3.8 bed-days. Applying the estimated costs per hospital bed-day of £250 to the change in the total number of hospital bed-days per quarter, for each cohort in Lambeth in the post-policy period on average, reveals the increased public health cost to be: (i) £6802 among those aged 15–24 on the eve of the policy; and (ii) £13,136 among those aged 25–34 on the eve of the policy. Summing across four quarters we derive the conservative public health cost of the depenalization policy to be £79,752 per annum, on average across all the post-policy years in the sample. Aggregating these cohorts across four quarters then suggests the natural decrease in costs associated with hospital bed-days is £4935. Hence the increase in bed-days attributable to the policy more than offsets this natural decrease in hospital bed-days attributable to London-wide time trends. Of course, this calculation still underestimates the total public costs of the increased hospital bed-days within Lambeth due to the policy because of the existence of many additional channels that we have ignored. First, we have ignored any additional demands placed on other parts of the national health service unrelated to hospital inpatient stays, as a result of the depenalization policy. These include demands through outpatient appointments, hospital emergency departments, and through treatment centers. Indeed, the existing evidence from the US on the link between the availability of cannabis and health relate to emergency or treatment costs: Model (1993) find that the de facto decriminalization of cannabis in twelve US states from the mid-1970s significantly increased cannabis-related emergency room admissions. Chu (2012) similarly finds that the passage of US state laws that allow individuals to use cannabis for medical purposes leads to a significant increase in referred treatments to rehabilitation centers. Second, we have ignored any cost to individual users of being hospitalized. Such events almost surely impact individual welfare, especially given the robust association found across countries in the gradient between health and life satisfaction.","We evaluate the impact of a policing experiment that depenalized the possession of small quantities of cannabis in the London borough of Lambeth, on hospital admissions related to illicit drug use. Despite health costs being a major social cost associated with markets for illicit drugs, evidence on the link between how such markets are regulated and public health remains scarce. Our analysis provides novel evidence on this relationship, at a time when many countries are debating moving towards more liberal policies towards illicit drugs markets. We have exploited administrative records on individual hospital admissions classified by ICD-10 diagnosis codes. We use these records to construct a quarterly panel data set by London borough running from 1997 to 2009 to estimate the short and long run impacts of the depenalization policy unilaterally introduced in Lambeth between 2001 and 2002. We find that the depenalization of cannabis had significant longer term impacts on hospital admissions related to the use of hard drugs. Among Lambeth residents, the impacts are concentrated among men in younger age cohorts. The dynamic impacts across cohorts vary in profile with some cohorts experiencing hospitalization rates remaining above pre- intervention levels three to four years after the depenalization policy is first introduced. We combine these estimated impacts on hospitalization rates with estimates on how the policy impacted the severity of hospital admissions to provide a lower bound estimate of the public health cost of the depenalization policy. Our analysis contributes to the nascent literature evaluating the health impacts of changes in enforcement policies in the market for illicit drugs. The depenalization of cannabis is one of the most common forms of such policy either implemented (such as in the Netherlands, Australia and Portugal) or being debated around the world (such as in many countries in Latin America). The practical way in which the localized depenalization policy we study was implemented is very much in line with policy changes in other countries that have changed enforcement strategies in illicit drug markets and as such we expect our results to have external validity to those settings. However unlike those settings, we are able to exploit a (within-city) borough level intervention and so estimate the policy impacts using a difference-in-difference design, as well as exploring differential impacts across population cohorts, where cohorts are defined by gender, age, previous admissions history, and borough of residence. This is different from much of the earlier research that, with the exception of studies based on US or Australian data, can typically only study nationwide changes in drug enforcement policies such as depenalization, and have therefore had to rely on time variation alone to identify policy impacts (Reuter, 2010). The administrative records we exploit allow us to provide novel evidence on how the impacts of such policies vary across population cohorts, over time within a cohort, and how they interact with potential changes of residence of drug users. Clearly, such policy impacts are unlikely to ever be estimated using randomized control trial research designs. We have used a difference-in-difference research design exploiting an unusual policy experiment in one London borough that allows us to exploit within and across borough differences in health outcomes to identify policy impacts. The key concern with such a research design is to distinguish policy impacts from time trends. To do so, we have used the detailed administrative records to present evidence on how different cohorts (by gender, age and previous admission history) are differentially impacted by the policy, how the results are strengthened when controlling for time trends, and checked for the presence of trends in the pre-policy period. Our results suggest policing strategies have significant, nuanced and lasting impacts on public health. In particular our results provide a note of caution to moves to adopt more liberal approaches to the regulation of illicit drug markets, as typically embodied in policies such as the depenalization of cannabis. While such policies may well have numerous benefits such as preventing many young people from being criminalized (around 70% of drug-related criminal offenses relate to cannabis possession in London over the study period), allowing the police to reallocate their effort towards other crime types and indeed reduce total crime overall (Adda et al., 2013), there remain potentially offsetting costs related to public health that also need to be factored into any cost–benefit analysis of such approaches. Two further broad points are worth reiterating. First, our analysis relates to the more general study of the interplay between the consumption of different types of drug. In particular there is a large literature testing for the “gateway hypothesis” that the consumption of one “soft” drug causally increases the probability of subsequently using a “harder drug”. The crucial challenge for identification is the potential for unobserved factors or heterogeneity that could drive consumption of multiple types of drug. Existing work has tried to tackle this problem by either: (i) instrumenting the gateway drug with a factor unrelated to the underlying heterogeneity, typically using cigarette and alcohol prices (Pacula, 1998; DiNardo and Lemieux, 2001; Beenstock and Rahav, 2002); or, (ii) using econometric techniques to model the possible effects of unobserved heterogeneity (Pudney, 2003; van Ours, 2003; Melberg et al., 2010). To be clear, in our analysis we make no attempt to test for gateway effects directly, but our contribution to this literature is to demonstrate that the markets for cannabis and hard drugs are concretely linked — be it because of gateway effects or some other channel — so that changes in policy that affect one market will have important repercussions for the other (DeSimone and Farrelly, 2003; van Ours and Williams, 2007; Bretteville-Jensen et al., 2008). Finally, our analysis highlights the impact that policing strategies can have on public health more broadly. It is possible that other policing strategies, such as police visibility or zero-tolerance policies, could also have first order implications for public health. These effects could operate through a multitude of channels including: (i) police behavior directly impacting markets and activities that determine individual health, such as the case studied in this paper; and (ii) police behavior affecting perceptions of crime and thus influencing psychic well- being. This possibility opens up a rich area of further study at the nexus of the economics of crime and health."],["We consider a two period model in which an incumbent political party chooses the level of a current policy variable unilaterally, but faces competition from a political opponent in the future. Both parties care about voters' payoffs, but they have different beliefs about how policy choices will map into future economic outcomes. We show that when the incumbent party can endogenously influence whether learning occurs through its policy choices (policy experimentation), future political competition gives it a new incentive to distort its policies - it manipulates them so as to reduce uncertainty and disagreement in the future, thus avoiding facing competitive elections with an opponent very different from itself. The model thus demonstrates that all incumbents can find it optimal to 'over experiment', relative to a counterfactual in which they are sure to be in power in both periods. We thus identify an incentive for strategic policy manipulation that does not depend on parties having conflicting objectives, but rather stems from their differing beliefs about the consequences of their actions. --------------------------------------------------------------------------------","Many of the most important public policy problems democratic countries face require cumulative efforts by successive governments to be successfully managed. Consider environmental policy (in particular regulation of stock pollutants such as greenhouse gases), social security reform, sovereign debt management, and public infrastructure development. None of these issues can be tackled in a single legislative term, and the total quantity of resources devoted to them will likely be the result of decisions taken by several governments. As such, the policies incumbent political parties choose to address these issues are heavily influenced by the incentives that the political system provides for them to make sound ‘long-run’ policy decisions, even if the effects of those decisions may only be realized once they have left office. The lack of future political control that is characteristic of democratic systems means that, for the purposes of setting ‘long-run’ policies, incumbents have incentives to manipulate their current policy choices so as to influence both who gets elected in the future and the policy choices future governments will make (Persson and Svensson, 1989; Aghion and Bolton, 1990; Tabellini and Alesina, 1990; Milesi-Ferretti and Spolaore, 1994; Besley and Coate, 1998; Persson and Tabellini, 2000; Azzimonti, 2011). These strategic incentives exist even if parties are not purely office seeking, but have interests that coincide with those of a group of voters, e.g. in models of partisan politics. These effects have traditionally been studied in models with heterogeneous preferences: parties are assumed to have intrinsically different preference parameters, which induce heterogeneous preferences over policies, and hence a strategic incentive for an incumbent party to manipulate present policy choices given that its reelection is uncertain. While heterogeneity in preference parameters undoubtedly accounts for some of the divergences between political parties’ preferred policies, heterogeneity in beliefs is likely to be an equally important factor. Milton Friedman famously argued that “differences about economic policy among disinterested citizens derive predominantly from different predictions about the economic consequences of taking action…rather than from fundamental differences in basic values” (Friedman, 1966). More recently, public surveys in the US demonstrate a strong polarization in the beliefs of Democrats and Republicans about a variety of policy issues, including, for example, the likely causes and severity of climate change (Leiserowitz et al., 2012; Borick and Rabe, 2012). Despite the empirical plausibility of belief heterogeneity, the consequences of relaxing the common prior assumption have been largely unexplored in the political economy literature on strategic policy choice.2 The crucial new feature of political competition induced by heterogeneous beliefs is that beliefs are dynamic, and potentially endogenous. Parties' policy preferences may change over time as their beliefs evolve in response to new information. Moreover this learning process may, at least to some extent, be under the control of the incumbent, who may choose policies with the express purpose of revealing information about their consequences in the future; learning may be ‘active’. Active learning — the idea that current policy choices influence how much is learned in the future — is an old concept in economics (e.g. Prescott, 1972; Grossman et al., 1977), which has been applied to problems in monetary policy (Bertocchi and Spagat, 1993), environmental regulation (Kelly and Kolstad, 1999), and firm behavior (Keller and Rady, 1999). It can be seen as a form of experimentation — we choose an action, observe its consequences, and so learn something new about the relationship between choices and outcomes. In addition, it is often the case that the more intensely we pursue a policy, the more we can separate the ‘signal’ from the ‘noise’, and the more we learn about its effects.3 Thus when learning is active, and parties have divergent beliefs that they update rationally, the incumbent party has a measure of control over its own, and its opponent's, future policy preferences. This gives rise to strategic incentives for policy manipulation that are entirely absent when parties merely have different preference parameters. Our core contribution is to elucidate the interaction between belief heterogeneity, active learning (or experimentation), and political competition, and how this affects the size of public programs with uncertain deferred benefits (or costs). We focus on how the interaction between these factors determines an incumbent's response to the intertemporal tradeoff inherent in such problems. We thus abstract from questions of taxation and redistribution, and consider a stylized model in which voters differ only in their beliefs about the benefits of the policy, and parties that represent the beliefs of groups of voters must decide only on the level of some policy variable. We show that the interaction between active learning and political competition gives rise to a new incentive for incumbents to distort their policy choices. This incentive pushes incumbents to choose policies that increase their chances of resolving uncertainty in the future, regardless of their beliefs: they will over experiment. The intuition behind this result is simple — since the preferences of parties with different a priori beliefs converge when learning occurs, incumbents avoid future competitive elections with an opponent very different from themselves by choosing policies that reduce disagreement. We demonstrate this mechanism in a two period model that combines the literature on intertemporal decision making under uncertainty and learning (Arrow and Fisher, 1974; Henry, 1974; Epstein, 1980; Gollier et al., 2000), with a simple but flexible model of political competition (Wittman, 1973, 1983; Roemer, 2001). To demonstrate the effects cleanly, the model assumes that parties care only about the voters' well-being, and disagree only in their beliefs. Thus, in the absence of belief heterogeneity all parties in our model would agree on the correct policy choice, which would also be the optimal policy for the voters. Yet even in the sanguine case where parties are well intentioned and have common objectives, heterogeneous beliefs and political competition will distort their policy choices. We show that when learning is active enough, all incumbents will over-experiment relative to a counterfactual in which they are sure to be in power in the future, regardless of their beliefs and the beliefs of their political opponents. Section 2 sets out the model structure. Section 3 examines how the interaction between active learning and political competition affects policy choices when beliefs are heterogeneous, without specifying the actual form of the political competition between parties. To build intuition, a simple model with binary policy choices is discussed first, followed by a more complex model with continuous policy choices. Section 4 specializes to a specific model of political competition: the Wittman model. In our version of this model parties know the distribution of the voters' beliefs, voters vote for their preferred platform, and elections are decided by majority rule. We show that our results hold under plausible primitive conditions on the parties' payoff functions in this case, which apply in both ‘full commitment’ and ‘no commitment’ versions of the model. We reflect on the application of our results to a variety of policy issues in Section 5, before concluding. Related literature ~~~~~~~~~~~~~~~~~~ While the consequences of heterogeneous beliefs and strategic experimentation for the policy choices of incumbents are (to the best of our knowledge) unexplored, several papers investigate some of these factors in other contexts. Piketty (1995) considers a model of social mobility and redistributive taxation, in which agents hold different beliefs about the relative importance of effort and social class in determining economic outcomes. The beliefs of different agents are updated based on their income mobility experience, and transmitted to their descendants. Piketty shows that belief heterogeneity persists in the steady state, and that experience of income mobility, and not simply income level, contributes to forming political attitudes. While heterogeneous beliefs are at the core of this work, it focusses on the voters' belief formation processes, and not on strategic policy experimentation by incumbent governments. Strulovici (2010) is explicitly concerned with strategic experimentation, but focusses on strategic voters, rather than strategic parties. In his model pivotal voters recognize that experimentation reduces their likelihood of being pivotal in the future — this results in under-experimentation in equilibrium. We focus on the behavior of strategic parties that manipulate their current policies in part to influence the beliefs of future voters. In contrast to Strulovici (2010), we show that when parties have good faith disagreements with their political opponents, they have an incentive to over experiment. Callander and Hummel (2013) consider a model that is in some respects close to ours. They examine the efficiency of political turnover, when the only link between successive governments is the information they possess. Incumbents can experiment strategically to influence the information that their successors will use to make their policy choices. They show that, due to the time inconsistency issues that are inherent in political systems with turnover, experimentation can improve the efficiency of policies, as it creates a channel for intertemporal influence. This informational channel of influence is also present in our work, but the political context differs. Parties have common beliefs but heterogeneous objectives in their model, and political turnover is imposed exogenously. By contrast, parties in our model have common objectives but heterogeneous beliefs, and the identity of future governments is determined endogenously via competitive elections. This allows us to study the interaction between endogenous political competition, policy experimentation, and heterogeneous beliefs. Finally, Hirsch (2013) considers a model of political organization in which a principal and an agent disagree about which policy to implement, but share the same objectives, and can engage in experimentation. Hirsch shows that it may be optimal for the principal to defer to the agent to motivate him to act, or to demonstrate to the agent that his beliefs are incorrect. While the fact that agents in his model differ only in their beliefs is common to our analysis, the roles of the players are exogenously assigned in his work. His model focusses on strategic delegation in a hierarchical organization, rather than strategic interaction between political parties. Despite the differences in context between our work and that of the last three papers mentioned above, a common overarching theme unites them. In all these cases learning provides a channel for influence, which is used to the advantage of a ‘first-mover’. In Strulovici (2010) this is the pivotal voter, in Hirsch (2013) it is the principal, and in Callander and Hummel (2013) and our own work, it is an incumbent government. Thus our work contributes to a wider recent research program which sees information as a source of strategic control in a variety of contexts.","We consider a two period model, and assume two political parties, indexed by i ∈ {G, B}. The parties are well-intentioned: they care only about the voters' well-being, and don't seek office for their own ends. Our choice of labels for the parties is motivated by an environmental interpretation of the model (‘G’ = Green, ‘B’ = Brown) which we will use to provide intuition at several points in the exposition, but the model is applicable much more widely. In the first period the incumbent party sets some policy variable e1, which gives rise to certain first period payoffs U(e1). Second period payoffs W(e2|e1, λ) depend on the policy e2 that is implemented in the second period, on the legacy of first period policy choices e1, and on an a priori uncertain parameter λ, which affects the optimal second period policy. The conditions we impose on U and W will be discussed below, for now we focus on the model structure. The function B(e) denotes known short-run benefits from industrial processes that emit the pollutant, and C(e1 + e2) denotes long-run costs (e.g. health impacts or productivity losses) resulting from the accumulation of the pollutant in the atmosphere. The magnitude of these future costs is uncertain, and depends on the realization of λ. Returning to our general model exposition, we assume that λ ∈ {λL, λH}, where λL < λH. The crucial feature of our model is that parties and voters have heterogeneous beliefs about the consequences of policy choices. In the first period, party i believes that λ = λL with prior probability qi. We assume without loss of generality that qG < qB. In our environmental example, this implies that the Green party puts more subjective weight on the ‘high damages’ state λ = λH than the Brown party — hence their labels. The voting population's beliefs are also heterogeneous, and each party's beliefs are assumed to be representative of some exogenously given subset of voters. The heterogeneity in the parties' beliefs is the only difference between them. A∗(e1, q) is thus the payoff a party with beliefs q expects to receive in the second period if the value of λ remains unknown in the future, and it has exclusive control over which second period policy is implemented. e2∗(e1, q) is the policy this party would choose in this situation. The parties are dogmatic, in that they do what they think is best for the voters given their beliefs qi, and don't account for the beliefs of those who disagree with them when making their policy choices. They are however rational, and realize that in the future new observations may be realized that provide information about the value of λ. They will interpret this new evidence in a rational Bayesian fashion, and update their priors. Moreover, each party knows that the other party will do the same. We compress this incremental learning process into a single period. To keep the learning process simple we assume that in the second period either the true value of λ is revealed (with probability f(e1)), or nothing is learned about the value of λ (with probability 1 − f(e1)).4 Crucially, we allow the probability of learning to depend on first period policies. If f′(e1) > 0 then learning is active — the more intensive are first period policies, the greater the chance of learning the value of λ in the second period. In this case, policy experimentation carries an informational payoff. Alternatively, if f′(e1) = 0, we say that learning is passive: policy choices have no informational consequences. Fig. 1 illustrates the timing of events in the model. At the beginning of the first period an incumbent chooses a policy e1. At the end of the first period either the true value of λ is revealed (with probability f(e1)), or nothing is learned (with probability 1 − f(e1)). If λ is revealed, the parties' policy preferences are identical in the second period — there is no difference between them as they hold the same beliefs. In this branch of the decision tree there is a ‘trivial’ election in the second period — it doesn't matter who gets elected, as both parties will choose the same policy. If however λ is not revealed, the parties' beliefs remain divergent in the second period. In this case even though the parties have common objectives, they offer different platforms, reflecting their different priors. Thus each party announces a policy platform e2i at the beginning of the second period, and voters decide between them in competitive elections.","Our main hypothesis is that the interaction between active learning and political competition gives rise to incentives for incumbents to ‘over experiment’ with their first period policies. This reduces uncertainty and disagreement in the future, and hence avoids costly political competition. In order to demonstrate this in our model, we need to examine the additional effects of active learning, political competition, and their interaction, on policy choice. Thus, we need to define baseline learning and political scenarios which we will compare to the active learning/political competition scenarios. We have thus set up two dimensions of variation in our model — passive vs. active learning, and political competition vs. the individual optimum. Evaluating the optimal policies in (10) and (13) under the two learning scenarios (11) and (12) leads to four policy scenarios. Table 1 summarizes our notation for the optimal first period policies in these four cases. The passive learning/individual optimum cases allow us to determine the additional effects of active learning/political competition, relative to these baselines. The interaction between active learning and political competition is captured by looking for differences between the effect of active learning (relative to passive learning) in the two different political scenarios. A simple model with binary policy options ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In order to build intuition, we consider a simple version of the above model in which the first period policy e1 can take only two values: e1 ∈ {0, 1}.6 The incumbent must either implement a policy (e1 = 1), or do nothing (e1 = 0) in the first period. Second period policies e2 may be discrete or continuous — all that we require is that optimal second period policies e2∗(e1, q) depend on the value of q. Active learning gives any incumbent party an additional incentive to experiment (i.e. choose e1 = 1) relative to the passive learning case, in both the individual optimum and political competition scenarios. However, this additional incentive is greater under political competition than in the individual optimum. The first of these inequalities follows from the convexity of the ‘max’ function (information has positive value), and the second from (9). Thus there is a greater incentive to choose e1 = 1 when learning is active than when it is passive when the incumbent faces political competition in the future. Similarly, there is a greater incentive to choose e1 = 1 under active learning (relative to passive learning) when the incumbent is certain to be in office in both periods. This result says that the difference between the incumbent's incentive to choose e1 = 1 under active vs. passive learning is larger when it faces political competition than in its individual optimum. There will thus be cases in which switching from passive to active learning induces the incumbent to switch from e1 = 0 to e1 = 1 under political competition, but not in the individual optimum. The converse, however, can never happen. If a switch from passive to active learning causes the incumbent party to change from e1 = 0 to e1 = 1 in the individual optimum, it must also do so under political competition. This simple result illustrates the incentive for the incumbent party to over experiment when it faces political competition from an opponent who shares its goals, but has differing beliefs. While active learning provides an additional benefit (relative to passive learning) to the e1 = 1 policy under both political scenarios, the difference between the relative benefits of e1 = 1 under active and passive learning is greater under political competition than in the individual optimum. This is so since, under active learning, the incumbent party increases its chance of avoiding an election with an opponent different from itself by choosing e1 = 1. It chooses its first period policy strategically to reduce disagreement in the second period. This result relies critically on the fact that the parties have heterogeneous beliefs. Beliefs are endogenous and amenable to manipulation, whereas preference parameters are not. Continuous first period policies ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The positive interaction between active learning and political competition is easily demonstrated in the binary case examined in Section 3.1. We now extend these results to a continuous model of policy choice. This turns out to be a more complex problem, for the following reason. The results we obtained in the binary model only required us to rank payoff levels under the different scenarios. When first period policies e1 are continuous (and payoffs W are non-linear in e1), the comparison between optimal first period policies under different learning and political scenarios involves not only the levels of second period payoffs, but also the derivatives of these payoffs with respect to e1. This additional complexity has long been recognized in the literature on the effect of learning on dynamic choice (e.g. Epstein, 1980; Ulph and Ulph, 1997; Gollier et al., 2000), which has focussed on conditions that are sufficient to determine the direction of the change in the optimal choice variable under different learning scenarios. Following in this tradition, we will state sufficient conditions for an analogue of the intuitive results obtained in the binary case to hold in a continuous model. To understand the conditions in the proposition, it is helpful to begin by examining a special case. Suppose that W(e2|e1, λ) is independent of e1. In this case the only way e1 influences second period payoffs is through the effect it has on the probability of learning f(e1); it does not directly affect the parties' payoffs in the second period. Thus the only linkage between the periods is informational. In this case the conditions (18–19) are satisfied as equalities as their constituent terms are all identically zero, and condition (20) is satisfied for any strictly increasing f(e1), as its right hand side is zero. Thus the conclusions of the proposition hold identically in this case (with political competition having no effect on e1 under passive learning in conclusion (b)). In words, second period payoffs under political competition when λ is unknown are always less than payoffs in the individual optimum when λ is unknown, which are in turn always less than payoffs when λ is known in the second period. These relationships imply the pattern of effects we observed in the binary policy case (we used them in (15–17)), and these effects carry over to the continuous policy case when information is the only linkage between the two periods. The core insight is that in this special case our intuition for how the interaction between active learning and political competition, which was based on comparisons of the levels of payoffs under different scenarios, is undisturbed by the derivatives of payoffs. Comparing these inequalities to those in (21), we see that the conclusions of Proposition 2 hold if the derivatives of second period payoffs with respect to e1 are ranked in the same way as the levels of the payoffs. While the message of the proposition is clear, the conditions (18–19) depend on endogenous quantities, and it is thus not possible to know when they are satisfied without putting more structure on the problem. This is a common feature of learning models (see e.g. Epstein, 1980). In the next section we consider two common models of political competition, and, in each, find primitive conditions on the payoff function W(e1|e2, λ) that ensure that the crucial conditions (18–19) hold.","In the model we consider, the voters' choices depend only on the platforms parties announce (i.e. they don't have a party affiliation), and the distribution of their beliefs is known to both parties. We consider two variants of the model — a full commitment case, and a no commitment case — and show that the same primitive conditions on the payoff function W(e2|e1, λ) imply that the conditions of Proposition 2 are satisfied in both cases. A median voter model (full commitment) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Thus when voters' beliefs are known and they have single peaked preferences, parties' platforms converge completely in the second period — they both offer the median voter's optimal policy.9 With this expression for the equilibrium value of the ‘no learning’ sub- game, we can seek conditions on the payoff function W which ensure that (18–19) of Proposition 2 are satisfied. The next result provides such conditions, without specifying a parametric form for W. (26) is the standard concavity condition, and in addition we assume that solutions to the second period optimization problem are interior, so that the constraint e2 ≥ 0 is not binding. This assumption simplifies our analysis, but our results are not crucially dependent on it. It is readily shown (see Appendix C) that the conditions (27) and (28) imply respectively that the optimal second period policy e2∗(e1, q), is decreasing in e1, and increasing in q. In our environmental example this implies that the greater is the level of first period emissions, the less parties want to emit in the second period, and similarly, the greater the weight they put on the low damages state λL, the more they want to emit in the second period. This property has important consequences for the condition (18), which gives rise to conclusion (b) of Proposition 1: under passive learning any incumbent reduces e1 under political competition relative to its individual optimum. This is really the novel condition of the proposition, as the other condition (19), which guarantees that active learning causes incumbents to increase e1 relative to passive learning, is well known; it can be seen as a special case of the sufficient conditions for signing the effect of learning on policy choice derived in Epstein (1980). Conclusion (b) is novel, so it is important to understand how the properties of the payoff function in Proposition 3 give rise to it. This can be seen by thinking about the strategic consequences of (34) in the political competition scenario, as we now explain. Thus from the inequalities (35–36) and (37–38), we see that if we increase e1, the distance between the median voter's optimum and either parties' optimum increases. However, reducing e1 brings the median voters' optimum closer to both of the parties' individual optima. Since it is the median voters' optimum that is implemented under political competition, and all parties have single peaked preferences over second period policies, all parties want this policy to be as close to their individual optima as possible. Fig. 2 illustrates this intuition graphically. The condition (34) thus ensures that regardless of whether the incumbent party's beliefs qi are greater or less than qm, it always has a strategic incentive to reduce e1 relative to its individual optimum. Exogenous election probabilities (no commitment) ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We can treat all these cases at once by allowing the probability of election to be an arbitrary constant. We have the following proposition: Suppose that the conditions on U and W inProposition 3are satisfied. Assume that the outcome of the political process is exogenously determined, so that πi(e2i, e2j) is an arbitrary constant in [0, 1]. Then the conclusions ofProposition 3continue to hold. Thus the results in Section 4.1, in which the election outcome was endogenously determined by the parties' platforms and parties were assumed to commit, carry over to the case in which election outcomes are exogenous (independent of the parties' platforms) and parties cannot commit.","The mechanism we have identified can be applied in many policy contexts. Before we discuss some examples however, we emphasize that we see the effect we have highlighted as only a partial contributing factor to actual policy outcomes. Our purpose has been to highlight an informational channel of intertemporal influence and its effects on the policy choices of incumbents. The model is intentionally idealized, in order to demonstrate the new incentives for policy manipulation that belief heterogeneity gives rise to. Previous literature (e.g. Persson and Svensson, 1989; Aghion and Bolton, 1990) examines strategic policy choice when parties' objectives are heterogeneous but they have common beliefs; our work assumes the opposite. The real world, of course, falls between these two stark cases. As a first example of the application of our mechanism, consider the case of public smoking bans. The benefits of reducing second-hand smoke include lowering the risk of lung cancer, reducing health care costs, and improving worker productivity. The costs of bans are born by restaurants and bars who may see a decline in profits. These costs are uncertain before the policy has been implemented. Indeed, the impacts of smoking bans on the restaurant business are still debated (Hyland et al., 1999; Adams and Cotti, 2007). Given this uncertainty, and conflicting sources of information, different public representatives are likely to hold different beliefs about the consequences of these laws. Despite conflicting views, public officials have proven willing to experiment with smoking bans, as their diffusion from a few cities and states (San Luis Obispo, California in 1990, California in 1998, New York City in 2002) to many states and countries illustrates (Adams and Cotti, 2007; Eriksen and Chaloupka, 2007). Congestion charges provide another application. These taxes provide an immediate social benefit by reducing traffic congestion and improving air quality in city centers. However, they also impose a priori uncertain costs on residents who commute to the city center by car, and hence on centrally located retail businesses. Despite disagreements about the projected impacts of these policies, city administrations have rolled them out. The charges are usually introduced in an initial experimental phase, which enables affected parties to learn about their consequences. This occurred in Stockholm in 2007, when the charge was implemented for an initial seven month trial period, before being expanded and made permanent (Winslott- Hiselius et al., 2009). Similarly, the London congestion charge in 2003 was adopted after an 18-month public consultation period (Leape, 2006). A further example is provided by policies that aim to regulate national emissions of atmospheric pollutants. Policies that place an annual cap on emissions, such as the cap-and-trade mechanisms of the US Acid Rain Program (established in 1990 to curtail sulfur dioxide and nitrogen oxide emissions), or the European Emissions Trading Scheme (established in 2005 to reduce greenhouse gas emissions), may reduce the future damages that arise from the accumulation of pollutants. However, different political actors hold different beliefs about the magnitude of these damages, which are highly uncertain (Borick and Rabe, 2012). Since forcing the atmospheric system with emissions allows us to observe how it responds, emitting more today helps to reduce uncertainty in the future (Kelly et al., 2005). Thus our model suggests that even well-intentioned incumbents have an incentive to set less stringent caps than they would prefer to, so as reduce uncertainty and disagreement in the future. Unlike more conventional explanations of the difficulty of passing stringent abatement policies, which appeal to partisan motives, collective action problems, and the influence of special interest groups, this explanation does not require political parties to act solely in their self-interest; it is simply a consequence of disagreements about matters of fact.","Our analysis has identified a novel incentive for incumbents to manipulate their policy choices, when beliefs are the primary source of disagreement between political parties. This stems from the interaction between active learning — the ability to endogenously influence future information revelation through current policy choices — and political competition. When learning is active incumbents can control the degree of disagreement in the future. Since the incumbent party avoids a costly election with an opponent very different from itself if information is revealed and beliefs converge, it has an incentive to increase the chances of resolving uncertainty in the future, regardless of its initial beliefs. This effect relies crucially on the fact that, unlike intrinsic preference parameters, beliefs are endogenous, and thus subject to manipulation. The mechanism we identify is widely applicable, and may have explanatory power for any policy issue about which voters and parties have diverse beliefs, and where learning occurs through policy implementation. Unlike conventional explanations of distortions to public good provision due to the influence of politics, our result does not rely on rent-seeking and the influence of special interest groups (Aidt and Dutta, 2007; Bohn, 2007; Battaglini and Coate, 2008), or the institutional structure of government (Persson and Tabellini, 1999; Lizzeri and Persico, 2001; Acemoglu and Robinson, 2001). It can be thought of as a parsimonious causal mechanism that assumes the best of political actors, yet predicts that they will still do more to reduce uncertainty than they would like to."],["This paper reports on an experiment designed to test whether people's preferences change to become more alike. Such preference conformism would be worrying for an economics that takes individual preferences as given (‘de gustibus es non disputandum’). So the test is important. But it is also difficult. People can behave alike for many reasons and the key to the design of our test, therefore, is the control of the other possible reasons for observing apparent peer effects. We find evidence of preference conformism in the aggregate and at the individual level (where there is heterogeneity). It appears also to be more consistent with Festinger's epistemic account of why it might occur than that of Social Identity Theory. --------------------------------------------------------------------------------","By suggesting in Middlemarch that people conform to the behavior of their own group so as to be able to distinguish themselves from members of other groups, George Eliot appears to anticipate Social Identity Theory (see Tajfel and Turner, 1979; Akerlof and Kranton, 2000). In so far as such conformism arises because people's preferences become alike, her suggestion is also potentially troubling for economics. In welfare economics, for instance, it is the identification of an individual with their preferences that makes the Pareto criterion appealing. To the extent that individual preferences are endogenous and bend to those of others in some process of preference conformism, the appeal of the Pareto criterion is diminished. Likewise in both micro and macroeconomics, individual preferences are often taken as the bedrock upon which economic modeling proceeds because they are assumed not to change: ‘de gustibus non est disputandum’ (see Stigler and Becker, 1977; Lucas, 1976).1 This paper is concerned with whether there is evidence that people exhibit preference conformism. As Social Identity Theory attests, the possibility of motivational conformism has long been important in the other social sciences. David Riesman's famous book from the 1950s, The Lonely Crowd, that gained him a place on the front of Time magazine, was based on the diagnosis that modern societies, and particularly the US, were distinguished by the rise of a conformist attention to one's peers. In making this argument he was developing a point first made by de Tocqueville about the US. Hannah Arendt's equally, if not more, famous book from the 1950s, The Origins of Totalitarianism, made conformism a part of the explanation of how ordinary people were swept along by the rise of the Nazi party. Edward Said's celebrated late 20th century book, Orientalism, turns in part on a similar observation to that of Elliot: that individuals identify themselves through the contrast they find (or imagine) with those outside their group. Whether preference conformism occurs also matters for some evolutionary accounts of group selection. This is because, when individuals conform to the behavior of their group, the differences in individual behavior, on which selection turns, become the same as those that distinguish groups (see Boehm, 1999). If for these reasons, it is important to know whether people exhibit preference conformism, it is also difficult. This is because people behave alike for many reasons and not just because people's preferences become alike. People often open their umbrellas when it rains – a common event triggers the similarity in behavior without any change in preferences. It also pays to do what others do when playing a coordination game. Equally people may extract information about the state of the world from the actions of others in what proves to be an information cascade yielding similarity in behavior (e.g., see Bikhchandani et al., 1992). Again, even though such cascades can sometimes lead one astray, there is no doubt that it can be rational, in the sense of acting upon exogenously given preferences, to be guided by others in this way. Individuals may also act alike when they have a preference for status or social approval that comes from behaving in accordance with a social norm (see Jones, 1984; Bernheim, 1994). A preference for following ‘the fashion’ can have the same effect, as can a social preference like inequality aversion in some settings (see Lahno and Serra-Garcia, 2015). The challenge in testing for preference conformism, therefore, is to disentangle these possible ‘untroubling’ sources of similarities in individual behavior. This is often difficult (see Angrist, 2014) and explains why we have adopted an experimental approach. A suitably designed experiment can, in principle, control for the ‘untroubling’ causes of conformism. If conformist behavior is still revealed, then it points to the troubling kind of preference conformism. Our experimental design has six novel features for this purpose. First, we focus on non-strategic decision problems to avoid the possibility that any similarity in behavior is triggered by reciprocation.2 There are some experiments that examine whether people follow norms in social dilemmas (e.g, see Kimbrough and Vostroknutov 2016; Carpenter, 2004; Gachter et al, 2013). Such norm-following is potentially relevant because it can be interpreted as social preference conformism. However, such behavior is also consistent with a simple form of reciprocation and, to avoid this confound, we focus on non-strategic decisions where reciprocation does not naturally arise. The use of non-strategic decisions also helps avoid another possible confound: the in-group bias in social preferences that has been revealed in social dilemmas (e.g. see Chen and Li, 2009; Hargreaves Heap and Zizzo, 2009) might also produce apparently conformist behaviors. There are some non-strategic decision experiments that examine peer effects. Cason and Mui (1998) find that peer information has an effect in a dictator game, but this information appears to be a prime for social preferences and not for conformism as such. There are also experiments involving choices over independent lotteries, where there is evidence of peer effects. That is, information about what others have chosen (which cannot convey any information about the state of the world) appears to influence individual lottery choices (e.g., see Cooper and Rege, 2011; Goeree and Yariv, 2015; Lahno and Serra-Garcia, 2015; Gioia, 2017).3 The peer effects are mixed in the sense that some subjects follow and some avoid what others do. Those who follow typically predominate (although see Corazzini and Greiner, 2007, where non-conformists predominate, and Duffy et al., 2015, who find that, when subjects must choose private or social information and the optimal choice varies, there is a balance between the ‘lone wolves’, who err on the side of private, and the, ‘herders’, who err on the side of social information). When people follow what others do in these experiments it could be because they have a preference for sharing in misery or inequality aversion rather than because they have some tendency to preference conformism. The evidence on this is again mixed. While Corazzini and Greiner (2007) and Goeree and Yariv (2015) reject the inequality aversion interpretation, Cooper and Rege (2011) find support for the sharing in misery interpretation (what they refer to as a form of ‘social regret’ motivation) but Gioia (2017) rejects this possibility.4 Our next feature is a response to this ambiguity over interpretation. Second, we have a range of different types of non-strategic decision problems. In addition to the lottery decisions discussed above, there are dictator-like distribution decisions and choices over ordinary goods. This is new and important because social regret or inequality aversion might apply to one type of problem, but it would not obviously apply to all problems. Thus if we find that individuals exhibit conformism across all types of decision problem, then, on Ockham's Razor grounds, it is more likely that that they have a general tendency towards preference conformism rather than several idiosyncratic preferences which explain their conformism across all types of decision problem. Third, the decision problems are chosen so that the value of the options do not depend in any obvious way on a state of nature that might be revealed by other people's decisions. This avoids the possibility that behavior follows an information cascade. Fourth, we only give individuals peer group information in the form of what someone in their group has done in the past. This avoids peer information being developed endogenously within the laboratory and so gives control. It also reduces the possibility that decision-specific preferences like social regret and inequality aversion might be triggered because the peer information refers only to ‘one person’. The peer information is in this sense weak. The next design feature adds to the slightness of the group information and so contributes, with this feature, to making this a ‘tough’, in Popper's sense, test of preference conformism.5 Fifth, we assign individuals to either the ‘red group’ or ‘blue group’ (e.g. see Hargreaves Heap and Zizzo, 2009). The groups are artificial and this is likely to weaken any influence that they have on individual behavior. It also means that the likelihood of subjects drawing in an habitual manner on a desire for social approval or status from behaving in accordance with a social norm that might come from membership of the same natural group outside the laboratory is much weaker. There are, for instance, no ‘red group membership’ habits outside the lab that might be drawn upon unconsciously inside. Six, individual decisions remain private. This militates against the generation of any feelings of social approval or status within the lab because they depend on others knowing about your actions (although you may still derive a sense of self-image). The two group feature of our design is also noteworthy in two respects. It allows us to test for whether any influence of peer information on behavior in the lab arises through an experimental demand effect. We explain this in more detail below. The feature also enables us to distinguish between two accounts of preference conformism in psychology. In Social Identity Theory, it may seem puzzling that individuals gain a sense of identity by acting like other members of their group, but not when there is more than one group and groups behave differently. Individuals acquire a distinct (social) identity through the differences in the way that groups behave. As a result, individuals not only pay heed to what others in their group are doing, they also take a cue on how to behave from those outside the group by avoiding in some degree what they do (see Gino et al., 2009, for experimental support and, from popular culture, Dr Seuss's ‘Sneetches’). Against this view, Festinger (1954) interprets the attention to what others do differently: its origins are epistemic and come from an individual drive to evaluate one's uncertain opinions (and abilities). We call this ‘cognitive appraisal’. It arises when a gap between one's own tentative opinion/tastes and that of others suggests a possible incorrectness or weakness in them which leads one to revise them in the direction of what others value. He argues, on the basis of a variety of experiments, that the strength of this social influence depends on the strength of a person's ties to their group. Thus what people in another group do would be less relevant than that of one's own group, but unlike Social Identity theory, their behavior would not be a cause for doing different.6 In the next section, we describe the experiment in more detail and set out the hypotheses that we test. Section 3 gives our results. We find evidence of preference conformism in the aggregate and at an individual level (and in the latter there appears to be heterogeneity). While a part of this may be the result of a demand effect, it cannot be wholly attributed to such an effect; and the evidence tends to support Festinger's account of why such conformism might occur. Section 4 discusses these results and concludes that Mill on liberty, who also worried about the prospect of conformism (see quote above), might be a better source, at least in welfare economics, for framing policy interventions once preference satisfaction loses some of its attraction.","Individuals came to the laboratory and were told that they would be making the same set of decisions in two stages. Instructions about the second stage were presented only after the conclusion of the first stage. There were 9 separate Decisions, as set out in Table 1. In each case there are 5 options and the individual had to choose one. The first 3 Decisions are simple choices over objects. Two of these refer to familiar objects (flowers and drinks) and one contains unfamiliar ones, the selection of a country from countries, mainly the eastern provinces of the old Soviet Union, ending in ‘tan’. We refer to these as the ‘label’ decisions. The next two problems are allocation, dictator-like decisions. With two commonly identified social preferences for efficiency and inequality aversion, the options in Decision 5 can be ranked (efficiency is the same but the degree of inequality changes), but the options in Decision 4 cannot be ranked in general if both considerations are in play (because there are changes in efficiency and inequality). The final 4 Decisions refer to risky options: they are the lottery decisions. The options in the last two can be ranked according to their risk. In the first stage, subjects were allocated to either a Red or a Blue group and have to choose one option in these 9 Decisions. In the second stage, subjects made the same 9 Decisions (in the same order) on five occasions (i.e., they made choices in each Decision five times consecutively). They also received information about what others had chosen. Treatments are distinguished according to this information as follows. Treatment 0 (Baseline) = no peer group information Treatment 1 = information on a person's choice from own group in the past. Treatment 2 = information on a person's choice from the other group in the past. Treatment 3 = information on a person's choice from both groups in the past. In Treatments 1, 2 and 3, the information about own or other group choice was phrased on the screen as ‘someone in your group chose….’. We call this the own and other peer signal and we identify a conformist peer effect when the likelihood of an individual choosing an object is increased by an own peer signal of that object. Treatments 1 and 3 supply own peer signals and so provide a test for this effect. We selected two signals (one for each ‘own’ group signal) on the basis of how people chose in each group in Treatment 0. In particular, in the Decisions that cannot be ranked (1, 2, 3, 4, 6 and 7), we selected the two options that were chosen with the smallest frequency. In the Decisions that can be ranked (5, 8 and 9) we selected the two with the biggest distance between them subject to the constraint that both were selected less than 25% of the time. In other words, we aim to have signals in the Treatments 1–3 that are rarely selected in Treatment 0 ensuring that any resulting effect of conformism should be ‘surprising’ given the individual choices. The two signals thus selected were randomly assigned: i.e. one to each group as their ‘own ‘signal (see note below Table 1 for details and the online Appendix for a screenshot of how signals are presented to subjects). Treatment 0 is also important for our test for the influence of these peer signals in Treatments 1-3 because it provides a control for any systematic changes in decision making that arise from the mere fact of repetition and which could give the appearance of a peer effect. Cooper and Rege (2011), for example, find that there is a reversion to the mean in their experiment with the result that choices naturally gravitate upon repetition towards the average of what has been observed in the past. Hypothesis 1a follows. (peer effects): After controlling for any change due to repetition, individuals tend to follow their own peer signal. Any such tendency in the choices that we observe could arise either because all individuals have such a (possibly random) tendency to follow others or because some individuals have such a tendency. Hypothesis 1b follows. (individual consistency in peer effects): After controlling for any change due to repetition, individuals who follow their own peer signal on one or some Decision(s) are more likely to follow their own peer signal on the other Decisions. The scope for subjects to make judgments regarding status and social approval seem more likely in Decisions where the choices can be ranked because distance between options can be meaningfully measured. Thus, in so far as we have not entirely eliminated the possible influence of a preference for status/social approval through the private nature of the decisions and the artificial creation of groups, we would expect if such preferences are in play, the peer effects will be stronger in these problems. H1c follows. (social status/approval peer effects): After controlling for any change due to repetition, individuals are more likely to follow their own peer signal in Decisions where choices can be ranked. If a peer influence is revealed in the test of H1, then it could arise from a ‘demand effect’: the subjects could follow the signal regarding what others do, not because it is what others do, but in response to the fact that this is a signal from those who have designed the experiment. To test for this possibility, we compare the ‘peer’ effect in Treatment 1 with that in Treatment 2. There is a single signal from the experimenter in both cases and if this is all that matters, then we should expect a similar effect on behavior. If however, it is the peer group aspect of the signal that matters, then we expect the influence to be stronger when it comes from own group (Treatment1) than from the other group (Treatment 2). (demand effects only): After controlling for any change due to repetition, the peer effect is the same in Treatment 1 as in Treatment 2. The next two hypotheses concern two views on the origins of conformist preferences and we test them using Treatment 2 and 3 where there are other group signals. (Social Identity Theory): After controlling for any change due to repetition, individuals tend to avoid the other group signal (i.e. choose it less frequently). (Cognitive appraisal): After controlling for any change due to repetition, individuals tend to follow other group signal but less frequently than their own group signal. Experiment was conducted at the Economic Science Institute laboratory of the Chapman University. 60 subjects drawn from the general student population were randomly allocated to each one of our four treatments (Total 240 subjects). Before each stage, subjects read the corresponding instructions on their screen (see Appendix A). In each session, subjects are randomly assigned to one of four orders of the Decision sequence. These were predefined and drawn randomly at the time of the experiment (see online Appendix). Subjects were assigned randomly to either the Red/Blue group. Within each Decision, options were presented in a row and positions were randomized across periods and subjects. Decisions 1–3 (labels) were incentivized weakly (in a gift exchange manner) with a fixed payment of 10 ECUs. In Decisions 4 and 5 (dictator), subjects retained the number of ECUs that they did not allocate to the other person and the ECUs they were allocated as recipient. In Decisions 6–9 (lottery), subjects received the amount of ECUs resulting from a computerized random draw in their chosen lottery. Every choice was paid in ECUs at the end of the experiment. The exchange rate of dollars to ECUs was $1 = 75 ECUs. Subjects made an average of $15 (including $7 as show up fee) for an average 40 minute experiment. This means that the pay-off from any individual choice, while non-zero, is modest. These modest pay-offs together with repetitions could encourage portfolio effects with individuals spreading choices. This would produce changes between Stage 1 and 2, but there would be no pattern of conformism in them.","We first check for possible order, position and color effects. Subjects faced Decisions in one of four different orders. A chi-square test rejects at the 5% level the null hypothesis that the frequency of a particular option is the same across orders in only 7 out of 45 possible cases (so, the tests cannot reject the null that frequencies are the same across the different orders in 38 out of 45 cases). Thus, there seem to be no systematic order effects. It is possible that the position of the options has an effect (e.g., subjects tend to choose the option on the far left). A chi-square test pooling data across subjects and Decisions does not reject at the 5% level the null hypothesis that the frequency of individual choices is the same across positions (p-value: 0.482).7 Thus, there seem to be no systematic position effects. Subjects drawn from the same population are randomly allocated to either a Red or a Blue group. For each Decision in the first stage, a chi-square test rejects at the 5% the null hypothesis that individual choices are the same across groups in 1 out of 9 cases (Decision 6).8 Thus, there seem to be no systematic effects associated with Red versus Blue. Finally, there should be no difference across treatments in terms of choices in the first stage. For each Decision, we compare the distributions of the four treatments and a chi-square test cannot reject the null hypothesis that choices are not different across treatments in 9 out of 9 Decisions. We turn now to the test of our 4 hypotheses. Table 2 gives the mean number of times (out of 5) with which subjects in different Treatments follow their ‘own group signal’ and ‘the other group signal’. In Treatment 0 subjects are not provided with this information, we provide it here for the purposes of comparison with the other Treatments. So, for example in Decision 3, subjects in the second stage of Treatment 0 actually choose the item that we used as an own group signal in Treatment 1 less than would be expected by chance (and average 0.82 times out or 5); whereas in Treatment 1 where they receive this signal they choose the item more frequently than you would expect by chance (an average of 1.72 times out of 5). A Mann Whitney test compares the mean frequency of ‘following own’ and ‘following other’ in stage 2 of Treatments 1–3 with the control frequencies of the choice of these options in Treatment 0 (that is when choices occur for entirely personal idiosyncratic or random reasons because there is no information about any other subject's behavior). Two things are apparent in this overall summary. First, there is always a tendency to follow ‘own signal’ in Treatment 1 (as compared with Treatment 0 choices) and this is significant at 5% or better except for Decision 2 where it is not significant and Decisions 1 and 6 where it is significant at only 10% level. This counts in favor of H1a. The same tendency can be found in Treatment 3, but it is statistically weaker (we pick up on the fact that it is weaker when discussing Result 5). Against, H1c and the influence of status and social approval, the Decisions that can be ranked (5, 8 and 9) do not appear to have stronger peer effects than the other Decisions in either a quantitative or statistical significance sense. These statistical tests on the aggregate data are potentially subject to a multiple hypothesis testing critique. As a result, we reproduce the analysis in Table 3 at the level of type of Decision (label, dictator and lottery). The same pattern emerges. The frequency of following ‘own signal’ is statistically significantly higher in Treatment 1 than the choice of the same options in Treatment 0 in each Decision type; and these differences are still statistically significant after the conservative Bonferroni adjustment of the test statistic for multiple hypothesis testing.9 We now turn to the evidence on these points from individual random effect Poisson regressions in Table 4. In these regressions, we consider the possible influences on the count of an individual ‘following own’ signal and ‘following other’ signal for a given Decision. In each case there are several specifications. They share a control for the initial choice in stage 1 (whether it coincided with ‘own signal’ or ‘other’ signal), dummies for each Treatment and dummies for Decision types. First, we note, in favor of H1a that the coefficients on Treatment 1 dummy and Treatment 3 dummy (where there are own group signals) are positive and significant in the ‘follow own’ signal equations. The sizes of the coefficients are also interesting. It is bigger in Treatment 1 than Treatment 3, where they are respectively 70% and 50% of the size of the coefficient on initial choice. The effect of conformity in this sense may be smaller in both cases than the influence of consistency coming from ‘initial choice’, but, with 50–70% of the influence of consistency, these conformity effects are quantitatively non-trivial. (Again we pick up on the fact that the own group effect appears weaker in Treatment 3 than Treatment 1 below.) Second, there is weak evidence in favor of H1c as the dummy for the ranked Decisions is positive and significant at a 10% level.10 (consistent with H1a): Individuals are more likely to choose an object when they know a member of their own group has done so in the past. (mixed in relation to H1c): There is only weak evidence in individual regressions and none in the aggregate data that following own signal is more likely in Decisions where the options can be ranked. Turning to H1b, we classify a subject as a ‘follower’ in a Decision in Treatment 1 if they follow their own group signal when they did not initially choose this option, on at least half occasions that they encountered this Decision with the signal. We now use this individual Decision determination to create a ‘follower’ index for each type of Decision (Label, Dictator, Lottery): it is given by the proportion of Decisions in that type of Decision that the person is classified as a ‘follower’ (e.g. for Labels, the index takes on 0, 0.33, 0.67 or 1 as there are 3 Decisions). We now test H1b by considering whether knowing this ‘follower’ index for one type of Decision helps predict a subject's propensity to follow the own group signal in the other types of Decision. We do this by entering the ‘follower’ index as an explanatory variable in the regression in Table 5 on the individual frequency of following own group signal, along the lines of the earlier individual regressions in Table 4, except we now drop the observations from the Decisions that we have used to calculate the Follower index. As it helps in addressing the next hypothesis, we can do the same for the subjects in Treatment 2 to construct a ‘follower’ index of the ‘other group’ signal and we can include them in this regression and allow for a possible difference in frequency and the influence of the ‘follower’ index by a Treatment dummy that takes a value of 1 (0) for Treatment 1 (2). Table 5 gives the results (column 1 tests for the influence of the Label ‘follower’ index on the other decisions, column 2 for the influence of the Dictator ‘follower’ index, etc.). The FollowerLabel, FollowerDictator and FollowerLottery by themselves are negative but not significant. The interaction with the variable treatment presents positive coefficients (and similar magnitude to the coefficient on the control for Initial choice) but it is only significant for the FollowerLottery. When we test for whether the sum of the Follower and the Follower interacted with the Treatment dummy are significantly different from zero, we find that the p-values are 0.1778, 0.0349 and 0.0545 for FollowerLabel, FollowerDictator and FollowerLottery, respectively. This means, with these definitions, that the extent to which a person is a ‘follower’ in Treatment 1 in either the dictator or lottery type of Decision in Treatment 1 helps predict (positively) their frequency of following the own group signal in the other Decisions in Treatment 1. However, it is not helpful in predicting decisions in Lottery or Dictator decisions to know the extent to which a subject is a ‘follower’ in Label decisions. In contrast in Treatment 2, none of the follower indices are themselves significantly different from zero (i.e., they only become significant when interacted with the Treatment dummy). The assumption concerning who is a follower in the construction of our Follower index is, of course, arbitrary and so these results are only illustrative. However we check for their robustness in two ways. First, we select a Decision at random for each subject and apply the same rule to classify that person as a ‘follower’ or not. In an analogous fashion, we then examined whether this classification helped predict the frequency of following the own group signal in the remaining 8 Decisions. The precise results depend on the original random selection of the Decision for the classification, but we report a typical result in the online Appendix (Table B4). The coefficient on ‘follower’ is again significant and roughly 55% of the size of the coefficient on the influence of an initial choice of this object. Second, we re-estimate Eq. (5) using OLS with clustered errors at the individual level. The results are stronger in terms of the predictive power of all three Follower variables in Treatment 1 Decisions and are reported in the online Appendix (Table B5). (consistent with H1b): Knowing that a subject is a ‘follower’ in either dictator or lottery Decisions in Treatment 1 helps to predict their likelihood of following their own group signal in the other Decisions. Before we turn to the remaining hypotheses, we explore tentatively what conformism looks like in our sample in Treatment 1 where the evidence of following the own group signal is strongest. First, how many conformists are there? Suppose we use our ‘follower’ classification in each Decision and adopt some plausible cut-off number of Decisions where being a ‘follower’ makes you a ‘conformist’. Suppose, for instance, we define a ‘conformist’, for this purpose, as someone who is a ‘follower’ on at least 4 of the Decisions.11 There were 23 who followed the own group signal with at least this frequency, but we exclude 7 of these subjects from being ‘followers’ on our definition because they also chose their own group signal initially. Our test does not allow us to distinguish whether their behavior reflects consistency or ‘following’. Thus there are 16 clear, with this definition, ‘conformists’ in our sample of 53 where the test can distinguish (or possibly a maximum number of 23 from the sample of 60, if we put all 7 into the sample as conformists). In other words, anything between 25% and 40% our population could be ‘conformist’. Of course, the definition is arbitrary. But it is not implausible and the numbers showing these ‘signs’ of conformism are non-trivial: they are not a small fringe even if they are in a minority. Second, what are the characteristics of the conformists? One way of answering this is to develop an analogous definition of ‘consistent’ choice to identify people who show similar ‘signs’ but this time of ‘consistent’ choice: in this instance, they follow their initial choice at least twice on at least 4 Decisions. 41 of our subjects are ‘consistent’ in this sense. These definitions of ‘conformist’ and ‘consistent’ are, of course, illustrative but they are interesting because they allow for overlap. So, people could both show signs of ‘conformism’ and ‘consistent’ choice under them. Indeed, 12 of our 16 subjects who show ‘signs’ of ‘conformism’ also show ‘signs’ of ‘consistent’ choice. In other words, most of the ‘conformists’ also shows signs of what from the point of economics looks likely perfectly normal behavior. They are not otherwise erratic choosers.12 This is one way of thinking about whether Arendt's conjecture that many, otherwise ordinary people are prone to conformism. Alternatively, we could use our ‘follower’ indices as the measure of conformism and see whether individual differences in the value of these indices are related to gender or attitudes to risk and inequality. We have measures of the latter attitudes from stage 1 choices in Decisions 4, 5, 8 and 9. None of these individual controls was significant (see Table B7 in the online Appendix). In this sense, there is no obvious individual characteristic associated with conformism in our sample. Our next result concerns the relative influence of the ‘own group’ as compared to the ‘other group’ signal. We first present the supporting evidence for the substance of this result on the relative influence and then comment on the interpretation of this evidence in relation to H2 (i.e. the hypothesis that what we have observed in Result 1 is due to a demand effect). In the aggregate data of Tables 2 and 3, there is evidence that individuals also follow the other group's signal in Treatment 2 when this is the only signal: that is, the frequency of ‘following other’ increases between Treatment 0 and Treatment 2. However, this is a weaker effect than the one reported in support of Result 1 on the own group signal in Treatment 1. The difference in the Treatment 2 as compared with Treatment 0 is only statistically significant in 3 Decisions (and one at 10% levels) in Table 2 and one type of Decision in Table 3. This compares with the much stronger evidence on following ‘own group’ signal in the aggregate data in Treatment 1, noted above. The individual regressions in Table 4 also support this difference. The coefficient on Treatment 2 dummy in the ‘follow other’ signal equation is positive and significant in Table 4, but it is smaller than the analogous coefficient from Treatment 1 on the influence of the ‘own group’ signal and this difference is statistically significant (p-values: 0.0495 in (1) and 0.0192 in (2)). Further in Treatment 3 where the subjects receive both signals, the coefficient on ‘follow own’ is positive and significantly different from 0; whereas the coefficient on ‘follow other’ is not significant. Again, the ‘own group’ effect is more powerful than the ‘other group’ signal. The final piece of evidence on the difference in behavior in relation to own and other group signals comes from Table 5 (and is summarized in Result 3). Following own group signal can be useful in predicting future behavior in Treatment 1 but following the other group signal is not useful in Treatment 2. We state that Result 4 tells against H2. This is for the following reasons. The evidence of following the other group signal in Treatment 2 is consistent with a demand effect because subjects could be responding to a piece of information provided by the experimenter. It is also consistent with the Festinger's Cognitive appraisal hypothesis because, on this account, subjects will treat other group information as potentially indicative of what to do. Accordingly, the evidence from Treatment 2 might be thought at best to set an upper bound for the demand effect. Consequently, if the influence of the own group signal in Treatment 1 exceeds that of the other group signal in Treatment, the tendency to conformism in Treatment 1 cannot be entirely explained by a putative demand effect. We arrive at the same conclusion when we consider the evidence on individual behavior in Treatment 3 given in Table 4. Treatment 3 is the least likely to produce a demand effect because the experimenter here provides two opposing pieces of information. In this sense, there is no clear lead or prompt from the experimenter through the experimental design. If, as we find, subjects follow the own signal but not the other signal in these circumstances, it suggests that they find the own signal more salient for reasons other than mere experimental suggestion because this applies equally to the other group signal. (against H2): Subjects are more likely to follow ‘own’ signals than ‘other’ group signals. Our final result concerns the direction of the influence from the other group signal. To summarise the evidence that has already been presented on this in support Result 4, it is never significant and negative. It is either positive and significant in Treatment 2 or insignificant in Treatment 3. In short, there is no evidence that subjects avoid the other group signal (as in H3) and some evidence that they are influenced in the manner suggested by H4. The key to the interpretation of this Result, however, in relation to H3 and H4 is whether the positive effect of the other group signal in Treatment 2 is wholly a demand effect. If it is not and there is some genuine following of the other group signal, then this counts against H3 but is consistent with H4. On the other hand, if the positive effect in Treatment 2 is entirely a demand effect, then this evidence neither tells in favor nor against H3 and likewise H4 (because it is just a demand effect). To help resolve this question, it is potentially helpful to look across all three treatments for a measure of congruence between the results of each. We have already noted that the coefficient on ‘follow other’ in Treatment 2 in Table 4 is significantly smaller than that own ‘follow own’ in Treatment 1. This difference, if we assume for this purpose that the ‘follow other’ influence arises wholly from a demand effect, captures the genuine influence of conformism. This suggests an influence of conformism in Treatment 1 equal to c.0.25 or c.20% of the effect that comes from the ‘initial choice’ control. However, in Treatment 3 where demand effects are much less likely, the size of the coefficient on ‘follow own’ is around 0.5 or 50% of the influence from the ‘initial choice’ control. This higher figure for the genuine conformist influence of the own group in Treatment 3 suggests that the ‘follow other’ coefficient in Treatment 2 cannot be wholly a consequence of a demand effect (i.e. we need to subtract something less than 0.5, the whole Treatment 2 coefficient, from the Treatment 1 coefficient to get a residual equal to 0.5 for the genuine conformist effect suggested in Treatment 3). To be specific, a demand effect equivalent to c.0.2 on the coefficient in Treatment 1 would yield a similar residual coefficient for the genuine influence of conformism in this Treatment to that found in Treatment 3 (i.e. 0.5). But if c.0.2 on the coefficient captures the demand effect, this leaves 0.3 on the Treatment 2 coefficient as genuine following of this signal.13 For this reason, we conclude that this Result, although not decisive, mildly favors H4 over H3.14 (mildly favoring H4 over H3): Subjects do not avoid other group signal.","In what is a ‘tough’ test in Popper's sense, we find evidence of a peer effect on behavior (Result 1). This evidence points to preference conformism because the other possible sources of behaving alike are unlikely to be important in these decision tasks. For example, it is difficult to be motivated by status and social approval when actions remain private; and there is only weak evidence that these peer effects are stronger in decision problems that can be ranked where the scope for such sentiments is stronger (Result 2). In addition, we find that individuals’ tendencies to conformism exhibit a measure of consistency across the Decisions (Result 3). This is important. The evidence of peer effects could have arisen in our experiment because individuals had some random propensity to follow the signal of what some other person was doing. Indeed, this might be the way that an experimental demand effect would operate. However, the evidence on consistency in conformity tells against this. So does the direct evidence on demand effects: while there is some evidence that is consistent with a demand effect, it is patchy and even if taken at face value, it would not account for all the conformism we observe (Result 4). Although this is important new evidence for the existence of preference conformism, it is perhaps not so surprising. Imitation is a well-known form of learning and in other experiments imitation captures a relatively large fraction of behaviors (e.g. see Apesteguia et al., 2007). Furthermore, there are evolutionary models where preferences are selected because they have survival value: i.e. preferences are, in effect, imitated and become endogenous (e.g. see Huck and Oechssler, 1998). Where the decision tasks that people face require equilibrium selection or shared social preferences to achieve more efficient outcomes, such a process of endogenous preference formation will yield in important respects shared preferences: that is, a form of preference conformism (see Bowles and Gintis, 2011). In turn, habits of preference imitation that have evolved to solve social dilemmas may easily carry over to other types of decision problem. There are also two prominent theories in psychology that predict forms of preference conformism. Our evidence mildly seems to favor Festinger's over Social Identity theory. Preference conformism is potentially troubling for some parts of economics and so these results are important. How troubling, of course, depends on the extent of such conformism and the nature of the trouble caused. On the numbers, our illustrative basis for classifying subjects as ‘conformists’ is, in this respect, arbitrary and so our figure of anything between 25% and 40% could easily change. Nevertheless, this way of coming up with a number and the fact that the coefficient in the individual regressions on the own group signal is as much as 70% of the size of the coefficient on whether someone initially revealed a preference for this object when there was no peer information suggests that the magnitude of conformist influence is non- trivial. What, then, is the nature of the trouble? It is not deeply worrying for positive economics because modeling depends on some plausible foundational assumptions regarding behavior. What is required for this is some behavioral regularities and not that they should be derived from a model of preference satisfaction alone. In this respect the evidence from this experiment joins the growing evidence from behavioral economics that there are many regularities in behavior that are not (or not easily) derivable from the preference satisfaction model; and positive economics should take this into account. The importance of this paper in this regard is that this particular kind of non-rational choice regularity in behavior has not hitherto featured significantly in the list of behavioral economic insights. The trouble is possibly deeper for welfare economics, because the appeal of the Pareto criterion depends on taking individual preferences as given. There are alternative criteria, however. Sugden (2004) is one example of an attempt to define an alternative metric of opportunity and connect this to the operation of markets when preferences are not well defined. Another comes from J.S. Mill. He is, of course, well known as an interpreter of utilitarianism and this may seem to push him in the direction of preference satisfaction, but his most famous work, On Liberty, presents an alternative view. In On Liberty, Mill wants to advance a constitution of liberty because this enables individuals to acquire their own character - that is, their individuality through the exercise of freedom of thought, discussion and ‘experiments in living’. He did not wish here to ground policy on satisfying given preferences because people's preferences were naturally something in the making. His policy approach, like that of Buchanan (1986) later, is constitutional. Policy should be focused on establishing the rules governing outcomes and not the outcomes themselves. This, then, is the other respect in which this experiment is important. It is an encouragement to thinking more about some of the alternatives to the Pareto criterion approach in welfare economics. Nevertheless, even for constitutionalists like Mill, the evidence from this experiment is still worrying. He thought, as the quote at the beginning suggests, that conformism was the enemy of the development of individual character. One encouraging aspect of the evidence from this experiment in this respect, however, is Result 5. It seems that there is stronger evidence for the cognitive or epistemic source of conformism than the Social Identity one. This helps provide a crumb of comfort for Mill in our results. He thought that liberty would work against conformity. In so far as the exercise of liberty provides diversity, then our results support this conclusion. When there is an ‘other group’ as well as the ‘own group’ signal, the ‘own group’ signal effect on behavior becomes weaker."],["We propose that religion impacts trust and trustworthiness in ways that depend on how individuals are socially identified and connected. Religiosity and religious affiliation may serve as markers for statistical discrimination. Further, affiliation to the same religion may enhance group identity, or affiliation irrespective of creed may lend social identity, and in turn induce taste-based discrimination. Religiosity may also relate to general prejudice. We test these hypotheses across three culturally diverse countries. Participants׳ willingness to discriminate, beliefs of how trustworthy or trusting others are, as well as actual trust and trustworthiness are measured incentive compatibly. We find that interpersonal similarity in religiosity and affiliation promote trust through beliefs of reciprocity. Religious participants also believe that those belonging to some faith are trustworthier, but invest more trust only in those of the same religion—religiosity amplifies this effect. Across non-religious categories, whereas more religious participants are more willing to discriminate, less religious participants are as likely to display group biases. --------------------------------------------------------------------------------","In this paper, we investigate the role of religion-based discrimination in trusting and in trustworthy behaviour when interacting with people from various social groups or cultures. Understanding the role of religion is important, because conflict between and within different religions is rising globally (The Institute for Economics and Peace, 2014; Grim, 2014) and fast becoming a defining feature of the post-cold war world order (Huntington, 1996). A standard manifestation of this religious conflict is inter-religious strife. Another, newer dimension involves religious radicalisation and extremism which can turn individuals against their compatriots and moderate fellow adherents. However, despite its ubiquity, importance and controversy, economists have only recently developed an interest in the effects religion has on economic outcomes (Iannaccone, 1998; Guiso et al., 2006; Tan, 2006). Religion can influence economic behaviour in at least two ways, by creating differential social group identities (Jackson and Hunsberger, 1999) and through individual differences in religiosity, i.e. the strength of an individual׳s religious attachment or commitment to a particular faith commonly measured as religious belief, ritual and experience (Tan, 2006). Identity (e.g. Akerlof and Kranton, 2000; Chen and Xin, 2009; Currarini and Mengel, 2013) and acculturation (Guiso et al., 2003) generally affect economic outcomes and might act as conduits for the economic influences of religion. One economic approach to examining these effects is the experimental economics of religion, as critically discussed by Hoffmann (2013) and Tan (2014), where the influences of religious variables on various kinds of individual economic decision are studied systematically in controlled settings. Previous studies demonstrated the first effect, that individuals treat others differently in economic contexts based on same or different religious affiliation even when other social identifiers such as nationality and ethnicity are shared. For example, we conducted a laboratory experiment with student participants from different cross-cutting ethnic and religious groups in Malaysia (Chuah et al., 2014). While participants cooperated relatively more within their own ethnic groups irrespective of religious affiliation, having the same religion as well enhanced their cooperation further. Conversely, participants divided by different ethnic identity cooperated more when they shared religious affiliation. A field experiment where both Indian Hindus and Muslims in Mumbai trusted members of their own religious groups relatively more (Chuah et al., 2013) lends further support. However, our work as well as that of other researchers failed to demonstrate the second effect, of religiosity, directly. In two experiments participants of higher religiosity were equally cooperative (Chuah et al., 2014) or trusting (Tan and Vogel, 2008) than others. These results suggest that religiosity, in reflecting an individual׳s socialisation into and internalisation of particular religious precepts (e.g. Ryan et al., 1993) does not independently affect consequent behaviour. However, both studies provided hints of a second avenue by which religiosity might influence decision making as a vehicle for taste-based or statistical discrimination. One such hint is that among the entirely Christian participant pool of Tan and Vogel (2008), those of known higher religiosity receive greater trust from others, and especially (but not exclusively) from those who share this trait. The second hint is that high religiosity amplified the higher cooperation which Chuah et al.׳s (2014) multi-cultural participants paid their religious fellows. In this paper, we propose that religious identities serve as cues on the nature and degree of connectedness between interacting individuals, and thus religion influences strategic behaviour, in particular trust and trustworthiness on which we focus here. In trust games (Berg et al., 1995; Johnson and Mislin, 2011), a sender decides how much to trust a receiver by sending an amount of money. The receiver receives thrice the amount sent and decides how trustworthy to be in returning a proportion of it. In equilibrium, by backward induction, assuming that receivers are rational and money- maximising, senders anticipate nothing in return, and so send nothing. Social connectedness is a psychological concept describing the closeness of people e.g. family or acquaintance, friend or foe (Aron et al., 1991; Gächter et al., 2015). We call closeness in religion-based relationships religious connectedness. Consistent with research on social connectedness in general (Laurenceau et al., 1998), we argue that individual religiosity operates through religious connectedness to affect trust. Religious connectedness increases with the duration and frequency of interactions, knowledge of others, the extent of (mutual) self-disclosure, and the number of people in the other׳s network one is also connected to. Religious beliefs, rituals, experiences and activities that unite or divide people facilitates this. We consider four forms of religious identity: (1) a connection at the fundamental level of individual religiosity; (2) group membership based on religious affiliation to the same creed; (3) religious affinity arising from the mere affiliation to some religion, regardless of creed; and (4) religious anonymity, where religiosity effects operate on the wider societal level of prejudice across social identities including non-religious ones. In turn, we examine four corresponding religious discrimination effects on trust and trustworthiness. The first is statistical discrimination (e.g. Mueser, 1999; Anderson et al., 2006), where more religious people are generally believed to be trustworthier and treated accordingly. The second is that religiosity amplifies intergroup bias on the basis of religious affiliation. Intergroup processes including taste-based outgroup discrimination or ingroup favouritism are strengthened by an individual׳s identification with the group (Farnham et al., 1999; Smurda et al., 2006). The third is that religiosity is used as a social identifier of affinity which unites religious people regardless of creed. The fourth is that religiosity is a correlate of greater general prejudice, i.e. discrimination based on social identity differences even in non-religion categories (e.g. Hunsberger and Jackson, 2005). For this purpose, we conduct a trust game experiment where participants can incur a financial cost in order to discriminate between co-participants of different religions and other social identities. We extend the trust game by allowing participants to make decisions conditional on the social identities of co-participants they might face. We then measure participants׳ religiosity and consider their religious affiliations, their responses to co-participants of diverse religious affiliations, and corresponding beliefs regarding co-participants׳ actions. In particular, we study how trustworthy senders think receivers are or how trusting receivers think senders are. We also test how much senders invest trust or receivers reciprocate trust. Further, we analyse whether these beliefs and actions relate to the religiosity and religious affiliation of sender and receiver. This informs us on the relevance of statistical and taste-based motives of discrimination, and whether religiosity per se is related to general prejudice, i.e. on the basis of even non- religious categorisation. Our design has a number of novel features. In many previous experiments, discrimination was observed in a particular context such as gender or ethnicity. In contrast, we are able to measure discrimination based on different social identifiers which vary within a multi-national participant pool. This allows us to measure discrimination tendencies in a more general way, and to compare these across different social identifiers. Further, we measure discrimination in participants׳ intention or willingness to discriminate as the resources they are willing to use in order to be able to make decisions contingent on the characteristics of their co-participants. This provides a graduated measure of discrimination intentions, elicited in an incentive compatible way in line with the costliness of discrimination in many real world settings and economic models (see Mueser, 1999). We discuss the literature and motivation in greater detail in Section 2. We outline our experiment and hypotheses in Section 3. Results are reported in Section 4, before concluding in Section 5.","Apart from its role in inter-religious conflicts across the world, high religiosity within all creeds plays an important part in a number of pressing contemporary social debates surrounding home-grown terrorism, abortion, contraception and gay rights. These have clear economic consequences. For example, Indiana׳s Religious Freedom Restoration Act allows trade to be refused on religious grounds, while provisions for religious exemptions from public immunisation programmes (in force in 48 U.S. states) can generate negative externalities on an epidemic scale. This provides economists with a clear motivation to examine the effects of religiosity in economic settings using economic methods. A few experimental economics studies have examined the effects of religiosity (a.k.a. religiousness, which measures an individual׳s attachment or commitment to a particular faith) on economic behaviour. Most use religious service attendance measures as a proxy and relate this to prosocial behaviour in experimental games.1 Generally, previous research has found little evidence for the relationship between religiosity variables and behaviour in the trust game. Fehr et al. (2002) found no effect of the church attendance of German household survey respondents on their decisions in a trust game. Karlan (2005) measured religiosity in terms of months since last religious service attendance and related this variable to public good contributions and trust game decisions in a field experiment in rural Peru. It was inversely related to public good contribution but only at the 10% level of significance. Attendance also did not explain trust game decisions in this study directly. However, participants with less frequent attendance were sent greater amounts for unexplained reasons. Anderson and Mellor (2009) measured the frequency of religious service attendance to serve as a proxy for religiosity. This variable was not significantly related to public good game contributions of older adult U.S. participants. Anderson et al. (2010) subsequently found a positive effect with college student participants, but only when comparing the corner cases of high and low attendance. Trust game behaviour here was unrelated to the attendance measure. Tan (2014) argued that one reason for the mixed results in terms of effect significance and direction could lie in the multi-dimensional nature of religiosity that simpler measures do not capture, e.g. based on attendance alone. Unidimensional religiosity measures like these are unsatisfactory as they fail to tap into the different motivations behind and expressions of religious attachment (Spilka et al., 2003, p. 28; Hill and Hood, 1999, p. 5), which can manifest behaviourally in opposite directions (e.g. Tan, 2006). For example, intrinsic spiritual or quest motives for religious attachment are sharply differentiated from extrinsic ones such as seeking social group identification. In response psychologists of religion have developed a now widely accepted approach (DeJong et al., 1976) which measures individual religiosity in terms of five dimensions, religious knowledge, practice of religious activities, belief in religious precepts, personal mystical experience and consequences of religion on behaviour (Glock and Stark, 1965). We used such multi- dimensional religiosity measures in a number of previous experimental economics studies with promising but still inconclusive results. Tan (2006) found the different components of a multi-dimensional measure to significantly affect dictator game offers or ultimatum game responses but in opposite directions. Chuah et al. (2009) used principal components analysis to derive a multi-dimensional religiosity scale using 15 items from the World Values Survey (see Inglehart, 1997) which was negatively and (marginally) significantly associated with ultimatum game offer sizes among Malaysian and UK participants. In the study by Tan and Vogel (2008) on German University students, higher religiosity receivers were trusted more especially by fellow high-religiosity senders. Receivers of higher religiosity returned greater amounts and especially to more religious senders. The results of Tan and Vogel (2008) suggest that religiosity can have an indirect effect as a social identity that generates ingroup favouritism. However, this is inconclusive in that religiosity differences in this study did not explain why senders trusted more religious receivers more. Alternatively the result could evidence statistical discrimination towards highly religious people to the extent that they are generally held to be trustworthier. Finally, in Chuah et al.׳s (2014) prisoner׳s dilemma experiment, shared religious creed raised cooperation within a multi-cultural Malaysian student participant pool. In contrast, multi-dimensional religiosity as an independent variable in its own right did not explain cooperation. However, religiosity raised cooperation further when interacted with the shared creed dummy variable. This result suggests a further, again indirect effect of religiosity as an enhancer of ingroup bias based on shared religious affiliation. Alternatively, the result could reflect the greater general tendency of religious individuals to discriminate on the basis of different social identities including religious creed. Let us now consolidate these results as behavioural patterns from the perspective of religious connectedness, as outlined in the Introduction. First, individual religiosity can increase connectedness in three ways. First, the participation in ritual increases the duration and frequency of interactions between individuals. Second, increases in religious knowledge and indoctrination increase knowledge of others in the group, e.g. how they think they ought to behave (Tan, 2006). The latter relates to the access to relevant social category, and in turn the likelihood of using that social categories as stereotypes to guide behaviour such as trust (Tan and Vogel, 2008). Thirdly and indirectly, common beliefs and experiences engender familiarity and closeness, which then carry over to group identification and biases at the levels of similarity in religiosity (Tan and Vogel, 2008) or religious affiliation (Chuah et al., 2014). Such effects should weaken as religious connectedness weakens, via the above processes as well as a decreasing overlap in social networks. In the limit, we have interactions across group markers that are orthogonal to religion. If so, would individual religiosity lose its bite on discrimination? Measuring trust and religion ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Following previous studies we used a trust game as a behavioural measure allowing for the expression of discrimination (e.g. Fershtman and Gneezy, 2001; Holm and Danielson, 2005; Falk and Zehnder, 2013). As shown in Fig. 1, we used a binary version of the trust game because it is cognitively less demanding on participants, so as to reduce biases from fatigue in view of the 88 games each participant had to play. The sender and the receiver begin each game with 200 points. We test two parameterisations of the trust game. In the first, namely the “low stake game”, the sender decides whether or not to trust, i.e. to send 50 or 0 to the receiver. If the sender sends the money, the receiver receives three times this amount and decides whether or not to be trustworthy by returning 100 or 0. In the second, i.e. the “high stake game”, we increase the stakes by allowing the sender to send 150 or 0 to the receiver, and the receiver decides whether or not to return 300 or 0. Assuming that players are rational and money maximising, in equilibrium nobody sends any money. By backward induction, receivers will prefer more money to less and not return anything to the sender, i.e. not reciprocate. The sender anticipates this and prefers not to send anything to the receiver, i.e. not trust, because the payoff from withholding is higher than sending and not receiving anything in return. The subgame perfect equilibrium is that neither sender nor receiver sends any money. This forms the benchmark relative to which we can measure the trust and trustworthiness of senders and receivers, respectively. It follows that there is low (high) temptation for the receiver to send 0, and this implies a low (high) stake for the sender in trusting the receiver. The two games allow us to test our hypotheses within a wider domain of stakes. In order to obtain measures of discrimination, we administered the trust game under different social identity conditions using the strategy method (Selten, 1967). To keep sender and receiver tasks symmetric, in the experiment we allowed receivers to choose “return” or “not return” under the understanding that the decision only applies if the sender had chosen “send”. In practice, the sender׳s decision would not influence payoffs in the game if the sender does not send any money. To make this explicit, we displayed games on the screen as extensive form representations consistent with this strategy method setup (see Fig. 2). In the first two rounds of the experiment, all senders and receivers stated their decision of whether to send or not to send without knowing the social identities of their co-participants. One round was for the high stake condition and the other the low stake condition, in counterbalanced orders across participants. We call these actions default actions. In the other rounds that followed, participants stated their decision based on every possible co- participant׳s social identity type according to different social categories (see Table 1). There were 88 rounds in total. Using religious affiliation as an example of a category, every participant was asked whether they would send or not send to co-participants of every religious affiliation (type) we provided, i.e. Buddhist, Christian, Hindu, Jewish, Muslim, other or none. This process was repeated for every type of every category, presented in random order after the tasks without social identity were performed. We call these actions conditional actions. Each category thus constitutes an experimental condition. In each round where participants could base their decisions on the co- participants׳ social identities, they were provided with an additional endowment of 100 points from which they could spend an amount of their choice to increase the probability of implementing their conditional action instead of their default action. Each point increases the probability by 1%, and each point unspent accrues as experimental payoffs. This incentive compatibly elicits their willingness to discriminate (WTD). When calculating experimental earnings, we applied the participant׳s stated WTD for the condition in concern to set the probability that the conditional action rather than the default action was to be used, and randomly determined subject to this probability. As an example, consider a high stake game where a participant chooses to send 150 to co- participants of high religiosity, and 0 to other types of co-participants. Assume also that the default is to send 0. A WTD of 20 points means that if the participant is subsequently randomly matched with a high religiosity co-participant for the purpose of calculating experimental earnings, there is a 20% probability that the choice of sending 150 is implemented, and a complementary probability of 80% that the default action of sending 0 will be implemented. A WTD of 100 points means sending 150 to the high religiosity co-participant for sure, and sending 0 to a medium or low religiosity co- participant for sure. Higher WTD values increase the probability that discriminating decisions are used to determine earnings and therefore represent the decision maker׳s willingness to pay for social identity information to afford discriminating actions. This method of eliciting WTD is novel and has two advantages. First, it experimentally models the costliness involved in discrimination activities, e.g. it takes time and effort to find out another person׳s religiousness or political inclination. This introduces an externally valid dimension to the test. In retrospect, observed decisions in previous experiments without this feature, e.g. Tan and Vogel (2008), capture behaviour “as if” the participant confidently assumes or knows the co-participant׳s social type. Second, the costliness of discrimination is in a way a disincentive to discriminate that mitigates demand effects in terms of discriminating actions, and in doing so incentive compatibly reveals the demand of the individual who despite of this cost goes for it. That said, we should not and do not try to remove all demand effects from the experiment, for we are interested in those germane to the act of discriminating on the basis of social identity—to which we can clearly attribute as the cause of action. Fig. 2 shows the experimental interface employed to elicit decisions. The interface shown in this example is asking participant 39, assigned to the sender role (Person A) to make decisions in the religiosity category for a low stake game (Round 4). The game tree displays the actions and associated payoffs for participants in both roles. The dark shaded button indicates the benchmark decision this participant has already indicated previously, which cannot be changed (SEND). The participant must make trust decisions in the religiosity category by clicking on either the SEND or NOT SEND buttons for each possible co-participant religiosity type, namely “High”, “Medium” and “Low” religiosity. The participant then indicates what proportion of the 100 points to allocate towards implementing the relevant conditional choice, i.e. their stated WTD. Once all these decisions have been made, the participant clicks on CONFIRM to enter them and proceed to the next round, which involves a different category. We administered a pen-and-paper questionnaire after the completion of the trust game task to collect additional measures. Beliefs were elicited as participants׳ expectations of co-participant actions in the trust game. Participants were asked (in their roles and for every possible value in every social identity category) to state the probability that such a type of co-participant would choose to send. Participants were paid depending on how close these beliefs were to true distribution of choices observed in the experiment, and payments were computed according to the quadratic scoring rule (Selten, 1998). We also recorded each participant׳s own demographic characteristics for each of the social identity categories in order to classify them in terms of the values for each category shown in Table 1. Notably, we elicited individual religiosity according to the Glock and Stark (1965) dimensions using the denomination- robust 8-item instrument by Rohrbaugh and Jessor (1975) which yields our religiosity measure. It takes into consideration different dimensions of religion, namely belief, ritual, consequences, theology, and experience. It delivers an individual׳s overall score between 0 and 32 (see Hill and Hood , 1999). Hypotheses ~~~~~~~~~~ Piecing together the mosaic of results given by the literature from the perspective of religious identity and connectedness, we shall use our experiment to test the following four hypotheses. These explanations of behaviour are not mutually exclusive and could operate in concert, potentially coexisting or reinforcing each other. We cater for these possibilities in the analysis.Statistical discrimination Statistical discrimination Senders generally believe that receivers of higher religiosity are trustworthier, and statistically discriminate by being more likely to trust them more than receivers of no or lower religiosity. The first possibility for the expression of religiosity in terms of economic behaviour is statistical discrimination (e.g. Anderson et al., 2006) when a person׳s social identity contains information regarding particular behaviour tendencies that can feed into strategic considerations, e.g. beliefs of trustworthiness. Statistical discrimination in the trust game applies only to senders, as they must anticipate the likelihood that receivers will fulfil or abuse their trust if invested. Tenets such as charity, neighbourly love and the Golden Rule are common to all religions and may confer a trustworthy reputation on religious people (e.g. Spilka et al., 2003, p. 172). If statistical discrimination based on religiosity is present in the current experiment, all senders regardless of their own religiosity should be more likely to behave trustingly towards receivers of greater religiosity levels. Senders would therefore be more likely to send to receivers of higher religiosity, compared to receivers of lower religiosity, and this effect should increase with the sender׳s religiosity.Ingroup love Ingroup love Religiously affiliated senders are more likely to invest trust in receivers who are affiliated to the same religion, relative to receivers who are not religiously affiliated or affiliated to a different religion. This effect increases with the sender׳s religiosity. Religiosity is a fundamental measure of religiousness as an individual. It might vary across religious affiliations. In turn, it weakens connectedness, e.g. from variances in religious doctrine and prescriptions for behaviour. Further, it is arguably more subtle than religious affiliation, which may serve mainly as a badge of membership. It follows that while religiosity might be a weaker marker of religious connectedness, it could serve to amplify discrimination effects based on religious affiliation, which increases the salience of religious categories as social markers. Thus, the degree to which people exhibit biased intergroup behaviour is related to the strength of their identification with the group concerned, and in turn increases cooperation through stronger social preferences (Farnham et al., 1999; Chen and Xin, 2009). In particular, greater discrimination can result from a loss in (implicit) self-esteem in people who highly identify with a particular social group that is undergoing a threat, i.e. a perceived negative evaluation by others (Smurda et al., 2006). In the current context this hypothesis suggests that greater trust in co-participants of the same religious group is relatively stronger in more religious participants in either role. Such effects are reinforced by individual religiosity, which embodies closeness nurtured through joint participation in activities. This, in turn, increases trust by increasing religious connectedness through commitment to the creed, i.e. ingroup membership. Religiously affiliated senders would therefore be more likely to send to receivers belonging to the same creed, compared to receivers who are atheists of followers of other creeds, and this effect should increase with the sender׳s religiosity.Religious affinity Religious affinity Religiously affiliated senders are more likely to invest trust in receivers who are affiliated to some—regardless of which—religion. This effect increases with the sender׳s religiosity. The third possibility we test is that people consider their religious affiliation or religiosity a pertinent social identity and exhibit biased intergroup behaviour (i.e. ingroup favouritism or outgroup prejudice) towards others depending on whether or not they are also religiously affiliated to some creed—irrespective of whether or not it is the same one. For example, former Prime Minister of the United Kingdom Tony Blair articulated this thinking publicly at the Westminster Faith Debate on “Religion in Public Life” held in London on 24 July 2012,2 “I find a connection with people who are of faith, even though they׳re of a different faith to my own, precisely because there is a certain space, philosophically and emotionally, you can congregate around”. Put differently, this weakens the religious connectedness relative to that between individuals of the same creed. That said, religious affinity does not extend to group membership, and its effect should be relatively weaker. A religious affiliate would thus be more likely to send in the trust game to another who is affiliated to some religion—regardless of whether or not it is the same creed, and this effect should increase with the sender׳s religiosity.General prejudice General prejudice Religious senders are generally more biased, such that they are more likely to send to receivers with the same non-religious social identity. Finally, since the middle of the last century (Adorno et al., 1950; Allport, 1954), psychological studies have repeatedly identified links between individual religiousness and attitudes of prejudice. Such prejudice is counter to religious teachings of charity, forgiveness, love and compassion. This link is complex and dependent on a number of other factors including religious orientation, social desirability and doctrinal attitudes towards particular out-groups (Spilka et al., 2003, Chapter 14). Links between religiosity and prejudicial attitudes have been demonstrated repeatedly (Allport and Ross, 1967; Altemeyer and Hunsberger, 1992; Hunsberger and Jackson, 2005; Hunsberger, 1996; Jackson and Hunsberger, 1999). We consider the possibility that religious people are generally more disriminating in the context with the weakest religious connectedness. If this holds, we should find that senders of higher religiosity have greater WTD across all social identity categories or overall. We should also find that religious senders are more likely to send to the “ingroup” based even on non-religious categories. In experimental terms, we are thus testing for the effect of religion on the individual׳s inherent disposition to discriminate. Procedure ~~~~~~~~~ We ran the experiment at the China, Malaysia, and UK campuses of the University of Nottingham. All campuses use English as the medium of instruction, and share common degree structures and syllabi. This participant pool affords high direct comparability of data collected from these diverse cultures. The cultural diversity of our sample widens the study׳s domain of validity. Such diversity increases the number of subjects of each social identity type. Thus, there is a much larger number of ingroup and outgroup combinations, which we shall also use to test for the cultural sensitivity or robustness of our findings. We used a computerised interface in English with 545 student volunteers (273 senders and 272 receivers) recruited by poster and e-mail announcements for 90-min sessions of 20–40 participants. The experimental software was programmed in Visual Basic 6, and the computerised text was in English. Our experiment followed the standards of cross-cultural experimental economics (Roth et al., 1991; Herrmann et al., 2008). Instructions, comprehension quiz questions, belief elicitation and demographic questionnaire were provided in the respective local languages. The English version was always available to participants in China and Malaysia on demand. The original English version was first translated to Chinese and Malay, and then back translated to English to check for consistency. Any inconsistencies were resolved in consensus with the co-authors on this project. Translations were performed by three people who are not co-authors on the project, but are native speakers of Chinese or Bahasa Melayu and English. All of them have professionally worked in the respective two languages. The English version of the experimental instructions is found in the online appendix. Participants were randomly assigned to either the sender or receiver role throughout the experiment, and made trust game decisions first for socially unidentified co-participants and then for each of the social identity categories and types as described (see Table 1), for both the low and high stake conditions, in individualised random order. After all experimental sessions were completed, participants were randomly matched experiment-wide across the three locations, and one social identity category was selected randomly to determine earnings. The participants׳ total earnings were the points from the game, those remaining from the WTD endowment, and payments depending on the accuracy of their beliefs in one randomly chosen belief task, with the answer compared to the statistical return rate of the sample for the type of participant. We paid participants at the rates of Renminbi (RMB) 0.20, Ringgit Malaysia (RM) 0.08 and Pounds Sterling (£) 0.04 per point earned plus a show-up fee (RMB 25, RM 10 or £5, respectively). Earnings were collected a week after the final session to allow for experiment-wide participant matching over the three locations. We paid participants in the three locations RMB 63.68, RM 28.66 and £14.65 on average. Each session lasted approximately 1.5 h. The exchange rate between the three currencies we used was determined using the Big Mac Index published annually by The Economist magazine.","Before testing our four hypotheses we look at some basic features of the data. Table A1 provides the distributions of participant types of each category across the three locations, and a summary of mean WTD, beliefs and actions across conditions and types by roles and locations. Religiosity scores ranged from 0 to 30 and the average was 11.9. The mean age was 20.5 (standard deviation of 1.74). There are more females (122) than males (65) in the Malaysian sample, more males (137) than females (27) in the Chinese sample, and almost as many females (96) and males (98) in the UK sample. In ethnic and religious terms, China is most homogeneous with 162 ethnic Chinese, 134 atheists and 25 Buddhists, out of 164 participants in total. Malaysia and UK are relatively heterogeneous, with Chinese (106) and White (115) as majority ethnicities, and Buddhists (61) and Christians (56) as majority religions, out of 187 and 194 participants in total, respectively. In Malaysia and UK, the non-majority religions are all represented, apart from no Jewish participant in the Malaysia subsample. In the high (low) stake baseline games where decisions could not be conditioned on the social identity of co-participants, 38.1% (56.0%) of senders chose to trust, and 27.9% (43.0%) of receivers chose to reciprocate. Further details may be found in Table A1. Preliminaries ~~~~~~~~~~~~~ We first check for independent effects of religiosity on trust, to confirm the result from previous studies that forms our departure point. Our measure of religiosity is RELI, which is the mean centred to avoid multi-collinearity in our regressions below, following Marquardt (1980). There is no significant difference in the religiosity of senders who trust and those who do not in both the high (t-test, p=0.780, 2-tailed henceforth) and in the low stake condition (p=0.758), or for receivers in either the high (p=0.775) or the low condition (p=0.886). To corroborate, individual level random effects binary logit regressions controlling for beliefs, stake and gender show that religiosity does not significantly influence trust and trustworthiness (p=0.921 and p=0.375, respectively; see Table A1 for details). As there is no evidence for an independent influence of religiosity on trust and trustworthiness. Senders spent an average of 21.4 and receivers 22.0 out of a hundred points to increase the probability of implementing their conditional actions (i.e. WTD) in the religious affiliation condition, where actions could be conditioned on the co- participant׳s religious denomination. WTD rises with one׳s religiosity level at 19.3, 21.2, and 30.9 for low, medium and high religiosity, respectively. Senders spent an average of 20.9 and receivers 20.0 points on WTD in the religiosity condition, where actions could be conditioned on the co-participants level of religiousness. WTD rises with one׳s religiosity level at 19.8, 19.9 and 26.0 for low, medium and high religiosity, respectively. The same pattern holds for receivers at 19.5, 23.0, and 29.6 (18.7, 21.8, and 21.5), respectively, for low, medium, and high religiosity. Further, 58.6% (50.7%) of senders and 48.2% (43.9%) of receivers discriminate on the basis of religious affiliation (religiosity) in the sense that they choose different conditional actions for different types of co-participants.3 With information of religious affiliation (religiosity), 23% (22.7%) of behaviour differs from that in the baseline: 9.2% (10.1%) increase and 13.8% (12.7%) decrease trust. As described in Section 3.2, this widely observed discrimination can take a number of forms as expressed in our four hypotheses, which we test next. To control for and to test the interplay of effects from multiple variables and their interactions, we use multivariate analysis. In the following subsections, our regressions include individual-level random effects to control for the potential non-independence of multiple observations per individual. We never provided participants with feedback between choices so there is independence between observations across participants. We always control for low and high stake conditions (STAKE=1 for the high stake condition and =0 for the low stake condition), and for own gender (FEMALE=1 for females and =0 for males) due to known gender effects on trust game behaviour (Croson and Buchan, 1999).4 Our regressions always include individual religiosity RELI. Results are robust to the inclusion of WTD or dummy variables for location (these alternative models are reported in online appendix OA3). Statistical discrimination ~~~~~~~~~~~~~~~~~~~~~~~~~~ Statistical discrimination implies that senders believe that some types of receivers are trustworthier than others. These stated beliefs are given by the dependent variable BELIEF=0 to 1. According to Hypothesis 1, a sender, irrespective of her own social identity, uses the receiver׳s religiosity to form an expectation of their trustworthiness. Participants should therefore be willing to pay more than in identity conditions unrelated to any possible statistical discrimination. Our control condition is a “birthday” category where actions were conditioned on whether the co-participant was born on an even or odd day of the month. There, mean WTD is 17.7 and its confidence interval is 16.1–19.2. The mean WTDs of the religious affiliation and religiosity categories are 21.4 and 20.9, respectively, i.e. outside the interval. We also examine how beliefs regarding the trustworthiness of receivers vary with the decision maker׳s religiosity using a religiosity level variable RLEV. This variable was used in the experiment to elicit participants׳ beliefs and actions contingent on the co-participant׳s low (RLEV=0 if religiosity questionnaire score is 0–10), medium (RLEV=1 if score is 11–20) and high (RLEV=2 if score is above 20) religiosity.5 We test this effect on senders across all religious affiliations. Further, to test if being of a similar religiosity level reinforces statistical discrimination, we interact RLEV with RELI. Senders׳ beliefs that low, medium, and high religiosity receivers would act trustworthily are 0.33, 0.41 and 0.43, respectively, pooled over both stake conditions. Average beliefs and actions are shown broken down by participants׳ own religiosity levels in Fig. 3. Senders of diverse religiosities believe that receivers of higher religiosity are more likely to return (top left figure), and are more likely to send to them (top right figure). Receivers of diverse religiosities believe that senders of higher religiosity are more likely to trust (bottom left figure), and are as likely to return to senders of different religiosity levels. Regression analysis (Table 2) confirms that more religious people are trusted more by people across different levels of religiosity, as the RLEV coefficient is positive and significant in models 1–3. This result holds overall, for people without or with religious affiliation, as demonstrated by the regressions on the pooled sample (model 1) and subsamples disaggregated by people without (model 2) or with (model 3) religious affiliation. Further, the statistically insignificant RLEV×RELI coefficient in model 4 shows that senders of different levels of religiosity are as likely to believe that receivers of high religiosity are trustworthier, confirming that statistical discrimination holds across senders irrespective of religiosity. Statistical discrimination Senders of all levels of religiosity believe that receivers of higher religiosity are trustworthier, and behave consistently with this belief by trusting them more. Ingroup love ~~~~~~~~~~~~ According to Hypothesis 2, higher religiosity strengthens the identification of participants with the religious group they are affiliated to, and thereby amplifies ingroup biases based on affiliation. We use WSEND as the dependent variable. To test for ingroup biases, we define a dummy variable INGROUP that takes on the value of 1 when participants are making decisions conditional on participants that are of the same type as those for the category in concern. In this case of ingroup biases in religious affiliation, INGROUP=1 when co-participants are of the same religious affiliation, and =0 otherwise. When people have information about others, they use it to guide their actions. In turn, this feeds into behaviour. Thus, our models of WSEND include BELIEF to control for statistical discrimination. However, beliefs do not necessarily explain behaviour completely, for taste-based discrimination can also play a role.8 Thus, by controlling for the effect of statistical discrimination with BELIEF, INGROUP is a measure for taste-based discrimination, such that remaining ingroup effects are attributable to it. We include the mean centred measure of individual religiosity RELI as well as the interaction term INGROUP×RELI, which tests if ingroup biases are strengthened by the decision maker׳s religiosity. Tests are performed on data from the religious affiliation condition rather than the religiosity condition where there is no clear sense of group membership. Note that participants were not told their own religiosity level according to our survey measure nor asked to state their perception of their own religiosity in absolute terms or relative to other participants. Fig. 4 shows the percentage change in trust actions in WSEND conditional on the receiver׳s religious affiliation, relative to the baseline where decisions are made unconditionally. In UK and Malaysia, where most participants have religious affiliations, we observe increases in trust for the ingroup relative to the baseline, i.e. ingroup favouritism. In China, where most participants are atheists, we observe decreases in trust for the outgroup relative to the baseline, i.e. outgroup prejudice. We scrutinise this econometrically. Referring to Table 3, model 7 shows that senders are more trusting towards those of the same religious affiliation (INGROUP is positive and significant) and this effect increases with one׳s religiosity (INGROUP×RELI is positive and marginally significant). The figure in Table A3 shows that ingroups are consistently trusted more than outgroups by people across different religions. This finding is also robust to contextual differences across groups and societies.9 This ingroup effect does not hold for atheists but for religious affiliates (see models 8 and 9, respectively). We run the same tests on receivers and find only a marginally significant positive INGROUP effect on WRETURN of the pooled data (see model 10), which corroborates the taste-based discrimination interpretation. Thus, we find support for Hypothesis 2.10 Ingroup love Religiosity enhances the ingroup favouritism shown by senders towards receivers of the same religious affiliation. This effect is driven by people with religious affiliations. Instead, atheists discriminate on ethnicity, which can be proxied by religious affiliation. Evidence of ingroup favouritism by receivers is marginally significant. Religious affinity ~~~~~~~~~~~~~~~~~~ Hypothesis 3 posits that religious affiliation or religiosity can serve as social identities irrespective of creed. Result 1 suggests this, but a stricter test involves data from the religious affiliation condition where there is a clear demarcation of social identity for self and other. This test distinguishes itself from previous ones in that it considers the possibility that people trust each other more so long as they both have some religious affiliation, even if they are of different religious denominations. Fig. 5 plots the linear fit of sender׳s beliefs in the trustworthiness of the receiver as a function of the sender׳s religiosity in the absence (left) or the presence (right) of religious affinity, and hints at an effect of religious affinity, which we properly test with a multivariate analysis. To test this formally, we derive the dummy variable AFFILIATE, which takes on a value of 1 when a participant who is religiously affiliated faces a task where the other is also religiously affiliated, regardless of creed. It takes on a value of zero when either the participant is an atheist or the task involves trusting an atheist. Referring to Table 4, AFFILIATE is positive and significant in model 11, showing us that religious people believe that other religious people are trustworthier than atheists. However, it is insignificant in model 12, showing us that despite this belief they are not trusted more. Model 13 includes an AFFILIATE×RELI variable and finds that such beliefs are amplified by the sender׳s religiosity. Model 16 corroborates model 12 and further shows that there is no higher order effect on actions. The effect of religious affinity on actions is weaker than that of being affiliated to the same denomination: in models 12 and 16, INGROUP and INGROUP×RELI are positive and significant, while AFFILIATE and AFFILIATE×RELI are not. This supports the arguments presented in Hypotheses 2 and 3 that connectedness enhances group identification. Beliefs only partially drive behaviour on the basis of mere religious affinity. Beyond statistical discrimination driven by beliefs, taste-based discrimination holds only if people are affiliated to the same denomination—not just by mere religious affinity. We further scrutinise the negative and significant INGROUP effect and its interaction term in model 13, which implies that religiosity diminishes the belief effect for those from the same denomination. This peculiar result of lower beliefs of trustworthiness in the ingroup is driven by atheists, as shown by our regressions on data disaggregated by atheists and religious affiliates (models 14 and 15, respectively). It suggests that atheists are more suspicious of each other, even though it does not lead to lower trust. In contrast, religious affiliates ultimately trust the ingroup more. These behaviours suggest taste-based discrimination.Religious affinity Religious affinity Senders׳ religiosity enhances beliefs about religiously affiliated receivers׳ trustworthiness regardless of whether or not they belong to the same denomination, but they do not invest more trust despite this belief. General prejudice ~~~~~~~~~~~~~~~~~ Hypothesis 4 posits that more religious people discriminate more over a range of social identities including non-religious ones. Our univariate tests examine whether more religious participants have relatively higher WTD across the different social identity categories we use. We construct, for each participant, an average WTD as the unweighted mean WTD across all of them. The correlation between average WTD and religiosity is positive and significant across both roles (ρ=0.087, p=0.0449). This relationship is significant for senders (ρ=0.123, p=0.0442) but insignificant for receivers (ρ=0.045, p=0.4658). Further, the average religiosity of those whose WTD is zero throughout the experiment (μ=33.5, n=73) is significantly less than that of others (μ=40.5, n=457, p=0.01). We also examine the correlation between religiosity and WTD across social categories (see Table 5). Again, these correlations are generally insignificant for receivers. For senders, information on religious affiliation, religiosity and ethnicity are salient and serve as social identifiers that markedly separate participants. In turn, the correlations of religiosity and the WTD along these dimensions are robustly significant. Referring to Table 6, model 17 shows that WTD is positively related to religiosity, which suggests that more religious people are more prone to religious-based discrimination. Further, we test if religious participants are generally more prone to ingroup favouritism, i.e. even if social identities of co-participants are unrelated to religion. Fig. 6 shows that both religious affiliates and atheists generally favour the ingroup over the outgroup by trusting the ingroup more across different categories of social identity. Models 17–19 test WTD and ingroup biases on data concerning all non- religion conditions. As found above, WTDs increase with religiosity (model 17). This effect is marginally significant. For beliefs (model 18), we find a positive and significant INGROUP effect for senders overall, but no RELI interaction effect. For actions (model 19), we also find a positive and significant INGROUP effect for senders overall, but no RELI interaction effect. This result is robust to controls for respective conditions.13 General prejudice Religiosity is positively associated with the general willingness of senders to discriminate across a range of non-religious social identities. However, participants of different religiosity are as prone to ingroup favouritism.","Inter-religious interaction is an increasingly important social phenomenon. However, previous experimental work has yet to establish univocal evidence regarding its direct, independent effects on trust and trustworthiness. To better understand this issue we conducted a trust game experiment across three countries with participants of different religious denominations and levels of religiosity. Our experiment was designed to test four hypotheses for indirect effects of religiosity we derived from these previous studies. Taken together these hypotheses propose that religiosity affects economic behaviour indirectly by moderating (a) the way we treat others of the same and different social groups and (b) the expectations and behaviour those we interact with develop towards us. Our main findings can be summarised as follows. First, religiosity is a strong social identifier which is used as a basis of statistical discrimination by senders of varying religiosities (Result 1). Both religious and non-religious people believe that more religious others are more trustworthy. Second, we found that religiosity enhances the ingroup favouritism people show to others who share the same faith (Result 2). Senders of all religions believe receivers of the same faith to be more trustworthy and follow these beliefs with actions in step with their own degree of religiosity. Third, we found a religious fellow feeling or affinity between religious people across different faiths, i.e. irrespective of whether they share the same one or not (Result 3). This was expressed in the greater belief people with religious affiliation have in the trustworthiness of others similarly affiliated. As before, individual religiosity amplifies this effect. This kind of religious affinity, however, does not generate quite the same positive effect on actual behaviour. Fourth, while we found that religiosity is associated with a willingness to discriminate across non-religious categories, observed ingroup favouritism did not vary with religiosity (Result 4). Since the 1950s, Adorno et al. (1950) and Allport (1954) have postulated general religious prejudice, but have since been met with scant reliable evidence. In summary, we uncovered evidence that religion operates indirectly through social identities and religious affiliation, which are used as a basis for discrimination in trust games. Religious identity is one dimension that tells decision makers how they are connected to those with whom they interact. The nature and degree of discrimination observed generally depended on the nature and degree of connectedness between individuals. The behavioural patterns we observed across the four main results showed that the closer people are the more they trust each other. Religious ingroup effects on beliefs carry over strongly to actions, in contrast to the weaker effect when religiosity was known but religious affiliation was unknown, and when religious affiliation was known (but) to be of a different creed. These effects increased with one׳s religiosity, which is an indicator of how rooted one is in a particular social group. We believe that the diversity in our participant pool lends our results good domain validity. Our study is general, as opposed to creed-specific, also in its explanation for how religion affects behaviour. In addition to the evidence relating to our hypotheses we generally found that people are willing to pay for the chance to discriminate, be it for statistical or taste motives. We designed an incentive-compatible measure of the willingness to discriminate which was shown to be significantly related to our other variables. We believe that our measure may be deployed in other social identity contexts to guide policy related to discrimination in labour markets and other specific areas."],["A central assumption in economics is that people misreport their private information if this is to their material benefit. Several recent models depart from this assumption and posit that some people do not lie or at least do not lie maximally. These models invoke many different underlying motives including intrinsic lying costs, altruism, efficiency concerns, or conditional cooperation. To provide an empirically-validated microfoundation for these models, it is crucial to understand the relevance of the different potential motives. We measure the extent of lying costs among a representative sample of the German population by calling them at home. In our setup, participants have a clear monetary incentive to misreport, misreporting cannot be detected, reputational concerns are negligible and altruism, efficiency concerns or conditional cooperation cannot play a role. Yet, we find that aggregate reporting behavior is close to the expected truthful distribution suggesting that lying costs are large and widespread. Further lab experiments show that this result is not driven by the mode of communication. © 2014 The Authors. --------------------------------------------------------------------------------","Situations with asymmetric information are ubiquitous. Most of economic theory assumes that people misreport their private information if this is to their material benefit; behavior is only determined by the trade-off between financial gains from misreporting and monetary fines when misreporting is detected.1 In contrast, many recent models in various domains of Public Economics (and in Economics more generally) rely on the assumption that people can experience a psychological disutility which holds them back from misreporting, at least to some extent. These models invoke different underlying motives. Kartik et al. (2014), for instance, assume that people face an intrinsic lying cost and show that in this case the social planner can fully implement a much wider range of social choice rules compared to the standard Maskin (1977) case without lying costs (see, e.g., Matsushima (2008) and Dutta and Sen (2011) for similar assumptions). Many studies about incentive systems for doctors assume that doctors are altruistic towards their patients and thus do not always state the profit-maximizing diagnosis but rather treat patients honestly (e.g., Ellis and McGuire, 1986; Chalkley and Malcomson, 1998). The large literature on “tax morale” (e.g., Lewis, 1982; Cowell, 1990; Andreoni et al., 1998; Slemrod, 2007; Torgler, 2007) demonstrates that many tax payers misreport their income only a little bit or not at all. This literature is usually agnostic about the exact underlying motives but some studies cite efficiency concerns (e.g., Alm et al., 1992), patriotism (Konrad and Qari, 2012), religiosity (Torgler, 2006), fairness (Bordignon, 1993), conditional cooperation (Traxler, 2010) or honesty (Erard and Feinstein, 1994). To further improve these models and to provide an empirically-validated microfoundation, it is crucial to understand the relevance of the different potential motives. Additionally, understanding these motives could inform the design of more psychologically-realistic policies, e.g., in the area of tax enforcement, that have a higher potential of being successful. In this paper, we focus on intrinsic lying costs and investigate how widespread and how large lying costs are. The ideal data set to answer these questions would allow studying lying costs for a representative sample of the population and in an environment without the confounding effects of strategic interaction (including the levy of fines), reputational or efficiency concerns, or altruism. So far, the best evidence on lying costs comes from experiments conducted in tightly controlled laboratory situations. A robust result is that many subjects misreport their private information to their own advantage but that a substantial share of subjects refrains from reporting the payoff-maximizing type and that some are fully honest (e.g., Gneezy, 2005; Charness and Dufwenberg, 2006; Fischbacher and Föllmi- Heusi, forthcoming; de Haan et al., 2011; Houser et al., 2012; Shalvi et al., 2011; Wibral et al., 2012; Serra-Garcia et al., 2013). These studies are a strong first indicator that lying costs influence behavior. However, lab experiments do not allow for inferences with respect to the prevalence of lying costs in the overall population since they have been conducted almost exclusively with student samples (DellaVigna, 2009; Falk and Heckman, 2009). Also, decision making took place in an austere laboratory environment which might trigger behavior representative only of certain non-lab situations. It could thus be that there are systematic differences between behavior of students in the laboratory and behavior of non-student subjects outside the lab. To circumvent these limitations, we measure how people report their private information outside the laboratory by calling participants on the phone at their home. Participants were drawn randomly from the German population, yielding a representative sample. An incentivized experiment was embedded in the interview. The experimental setup is related to the design of Fischbacher and Föllmi- Heusi (forthcoming) and is extremely simple: participants were asked to toss a coin and report their type, i.e., either “heads” or “tails”. Reporting tails yielded a payoff of 15 euros, which participants could choose to receive in cash or as an Amazon gift certificate, while reporting heads yielded a payoff of zero. Participants thus had a clear monetary incentive to report tails regardless of their true type. It was obvious that the true outcome was only known to the participants, as they tossed the coin privately at home. In this setup, we cannot draw reliable conclusions about the truthfulness of any individual report. But we can learn about aggregate behavior by comparing the distribution of reports to the true distribution of a fair coin (50% tails) and to the payoff- maximizing distribution (100% tails). This indirect observation therefore allows us to study the behavior of subjects in a situation in which private information is kept truly private and in which subjects do not face any risk of detection.2 Moreover, the decision is non-strategic; altruism does not play a role as the money is not taken from any individual person; and reputational concerns are minimized since the interviewer is a stranger with whom no future interaction can be expected. If all our participants were rational money maximizers, we would expect that all of them reported tails. If behavior on the phone was similar to previous comparable laboratory experiments (e.g., Houser et al., 2012), we would expect about 75% of subjects reporting tails. In contrast to these predictions, observed behavior does not statistically differ from everybody reporting honestly. If anything, participants report the payoff maximizing outcome less often than expected under truthful reporting. This latter effect, however, is small and disappears in a second treatment in which participants were asked to report the total number of tails in four consecutive coin tosses and received 5 euros times the number of reported tails. The resulting distribution of reports in the 4-Coin Treatment is indistinguishable from the distribution under complete truth-telling. Moreover, while previous studies (e.g., Dreber and Johannesson 2008) have found correlations between individual characteristics, like gender, and truth-telling, we do not find any robust correlations between individual characteristics and reporting behavior. This is not surprising if almost all participants report truthfully. Reports are solely determined by chance, namely the coin toss, which cannot be related to any individual characteristic. Our results thus show that lying costs are pervasive and are influencing behavior regardless of gender, religious beliefs, education, or age. We complement our telephone study with two additional control treatments in the laboratory to better understand what shapes lying costs, in particular the effect of the mode of communication. In both lab treatments subjects reported the outcomes of four consecutive coin tosses. Incentives were the same as in the 4-Coin Treatment in the telephone study: 5 euros times the numbers of tails reported. In the first lab treatment, subjects had to report the outcome directly to an interviewer via the phone, mirroring our telephone study. We observe the same pattern of behavior as in previous lab experiment: subjects lie much more than in the telephone study. In the second control treatment, subjects reported the outcomes by clicking a number between 0 and 4 on the computer screen as in most previous lab experiments. We find that subjects who enter their report by clicking report slightly higher numbers but this difference is not statistically significant. The difference to the telephone study persists: the average report in each lab treatment is higher than in the telephone study. This shows that the mode of communication does not systematically influence reporting behavior strongly and is not driving the widespread truth-telling in our telephone study. We also elicit beliefs about the behavior of other participants and find in all four treatments that participants believe others to lie more than they actually do. Older participants (correctly) believe that lying is less prevalent. In the lab, higher beliefs are correlated with higher own reports. We find no evidence that being a student has a significant impact on behavior, or that the perceived time pressure on the telephone or the limited experience of the survey participants with the abstract design of economics experiments played a role. Our paper adds to the nascent literature studying lying outside the lab. Previous studies focused on particular groups: Bucciol and Piovesan (2011) study a sample of children and find that many of them lie, unless they are reminded to be honest; Cohn et al. (2013) study prisoners and find that they become less honest when reminded of their criminal identity; and Utikal and Fischbacher (2013) ask a small sample of nuns to report the roll of a dice and find significant downward lying. Studies looking at unethical behavior in less abstract environments include Azar et al. (2013) who find that the majority of customers in a restaurant do not return excessive change. Similarly, Bucciol et al. (2013) study free-riding in public transportation in Italy and find that 43% of passengers evade the fare. We add two features: we study a representative sample and we can investigate the underlying motives by conducting additional lab experiments using the same well-defined decision. Taken together, our results strengthen the doubts that previous lab experiments have cast on the assumption of zero lying costs: we find evidence for even higher lying costs in the telephone study. This suggests that studying the theoretical implications of such costs (e.g., Kartik et al., 2007, 2014; Doerrenberg et al., 2013) is a promising research avenue. At the same time, it is very likely that altruism, efficiency concerns, etc. are also important factors in the decision to pay taxes or how to treat patients, for example. Future research would need to investigate the relative importance of different motives that hold people back from misreporting and the interactions between motives. Our results also do not mean that lab experiments are uninformative about non-laboratory settings. However, the difference in behavior between our telephone study and our previous lab experiments rather shows how malleable reporting behavior can be. This opens many new questions about how exactly reporting private information depends on the decision-making context. Intuitively, different norms might apply when making such a decision at home, representing a private and familiar environment. Similarly, people could be more attentive to their own moral rules, e.g., abstaining from lying when at home.3 Irrespective of these differences between lab and field, our study establishes that lying costs are more important than previously assumed and are strongly influencing behavior across different decision environments. In the next two sections, we present the design of the study and our hypotheses. Section 4 contains the results. We discuss policy implications in Section 5.","The computer-assisted telephone interviews were operated by the Institute for Applied Social Sciences (infas), a private and well-known German research institute. They were conducted between November 2010 and February 2011.4 The average interview lasted 20 min (standard deviation: 5.5 min). Telephone numbers were selected using a random digit dialing technique: numbers were generated randomly based on a data set of all potential telephone numbers in Germany. Only landline numbers were used in this study, as 92.7% of German households have a landline number (Destatis, 2012). The selection of the participant within each household was also random: only the member of the household whose birthday was the most recent among all household members was eligible to participate. We restricted participation to those aged between 18 and 70 years at the time of the interview.5 The survey was split into two parts. The first part of the questionnaire consisted of questions relating to the participants' socio-demographic background and their risk and trust preferences. Risk and trust preferences were measured by using subjective self-assessments, using the general risk question of the GSOEP (“How do you consider yourself? Are you in general a rather risk-loving person, or do you try to avoid risks? Use a scale from 1, meaning that you are not at all willing to take risks, to 7, meaning that you are absolutely willing to take risks.” (Dohmen et al., 2011)) and the World Value Survey trust question (“Generally speaking: Do you think one can trust other people, or that one should rather be careful when dealing with other people? Please indicate your answer on a scale from 1 to 7, with 1 meaning that one should be careful when dealing with other people, and 7 meaning that one can trust other people.”). After this part, the experiment described below took place. After the experiment, participants were asked about their political preferences, their current living and financial situation, their religious beliefs, and their attitudes towards opportunistic behavior and everyday crime. At the very end of the interview, participants were asked to state their belief about other participants' behavior in the experiment. Before the experiment started the participant was reminded that the resulting data would be anonymized, and that infas and the University of Bonn guaranteed the correct payment. The interviewer then asked the participant to take a coin and explained the rules of the experiment: the task was to toss the coin and report whether heads or tails came up.6 If the participant reported heads, they received no payment. If the participant reported tails, they would receive 15 euros. Then, the participant was asked to toss the coin and report the outcome. We will call this treatment “1-Coin-Telephone.” 658 people participated in this version of our experiment. A translation of the exact experimental instructions can be found in Online Appendix A. In a second treatment, 94 people were interviewed and participated in the following variation of the experiment. Participants were asked to take a coin, toss it four times, and report the number of times that tails came up. For each time participants reported tails they received 5 euros. Thus, they could earn 0, 5, 10, 15, or 20 euros. We will call this treatment “4-Coin-Telephone.” Payment in both treatments could be received either in cash via regular mail or as an Amazon gift certificate code. The alphanumeric 14-digit gift certificate code was transmitted via email or directly on the phone at the end of the interview. In order to further investigate what influences behavior in the telephone study, in particular the mode of reporting, we additionally conducted two versions of the 4-Coin Treatment in the laboratory. Subjects were students of the University of Bonn studying different majors except Economics. They were seated at a desk with a computer in separate room-high cubicles closed off by curtains. As the experiment took only a few minutes, it was run at the end of the sessions of a different experiment (similar to Fischbacher and Föllmi-Heusi, forthcoming). In the preceding experiment subjects made abstract consumption or labor supply choices which involved no private information and no interaction with other subjects. When the experiment started, subjects were asked to take a coin that was placed in their cubicle, toss it four times, and report how often tails came up. For each time they reported that tails came up they received 5 euros, i.e., up to 20 euros, just like in 4-Coin-Telephone. Their earnings were paid in cash directly after the experiment.7 The only difference between the two lab treatments was how the reporting was done. In the first treatment, subjects had to state their report directly to an interviewer via the phone, mirroring our telephone study. After tossing the coin in their cubicle, they were asked to go one-by-one to an adjacent room and pick up the telephone that we had placed there. An interviewer on the other side of the line (whom subjects never met directly) would then ask for their experimental ID and the number of times the coin showed tails. We made sure that other subjects could not hear the conversation. The starting times for the coin tossing was staggered, such that subjects did not have to wait between coin-tossing and reporting. 170 subjects participated in this treatment which we will call “4-Coin-Lab-Tel.” This treatment serves to replicate our telephone study as closely as possible in the laboratory. In the second treatment, subjects reported their outcome by clicking a number 0 to 4 on the computer screen, similar to previous lab experiments. 180 subjects participated in the second treatment which we will call “4-Coin- Lab-Click”. This treatment serves to investigate whether the mode of communication, i.e., clicking on a computer screen versus reporting to a person via the telephone, influences reporting behavior.","The standard economic prediction in our setup is straightforward: depending on the treatment, people will report tails one or four times, respectively. This is the payoff maximizing outcome as there are no exogenous costs linked to misreporting, no possibility of detection and no fines. The setup is extremely simple and participants should have no trouble identifying the payoff maximizing choice. Moreover, the setup is highly anonymous, discouraging any reputational concerns because of repeated interaction. If, however, some participants incur a psychological cost or derive direct disutility from falsely reporting their private information per se we should expect both heads and tails to be reported in the experiment. There are a few recent theoretical papers that assume such a cost. For example, Kartik (2009) and Kartik et al. (2007) build on Crawford and Sobel's (1982) cheap-talk model and derive predictions for the case that some agents incur costs when misreporting their private information (see also, e.g., Saran, 2011; Kartik et al., 2014). Assuming some degree of heterogeneity in the incurred costs when misreporting, it is then a question of the trade-off between psychological costs and monetary benefits of misreporting how many participants will report heads and how many report tails. Participants in 1-Coin-Telephone have to make a clear, binary choice whether to lie or not; if lying costs are related to self-reputation or identity (e.g., Bénabou and Tirole, 2006; Akerlof and Kranton, 2000), lying in such a setting could impact self-reputation or identity more and thus make lying more costly. Participants in 4-Coin-Telephone can make a finer choice between being honest, exaggerating a little bit, or lying maximally; this could render small lies compatible with a positive self-reputation and thus enhance lying (Mazar et al., 2008). Such non-maximal lying has already been shown to be important by Fischbacher and Föllmi-Heusi (forthcoming). In the telephone study, participants tossed the coin at their home. It was thus obvious that the interviewer could not secretly observe the true outcome of the coin toss.8 If some participants in our lab experiments (erroneously) believed that the experimenter could observe the true outcome and believed (again erroneously) that misreporting would lead to some negative or unpleasant outcome, we would expect more truth-telling in the laboratory.9 Regarding potential differences in reporting behavior according to individual characteristics, we would expect that women are more honest than men (as already shown by Dreber and Johannesson, 2008; Houser et al., 2012). More religious participants would be expected to be more honest, since religious priming leads to less lying and more pro-social behavior (Mazar et al., 2008; Shariff and Norenzayan, 2007). Income could be positively correlated with honesty because of the lower marginal utility of the monetary rewards or negatively correlated because of reverse causality. A similarly ambiguous hypothesis can be derived for education or the social environment, e.g., the size of the community or family status. Along theories of endogenous social norms (e.g., Traxler 2010; López-Pérez, 2010, 2012), we would expect that higher beliefs about the reporting of other participants are correlated with own high reporting. Telephone study ~~~~~~~~~~~~~~~ In 1-Coin-Telephone, the distribution of actual reports is very close to the truthful distribution; participants report the payoff-maximizing outcome slightly less often than expected if everyone reported truthfully. In 4-Coin-Telephone, the distribution of reports is indistinguishable from the truthful distribution. Fig. 1 illustrates aggregate behavior (the dashed line corresponds to the expected distribution if every participant reported the true outcome of the coin toss). 55.6% of participants report heads as the outcome of the coin toss, yielding a payoff of zero, the remaining participants report tails yielding a payoff of 15 euros. The payoff-maximizing outcome is reported slightly less often than in 50% of the cases and although the difference is small in terms of effect size, it is significant (Binomial test, p = 0.004) Fig. 2 shows aggregate behavior in 4-Coin- Telephone. Again, reporting behavior follows the expected distribution under complete honesty very closely (the dashed line corresponds to the truthful distribution). In fact, the distribution of reported outcomes is statistically indistinguishable from the truthful distribution (Kolmogorov–Smirnov test, p = 0.61; binomial tests of the expected against the observed frequency, all five p > 0.13). In particular, and unlike in 1-Coin-Telephone where “too many” people report the payoff-minimizing outcome, there is no significant over-reporting of zero in this treatment.10 Looking at behavior in both treatments we can therefore summarize that the payoff-maximizing outcome is reported by much fewer participants than expected if no one incurred lying costs. It is also reported less often than suggested by previous lab experimental studies, which find some truth-telling but also many instances of the payoff-maximizing report. Instead, it is close to the distribution that would arise if every participant reported his or her type truthfully.11 Previous studies have shown that truth-telling correlates with observable characteristics, e.g. gender or religiosity (Dreber and Johannesson, 2008; Houser et al., 2012; Mazar et al., 2008; Shariff and Norenzayan, 2007). In contrast, if our conjecture that almost all participants report truthfully is correct, an individual's reported outcome will only be driven by their random coin toss; if this is the case, reporting cannot be correlated with any individual characteristic, as these are orthogonal to the chance move. Therefore, if we do not find such a correlation, our finding of (almost) complete honesty is supported. More specifically, we conduct regression analyses for the two experiments in order to examine whether there are systematic effects of individual characteristics on reporting behavior. First, we regress the report only on clearly exogenous variables such as age and gender, in a second step adding religious denomination. We then include income, the size of the city the individual lives in, and education dummies. Finally, we look at the effect of an individual's religiousness (interacted with denomination), their risk and trust preferences, and their belief about the reporting behavior of other participants. There is no significant correlation between reporting behavior and any individual characteristic. First, we look for potential group differences in terms of reporting behavior in 1-Coin- Telephone by conducting Probit regressions of the reported outcome on the respective characteristics (see Table 2 in Online Appendix E). No characteristic except for one's belief about others' behavior is significantly associated with reporting in the experiment: participants who think many other participants report tails dishonestly, are less likely to report tails themselves. This belief is, however, not significant if we include it as the only explanatory variable (p = 0.15). Note in particular that neither gender nor any religion-related variable is significantly correlated with reporting. Conducting the same regressions as in Table 2 using OLS leaves the results unchanged. Next, we check whether these results also hold in 4-Coin-Telephone. We run Ordered Logit regressions of the reported number of tails on the same explanatory variables as before. Table 3 in the online appendix illustrates the results from this estimation. Only the coefficient for trust is significant. This effect is, however, not robust to the inclusion of other explanatory variables. The effect is also not present in 1-Coin-Telephone. In contrast to 1-Coin-Telephone, the belief coefficient shows no significant association with reporting behavior in this treatment and the point estimate has the opposite sign. We will discuss the data on beliefs in more detail in Section 4.3. Two further aspects of our analysis are worth noting. First, when running OLS regressions using the same predictor variables as above, we find that only two of the 10 specifications have an adjusted R2 above 0 (below 0.004), all other adjusted R2 values are negative. Moreover, the resulting adjusted R2 tends to decrease in the number of included variables. This again underlines our conclusion: the tested predictor variables do not increase explained variance in the dependent variable compared to pure chance. Second, we also test the correlations between reported number and answers to the survey questions that we did not include in the main specifications of Tables 2 and 3. These include a person's citizenship and country of birth, various personal characteristics, a person's current job or educational situation and their current or recent position in the professional hierarchy, a person's willingness to tell white lies in different situations, a person's family status and living situation (whether one lives with a partner and the number of people belonging to the household), the frequency of church attendance, a person's political preference, and the individual's tendency to behave in an opportunistic way as well as the belief about others' willingness to behave like that. Testing these variables as predictors in Probit and Ordered Logit regressions in the two different data sets, akin to Tables 2 and 3, we find no robust association between any of them and reporting behavior. In particular, this means that students and non-students do not behave differently in our sample. This holds when we consider current students or include former students as well (Kolmogorov–Smirnov tests, all p > 0.409). It is thus not a student vs. non-student difference, e.g., a difference in education, age, cognitive skills, or socio-demographic background, which drives the difference between our results and previous lab experiments. Summing up, the overall picture is confirmed: no individual characteristic, whether exogenous or endogenous, is systematically associated with reporting behavior suggesting that almost all participants in our study tell the truth. It could still be that a subgroup of people, which we cannot identify with our background information, reports tails more often than actually true while another subgroup reports tails less often. This could result in the two effects offsetting each other, which would result in a similar picture of aggregate behavior. However, we consider this to not be likely as our analysis shows that this is not the case for any of the numerous subgroups that we can identify with our data. More importantly, such an effect would further need to recreate the distinct distributions of Figs. 1 and 2 which is implausible. Laboratory experiment ~~~~~~~~~~~~~~~~~~~~~ To further investigate the motivations underlying behavior in the telephone study, we conducted two 4-Coin Treatments as laboratory experiments. We will first discuss the 4-Coin-Lab-Tel treatment which keeps the mode of communication as in the telephone study: subjects had to report their result over the phone directly to an experimenter.12 Subsequently, we compare this treatment to 4-Coin-Lab-Click, in which subjects reported their number by clicking a button on a computer screen as in previous lab experiments. This second comparison will allow us to disentangle the influence of the mode of communication.13 Subjects in 4-Coin-Lab-Tel report substantially higher numbers than subjects in 4-Coin-Telephone. The upper panel of Fig. 3 shows aggregate behavior in 4-Coin-Lab-Tel: most subjects refrain from reporting the maximal outcome, forgoing on average 6.82 euros, quite a considerable amount compared to the average hourly student wage in Germany of about 10 euros. At the same time, behavior is significantly different from the distribution expected under truthful reporting, the dashed line in the figure (Kolmogorov–Smirnov test, p < 0.001; binomial tests, all five p < 0.009). This replicates previous findings in the lab: many subjects lie but often not maximally. Reporting behavior also deviates strongly from what we have observed in the telephone study: reports are significantly higher in 4-Coin-Lab-Tel than in 4-Coin-Telephone. In Table 1, columns 1 and 2, we regress the reported number of tails on a dummy for being in the lab, a dummy for 4-Coin-Lab-Click and controls for age and gender. The lab dummy is highly significant.14 We find the same result if we compare 4-Coin-Telephone and 4-Coin-Lab-Tel using a two-sample Kolmogorov–Smirnov test (p < 0.001). These results demonstrates that our 4-coin randomization mechanism does not drive the truthful behavior in 4-Coin- Telephone and that, by moving our telephone setup to the laboratory, we are able to strongly change behavior (as we showed within the telephone study, this is not driven by subjects being students per-se). How big is the additional effect if we also change the mode of communication? Subjects in 4-Coin-Lab-Click report slightly higher numbers than subjects in 4-Coin-Lab-Tel but this difference is not statistically significant. Only the report of 4 occurs significantly more often in 4-Coin-Lab-Click; the reports of 0, 1, 2, and 3 are not different across treatments. Reports in 4-Coin-Lab-Click are significantly higher than in 4-Coin-Telephone. The lower panel of Fig. 3 shows aggregate behavior in 4-Coin-Lab-Click. The distribution of reports is very similar to the one in 4-Coin-Lab- Tel, the average report being only slightly higher (2.78 in Click vs. 2.64 in Tel). The overall distribution and the average report are not significantly different across the two treatments (two-sample Kolmogorov–Smirnov test, p = 0.136; Ordered Logit in Table 1, columns 1 and 2, both p > 0.096). The share of subjects reporting 0, 1, 2 or 3 are also not significantly different (tests of proportion, all p > 0.100). However, subjects in 4-Coin-Lab-Click report 4 significantly more often (p = 0.007).15 At the same time, behavior in 4-Coin-Lab-Click is markedly different from 4-Coin-Telephone (two-sample Kolmogorov Smirnov test, p < 0.001, Ordered Logit in Table 1, columns 1 and 2, F-test, both p < 0.001). Overall, our data thus show that the mode of communication does not have a strong effect on behavior and cannot explain the difference between our telephone study and previous lab experiments. This result is further confirmed by Waubert De Puiseau and Glöckner (2012) who also find considerable truth-telling at home, though not as extreme as in our data, using an online panel in which participants answered questions at home by clicking on a computer screen. Houser et al. (2012) conduct a 1-coin lab experiment and find similar levels of lying as in our lab experiments, replicating the other side of our results. One could think that one reason why behavior in the telephone study differs is a perceived time pressure on the telephone which might make lying more difficult. However, we measure response times in the laboratory and do not find a correlation with the report (Ordered Logit, p = 0.108).16 If anything, the report in the lab is higher for short decision times. This mirrors results of Shalvi et al. (2012) who impose exogenous time pressure in a similar lab experiment and who find that subjects become less honest under time pressure. Taken together, these results suggest that behavior in the telephone study is not driven by perceived time pressure. We also find no correlation of the number of previous participations in other lab experiments with reporting behavior in the lab (p = 0.578), suggesting that the limited experience participants of the telephone study have with the abstract design of economics experiments does not play a role. It rather seems that different norms apply when making a reporting decision at home, representing a private and familiar environment, compared to in the lab where other, more selfish norms might be triggered.17 We showed above that women do not report differently from men in the telephone study. As one can see from Table 1, women do report lower numbers in the lab. This effect is only weakly significant in the sample of all three 4-Coin Treatments, i.e., also including 4-Coin-Telephone which dilutes the effect, and becomes significant if we restrict the sample to the two lab treatments (p = 0.027 and p = 0.046 in regressions akin to columns 2 and 4 of Table 1). Beliefs about other participants ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Previous studies (e.g., López-Pérez, 2010; Diekmann et al. 2011) have investigated the relationship of reporting behavior and the beliefs about what other people report. Since our telephone and lab settings generate strong differences in reporting behavior, we next examine whether there is a similar difference in beliefs and whether this could help explain the differences in behavior. In all four treatments, we elicited beliefs about the reporting behavior of the other participants. We will mainly focus on analyzing beliefs in the 4-Coin Treatments as the outcome variable is richer and we have additional treatments. In the 4-Coin Treatments, subjects were asked two questions regarding their beliefs about the behavior of other subjects in their treatment (the question referred to 1000 participants in 4-Coin-Telephone): “We are conducting this experiment also with 100 other participants. How many of these 100 participants do you think report tails more often than they actually tossed?” and “How many of these X overreporting participants do you think report that they tossed tails in each of the four coin tosses?”18 We will use the answers as direct measure of the belief about the share of liars and about the share of maximal liars. Using a very simple model, we can also combine the two answers to back out the implied belief about the average report in the population or the share of participants reporting the payoff-maximizing outcome. The model assumes, similar to Kartik (2009) and Mazar et al. (2008), that participants expect others to face a psychological lying cost which is increasing and convex in the size of the lie and which might be heterogeneous between participants. Online Appendix D describes the model and the belief measures in more detail. In all treatments, participants believe others to overreport more than they actually do. Fig. 4 compares average beliefs with average actual behavior for each treatment. We take as variable of interest the share of participants who report the payoff-maximizing outcome, i.e., 4 tails in the 4-Coin Treatments and tails in 1-Coin- Telephone. Since we expected participants to be unfamiliar with the true distribution of the sum of four coin tosses, we didn't ask directly for their belief about this share. We are able to calculate it, given the assumption of convex lying costs, from the two questions for the 4-Coin Treatments: it is the share of liars who report 4 (question 2) plus the share of honest 4's (the probability of a true 4 times (1 — answer 1)). Since we do not observe whether an individual overreports we cannot directly compare the two answers to actual behavior.19 We find in all four treatments that participants believe that others overreport more than they actually do. The differences are highly significant (t-tests, all p < 0.001). The same results obtain when we consider the average reported number as variable of interest.20 What shapes these beliefs? Older participants believe others to overreport less. Participants expect more overreporting when participants can enter their report by clicking on the screen. In Table 1, columns 5 and 6, we regress the answer to the second question on treatment dummies, a gender dummy and age (the table only considers the 4-Coin Treatments). We find that subjects in 4-Coin-Lab-Click believe others to overreport more than subjects in 4-Coin-Lab-Tel. Being in the laboratory seems to increase beliefs (column 5) but this effect goes away once we control for age (column 6). Participants in the telephone study are on average much older than the student sample in the lab and older participants expect others to overreport less. This means that the beliefs of older participants in the telephone study are closer to actual behavior than the beliefs of younger participants. The same effect of age is present in 1-Coin-Telephone (p < 0.001). Using the answer to the first question, or the belief about the average report or the belief about the share of participants reporting 4 does not change any of the results. In the lab, participants who believe that others report high numbers also report higher numbers themselves. We discussed above that there is no robust correlation between reports and beliefs in the telephone study. In Table 1, columns 3 and 4, we study the correlation of reports and beliefs for the 4-Coin Treatments in lab and field. We regress the reported number of coin tosses on treatment dummies, controls for gender and age and on the answer to the second belief question. We find that participants who believe others to report high numbers also report higher numbers themselves. If we exclude 4-Coin- Telephone from the analysis, the coefficient on the belief variable becomes even bigger and stays significant. One could interpret this finding as yet another indication that almost all participants are honest in the telephone study because, if some were not, we should also find a correlation with beliefs in the telephone study (similar to the gender effect we do not find). Furthermore, since beliefs are on average higher in 4-Coin-Tel- Click, the difference between the two lab treatments – which is barely significant in column 1 – becomes even smaller once we control for beliefs. These results are again robust to the exact belief measure we use. The direction of causality between beliefs and behavior is obviously unclear in our setting. On the one hand, it could be that a high belief induces participants to also report higher numbers. This would be in line with a notion that moral norms are endogenous to the beliefs people hold about the behavior of their peers (see, e.g.,Traxler, 2010; López-Pérez, 2010, 2012). Diekmann et al. (2011) provide causal evidence that higher beliefs lead to higher reports. If this is the mechanism for the correlation between beliefs and behavior, it is even more surprising that participants, in particular in the telephone study, decided to refrain from exploiting the opportunity to receive a considerable amount of money when they believed that many others would do so. On the other hand, the causality might run in the opposite direction if participants ex-post justify their own high report with a stated belief that others also overreport.","Using a representative sample of the German population we conducted telephone interviews during which respondents participated in an incentivized experiment. Depending on the treatment, they could earn money by reporting tails as the outcome of one or four coin tosses. We find that almost all participants report their coin toss(es) honestly: the distributions of reports are extremely close to the true distribution of a fair coin toss or four coin tosses, respectively. Moreover, reports are not correlated with any individual characteristic, including gender which has been shown to predict honesty in previous lab studies. We conducted additional laboratory experiments to study the motives underlying the behavior on the phone. While reports are generally higher in the lab than in the telephone study we find little evidence that the mode of communication (reporting directly to someone via the phone vs. clicking a number on a computer) influences behavior. Being a student has also no effect. Our results underline doubts about the generalizability of economic models which assume that people always lie maximally when it is financially beneficial. Apparently, people do not only care for the trade-off between financial gains from misreporting and the monetary fines when misreporting is detected (cf. Becker, 1968). Our results instead support models like Erard and Feinstein (1994), Kartik et al. (2014) or Doerrenberg et al. (2013) which assume that many people do not lie or do not lie maximally; intrinsic lying costs could be a potential microfoundation for these and similar models. The effect of patriotism and religiosity on tax morale (Konrad and Qari, 2012; Torgler, 2006), for example, could also work through an increased lying cost. The strong differences we find between telephone and lab environment suggest that lying costs are stronger in our setting outside the lab. It seems that different norms apply when reporting private information at home. Similarly, it might be that the familiar and intimate environment of one's own home reinforces one's personal identity and renders personal moral standards more salient. This is in line with recent evidence by Cohn et al .(2013) who conduct a similar experiment with prisoners. They find that priming prisoners with their criminal identity reduces honesty. Lab experiments, in turn, could be more representative of decisions for which people take on a particular role or identity in addition to their private identity. At the same time, this study does not imply that everybody always reports their private information truthfully. The level of lying costs seems rather to be influenced by the context in which people are asked to report their type (see also Mazar and Ariely, 2006; Mazar et al., 2008). The difference in behavior on the phone and in the lab shows how malleable reporting behavior can be. Our results therefore point to important policy implications: institutions, e.g., tax authorities, could make use of the context dependence of reporting behavior when designing decision- making environments. As we find strong evidence for widespread lying costs, appropriate mechanisms might be much less complex than those resulting when assuming that agents have no qualms about lying. It might be possible to change reporting behavior in simple and low-cost ways in the spirit of libertarian paternalism (Thaler and Sunstein, 2003). Further research is necessary to uncover what the crucial aspects of the decision-making environment are that induce truth-telling."],["The existence of an asymmetric history between bargainers can trigger self-serving beliefs about the fair settlement of a subsequent dispute, ultimately leading to bargaining impasse. In a two-stage bargaining experiment we demonstrate that dyads who share a history that produced wealth asymmetries between them are less likely to settle in a subsequent negotiation than when the same wealth asymmetry stems from negotiators’ independent histories. When negotiators share an asymmetric history, the individual who previously lost out in the first stage believes that s/he deserves compensation in the second stage, but the individual who prevailed in the first stage believes that compensation is not warranted. These divergent, self-serving views about a fair settlement – and the resulting irreconcilable demands – lead to bargaining impasse. We find, further, that unbiased judges side with the losers in the first stage; they believe that it is fair for them to be compensated in the second stage. Indeed, this is true, albeit to a lesser extent, even if the winner and loser had not directly interacted with one another – i.e., if their history is not shared. --------------------------------------------------------------------------------","Examples of a dispute in which one party self-servingly invokes the past – for example, that one put in a lot of overtime to meet a recent project deadline and thus deserves some “late mornings” – are easy to call to mind. These situations are not just unique to individuals. On the international stage, politicians often invoke history to justify demands that conflict with those of other countries. In discussions of European refugee quotas, for example, some argued that their country did not benefit from colonization of the country from which refugees originated, and therefore they should bear a smaller fraction of the burden of the refugee quota. Needless to say, such claims were rejected by leaders of the countries that had participated in, and likely benefited from, such colonization. Differences in the relevance of historical carbon emissions between developed and developing countries likewise complicate climate change negotiations over different countries’ fair burden in reducing greenhouse emissions. These are examples of a large spectrum of disputes that are intensified by, and in some cases specifically revolve around, self-serving views about the implications of the past for a current situation. At issue is whether and to what extent one party's prior outcome should have a bearing on the fair division of current resources. We are specifically interested in situations in which one party (the “loser”) lost to another party (the “winner”) in a previous allocation of resources. We ask whether, when the same parties enter into a new negotiation, each party views compensation for the prior loss as fair, and the extent to which divergent views about the implication of the past for the fair settlement of the current dispute are associated with increased likelihood of impasse. Although the history of interactions between players has been examined in prior game-theoretical research, the focus has been on strategic behavior to build reputation (e.g., Roth and Schoumaker 1983). By contrast, our interest in history focuses on whether and how different sequences of prior outcomes affect subsequent negotiations via divergent conceptions of fairness (e.g., Dezső et al., 2015). We argue, and experimentally demonstrate, that whether there is a self-serving invocation of an asymmetric history leading to impasse depends on whether negotiators’ wealth asymmetry stems from their shared (i.e., interdependent) or independent history. By a shared asymmetric history, we mean a situation in which the winner specifically won out over the loser in a prior distributive situation. In such situations, the disadvantaged party may feel that the advantaged party benefited at his/her expense. By contrast, when there is asymmetry from independent histories, one party was the winner and the other party the loser in the prior interaction, but the winner did not win at the loser’s expense. Results of our bargaining experiment demonstrate that when negotiators’ ex-ante asymmetry is due to their shared histories, the loser believes that s/he deserves compensation from the winner, but that the winner rejects this perspective. This divergence of negotiators’ beliefs about the fair solution increases the likelihood that they will fail to settle the subsequent negotiation. In contrast, when the winner's and loser's histories are independent, they are likely to hold more convergent views of fairness and be more likely to settle the current dispute. To obtain an impartial view of fairness, unencumbered by self-serving bias, we recruited an independent sample of subjects to provide an unbiased perspective on the fair solution to the dispute in the two asymmetric conditions – i.e., in which the history between the parties is shared or independent. As predicted, we find that the amount of compensation seen by these impartial judges as fair is greater for losers when the asymmetry stems from negotiators’ shared histories versus independent histories. Comparing players’ beliefs to those of the unbiased judges, we find that when negotiators share an asymmetric history, the fairness views of both winners and losers give the loser less than the amount judged as fair by the impartial judges. However, the losers’ fairness judgments are closer to the impartial judges than the winners’ judgments. When the histories are independent, on the other hand, winners’ views of fairness coincide with the views of unbiased judges, but losers believe they should be entitled to more compensation than do the impartial judges. In what follows, we briefly summarize relevant literature. Then, we present the framework of the bargaining experiment and our predictions. These are followed by a detailed description of the methods. Next, we present the results of the experiment. We conclude with a discussion of the broader implications of our findings.","According to equity theory, people prefer outcomes of joint activities to be equitable, meaning that rewards are proportionate to inputs, and, when they perceive inequity, will engage in efforts at restoring equity (Adams, 1966; Homans, 1974). Equity theory is supported by research showing that people who perceive themselves as having been treated inequitably are more likely to lie and cheat in an effort task to restore equity. In psychology, see, Gino and Pierce (2009), Greenberg (1993), John et al. (2014), Sharma et al. (2014), for instance. In economics, see, e.g., Fehr et al. (1993) and Houser et al. (2012). Greenberg (1993), for instance, reports that workers compensate themselves by stealing from a company's inventory if they feel unfairly underpaid. Houser et al. (2012) finds that people who received an inferior allocation in a dictator game, from a human rather than a random device, are more likely to lie in a subsequent task to increase their experimental earnings. In situations in which parties jointly create the to-be-divided resources and the individual contributions are known, division proportional to contribution is typically the salient norm, consistent with equity theory (Cappelen et al., 2017, 2013a, 2007; Karagözoglu and Riedl, 2014; Konow, 2000). There are, however, complex situations in which unequal contributions are due to uncontrollable factors such as bad luck. The question of what constitutes equity arises here and sets the stage for context-dependent views on a fair division (Konow, 2001). One solution is to compensate the disadvantaged party through redistribution to the extent that her/his lower contribution is caused by uncontrollable factors (Cappelen et al., 2013b; Fong, 2001; Konow, 2001). The degree of freedom in deciding whether and how the asymmetry should be dealt with, however, leads to the emergence of multiple fair solutions not only among stakeholders (Thompson and Loewenstein, 1992) but also among neutral judges with no personal interest in the distribution (Cappelen et al., 2007). As long as negotiators are flexible in their claims, agreement is likely to occur. There are, however, situations in which claims are fueled by inflexible and divergent beliefs about fairness (Birkeland and Tungodden, 2014; Camerer, 2003). When notions about the fair settlement of a negotiation are self-servingly biased, bargainers are convinced that their view of a fair settlement coincides with the unbiased fair solution and they refuse to compromise (Babcock et al., 1995; Birkeland and Tungodden, 2014; Konow, 2005). This central role of fairness beliefs in negotiations has been advanced in a theoretical framework proposed by Birkeland and Tungodden (2014). The authors incorporate fairness motives and the weights that people attach to them, and generate predictions of bargaining outcomes in different situations. Here, we focus on their proposition regarding principled disagreement, which occurs when negotiators hold incompatible beliefs about what would be a fair settlement, and insist on these beliefs to the point where they reach stalemate. In this paper, we combine the research on compensation-seeking after experiencing a loss with that addressing the key role in negotiation of beliefs about fairness. We ask whether a prior asymmetric distributive outcome between negotiators would be more likely to lead to incompatible beliefs about the fairness of compensation in a subsequent negotiation. Furthermore, we ask whether this situation is more likely to lead to impasse when the asymmetry is due to an allocation schema in which negotiators’ outcomes were interdependent than when their outcomes were independent. Rather than creating asymmetries between negotiators via lopsided potential contributions to the to-be-divided proceeds (as done, for instance, by Cappelen et al., 2013a), we establish the asymmetry between negotiators with respect to their initial wealth levels. Negotiators are subject to a lopsided allocation of proceeds generated through a real-effort task in a prior distributive situation, and we test if this history spills over onto their subsequent bargaining behavior. We view asymmetries in the bargainers’ prior outcomes as a context for the focal negotiation, in which the interpretation of the significance of the prior outcomes is likely to differ between them. The loser is likely to believe that the prior outcome has a bearing on the fair distribution, and that s/he is entitled to compensation. By contrast, the winner is likely to believe that the past is irrelevant to the present and that the outcome of the negotiation should depend only on factors relating to that negotiation. To the best of our knowledge, only two previous studies address the impact of ex-ante asymmetries on bargaining behavior. The first is a lab study from Camerer and Loewenstein (1993, Study 1 – Appleton-Baker). They demonstrated increased likelihood of impasse between buyers and sellers who had the opportunity to renegotiate a sales transaction after their BATNAs (best alternatives to negotiated agreement) were revealed. These BATNAs were treatment variables, randomly assigned to subjects at the beginning of the experiment, and subjects were not allowed to reveal them during the initial negotiation. After parties first negotiated the sales prices, their BATNAs were revealed and they were prompted to renegotiate the price. Here, the ex-ante asymmetry between bargainers arose from one party benefiting more than the other from the initial sales transaction in the light of the revealed BATNAs. The authors’ key finding is that pairs with greater ex-ante asymmetries were more likely to reach impasse in the renegotiation. Those subjects who realized that they made a disadvantageous initial deal upon BATNAs becoming public knowledge tried to get compensation in the renegotiation. However, their claims were unwelcomed by their bargaining partners who made advantageous initial deals. The authors speculated that this is because negotiators hold self-serving views about the implication of their sales histories on the current bargaining. Importantly, in this study, the ex-ante asymmetry was not randomly assigned but rather was partly the outcome of the initial bargaining, raising the issue of selecting for weak and strong bargaining skills. Therefore, it could not be determined if the bargaining impasse was due to the initial asymmetry, or rather to a factor that was associated with creating the asymmetry in the first place. In other words, it is possible that the same characteristic accounts for weak negotiating skills that put someone in an inferior position after the initial sales transaction and for later claiming compensation upon learning how bad of a deal one made. Moreover, as the authors did not elicit subjects' beliefs about the unbiased fair solution, one cannot tell whether and to what extent stalemates were driven by divergent and potentially self-serving beliefs about the fair compensation for the person who lost out in the initial transaction. The second relevant study is Experiment 2 reported by Dezső et al. (2015). Here, the authors did manipulate prior asymmetries between bargainers and found that when their asymmetric histories are linked together, they are more likely to reach an impasse than when the asymmetries are due to a prior interaction with another party. In this earlier study, as in the current study, bargainers arrived at the table with asymmetric wealth levels due to a previous, lopsided distribution of their joint proceeds. In the first stage of the experiment, both subjects took the same trivia quiz, and each of their scores was added together to determine total joint earnings. Then, the person who had obtained the higher score received all the joint earnings and the person with the lower score received nothing (in the event of a tie, the individual who received total earnings was determined by chance). Then, in stage two of the same-partner condition, winners and losers remained paired. However, in the different-partner condition, winners and losers were paired with someone else – always ensuring winner-loser pairs. In both conditions, the asymmetry was due to the winner having benefited at the expense of a loser, but, in the different- partner condition, the winner benefited at the expense of a different loser than of his/her stage two pair. The key finding from this study was that impasse was more likely in the same- than different-partner condition. The authors speculate that this is because losers and winners in the same-partner condition held incompatible views about the fair compensation to the loser. This, however, could not be addressed because negotiators’ beliefs about the fair division were not elicited, so it was unclear whether, and to what extent, impasses were driven by self-serving beliefs about a fair settlement. Furthermore, negotiators in the Dezső et al. (2015) study were uninformed about their actual stage one and two contributions (i.e., quiz earnings), obscuring whether they sought shares according to the proportionality rule. Additionally, allocations in the first stage were not random, but were determined by which individual got a higher score on the quiz. So there was, as in Camerer and Loewenstein (1993), a connection between the personal characteristics of the disputants (specifically, their skill at the task) and the outcomes they experienced. Moreover, at the beginning of stage one, subjects were uninformed about the forthcoming unfair allocation schema, which calls into question the extent to which frustration, surprise, and relief may have influenced the observed bargaining behavior. Finally, as fair divisions were not elicited from neutral judges, it is unclear if redistribution is generally viewed as a fair solution between negotiators with asymmetric histories. In addition to addressing the limitations of these previous two studies, we examine the nature and role of beliefs about the fair settlement in bargaining between parties with asymmetric history. To disentangle the effect of asymmetry from interdependency (i.e., losing to a beneficiary), we manipulate whether asymmetric prior wealth levels between bargaining dyads are due to histories that were shared, such that only one party could win at the expense of the other, or independent, such that both could have won or lost (in the event that one won and the other lost, the winner did not win at the expense of the loser). We call the former the “shared asymmetric history condition” and the latter the “independent asymmetric history condition.” We suspect that losers in the shared asymmetric history condition view it as fair to obtain compensation from the winner, but that compensation will be less of an issue in the independent asymmetric history condition. To investigate whether the unbiased fair solution involves compensating the loser for his/her prior loss, we employed a preliminary survey. Here we elicited judgments from a sample of subjects drawn from the same population as those recruited for the actual experiment. We described the experiment to them and asked them to take the position of a neutral judge and to provide their views about what would be a fair division of hypothetical joint proceeds that could arise in the first stage of the experiment. These judgments allow us to examine the magnitude of the self-serving bias by the parties in the experiment (winners and losers) and to obtain impartial judgments about whether and how history should play a role in the situation created by the experiment.","In stage one of the bargaining experiment, we manipulated pairs’ histories by means of a random allocation of proceeds generated in a real-effort task. In the shared history condition, one single coin flip for each pair of subjects determined which of the two individuals received remuneration for the task. Negotiators’ outcomes were therefore interdependent, as only one party (the winner) received his/her stage one earnings, while the other (the loser) did not get paid for completing the task. As a result of this stage one manipulation regime, we solely obtained shared asymmetric (winner-loser) history pairs in the shared asymmetric history condition. In the independent history condition, one coin was flipped for each player, which determined whether the player received his/her stage one earnings (and became a winner or loser). Depending on the outcomes of these coin flips, we realized the following three different types of pairs in the independent history condition: both players are winners, both players are losers, or one player is a winner and one other the loser. Only the latter group, consisting of a winner and a loser, is comparable in terms of material outcomes to pairs in the shared asymmetric history condition. The two independent symmetric history conditions, consisting of two winners or two losers, are uninformative for testing our hypothesis and are not included in the section on testing predictions, although we do report bargaining outcomes in those conditions. In a similar vein, although we also elicited fair divisions of joint proceeds from judges for both types of independent symmetric history pairs (two winners or two losers), we do not investigate these divisions in the hypothesis testing section. In stage two, players completed a ten-item knowledge quiz remunerated in a piece-rate fashion. For a similar approach see, e.g., Ball et al. (2001), Clark (1998), Gächter and Riedl (2006), and Hoffman et al. (1994). The earnings of both players were pooled, and they then negotiated over how to divide the pooled amount. The purpose of the real-effort tasks in both stages of the experiment was to create a sense of entitlement on the part of the subjects (Birkeland, 2013; Cherry et al., 2002). This had particular significance in stage two, because we expected that knowing one's own and one's bargaining partner's individual contributions to the joint proceeds would reduce the tendency of equal-split claims and make splitting in proportion to each player's contribution a salient fair division (e.g., Gächter and Riedl (2006) and Ochs and Roth (1989)). This design feature also allows us to define compensation. By compensation we mean giving to the loser beyond his/her contribution to the stage two pooled earnings – i.e., compensating the loser at the winner's expense. After both subjects in a pair finished working on the knowledge quiz, they learned their own and their bargaining partner's individual earnings and the pooled earnings. Then, they bargained over how to split their pooled earnings. Before the negotiation began, players stated their beliefs about their own fair share of the earnings expressed in Hungarian Forints (i.e., HUF). Subjects were not asked what they personally viewed as a fair share, which we could not have elicited in an incentive-compatible fashion, but instead were asked to predict the judgments of the impartial (in the sense of not having a stake in the negotiation) judges whose judgments about a fair division had already been elicited via the aforementioned preliminary survey administered. Subjects in the bargaining experiment were then rewarded based on how close their predictions came to the actual judgments of the impartial judges. This procedure is similar to that used in prior research on self-serving assessments of fairness (e.g., Babcock et al. 1995, Gächter and Riedl 2005).","Our key behavioral prediction is that impasse will be more likely to occur in the shared than in the independent asymmetric history condition. We anticipated that negotiators in the shared asymmetric history condition would form divergent, settlement-hindering beliefs about the fair settlement, while these beliefs would be more similar between negotiators in the independent asymmetric history condition, which would facilitate settlement. More specifically, we predicted that losers in the shared asymmetric history condition would perceive a larger amount of compensation as fair than losers in the independent asymmetric history condition or winners in any condition. Relatedly, we predicted that losers in the shared asymmetric history condition would believe that they were entitled to compensation relative to the proportion-to-input settlement, while winners would believe that compensation was not called for. We also expected that impartial judges would perceive compensation as fair for losers in both asymmetric history conditions, but that this would be greater for losers in the shared than in the independent asymmetric history conditions. We had no specific predictions about how much players’ and judges’ views of fair compensation would overlap, nor about whether winners’ or losers’ perspectives would deviate more from those of the unbiased judges.","The two-stage bargaining experiment was programmed in oTree (Chen et al., 2016) and conducted in twenty-eight sessions. Each session lasted approximately fifteen minutes, and participants received a 300 HUF show-up fee.1 Experimental screenshots in the original language and their English translations can be found in Appendix A. Procedure ~~~~~~~~~ Subjects were recruited from Corvinus University of Budapest, Hungary. After arriving at the lab facilities at the university, they read and signed the informed consent form and were seated at one computer arranged so that subjects could not see others’ screens. Assistants welcomed the subjects and informed them that they could discontinue the experiment at any point, in which case they would only be paid their show-up fee. Subjects were also informed that if they had any questions, they should raise their hand and address them to the experimenter privately. Then, they clicked on a link which started the experiment. In the first stage of the experiment, paired subjects (henceforth, players) were randomly assigned to either the shared or the independent history pair-level condition. After answering basic demographic questions, they were given a real-effort task entailing labeling five simple images. For example, if an image of a spoon was presented, they had to type the word “spoon.” Players in both pair-level conditions were truthfully informed before they began the image-labeling task that, after they labeled all five images, one image would be randomly selected for each player in a pair. If both parties within a pair correctly typed in the name of this image, they both individually earn 1500 HUF. Pairs in which at least one party failed at successfully completing the image- labeling task stopped the experiment after this stage and were only paid a show-up fee. Players were also told that, upon successfully completing the image-labeling task, a random device would decide whether they would be paid their 1500 HUF remuneration. At this point, they were only told about these earnings and they were actually paid out at the end of the experiment (this procedure was truthfully described in detail to the players). To establish the shared history manipulation in the shared history condition, one coin was flipped for each pair. That is, players’ outcomes were interdependent since only one of them could in fact receive her/his 1500 HUF earnings for the image-labeling task. The player favored by the coin flip (i.e., winner) received his 1500 HUF earnings, whereas the player not favored by the coin flip (i.e., loser) did not receive his earnings of 1500 HUF for completing the image-labeling task, but instead received 0 HUF. To establish the independent history manipulation in the independent history conditions, one coin was flipped individually for each player in the pair. That is, two coins were flipped simultaneously for each pair creating independence between the players’ outcomes. The outcomes of these two coin flips individually and independently determined for each player whether s/he in fact receives the 1500 HUF earnings from the image-labeling task – i.e., was a winner or a loser. The outcomes of these two coin flips created three different independent history conditions. When one of the two coin flips favored one player (i.e., winner) but the other coin flip did not favor the other player (i.e., loser), as in the independent asymmetric history condition, the winner received his/her 1500 HUF earnings while the loser did not receive her/his 1500 HUF earnings. When the two coin flips favored both players, as in the independent symmetric history winner-winner condition, they both received their 1500 HUF earnings for completing the image-labeling task. When the two coin flips favored neither party, as in the independent symmetric history loser-loser condition, neither player received his/her 1500 HUF earnings for completing the image- labeling task, and instead, they each received 0 HUF. At the end of stage one, the coins were flipped and players learned the outcomes of the coin flips corresponding to their and their partner's earnings. That is, players in all conditions were not only informed about whether they received their 1500 HUF but also whether their partner received his/her 1500 HUF for completing the image-labeling task. In the second stage, player pairings were kept the same as in stage one. At the beginning of this stage, players were reminded about their own and their partner's stage one history – i.e., the partner's coin flip outcome and earnings – and then were given instructions for stage two. Next, they individually worked on the same ten-item trivia quiz for which each correct answer yielded 150 HUF for each of them. Once quizzes were completed, they learned their own score on the quiz (i.e., how many correct answers they gave), their individual earnings in HUF, their partner's quiz correct score and earnings in HUF and their joint quiz earnings in HUF. As a proxy for their beliefs about their unbiased fair share, players then provided their best guess (expressed in HUF) of the mean judgement of five impartial judges of the fair division of their pooled stage two trivia quiz proceeds (see Appendix A for the exact wording). This belief elicitation procedure was incentivized as follows. Players were told, truthfully, that five neutral judges, who had been informed about both partners’ stage one history and both parties’ individual contribution to the joint stage two quiz earnings, had judged fair splits of their joint proceeds. Each player who answered within 10% of the mean of these judgments received an extra 300 HUF. To avoid wealth effects, however, they only learned whether they had earned this bonus at the end of the experiment. The beliefs were then compared to results from a preliminary judge survey. In this preliminary judge survey, we elicited five fair divisions for each one of every possible combination of pooled trivia earnings between players in all four pair-level conditions. Each combination was judged by five different judges, and these judgments were averaged, creating a large table of means that was used in the main experiment. Appendix A provides a detailed description of this preliminary judge survey in which fully informed neutral judges proposed fair splits between players. After participants stated their beliefs about their fair share (their prediction of the judges), they entered the bargaining phase, in which they were given three rounds to agree on how to divide their joint quiz proceeds. Negotiators simultaneously submitted their offers (i.e., how much they would like to take for themselves) and then learned how much their bargaining partner claimed. If offers summed up to the joint proceeds, then they agreed and received as much as they claimed. If their offers summed to less than the joint proceeds, they received their claims and the leftover funds were equally split between them. If offers summed to more than joint proceeds, then they entered the next round. If they failed to agree in the third round, the amount to be divided shrank by 20% and was randomly (with every division equally likely) distributed between parties. Once players were done with the bargaining (i.e., they either settled or reached impasse), they learned how much they had earned in the bargaining phase. Next, they responded to a seven-item survey presented in two clusters. In the first cluster, which was two items, one item asked them how much they agree with the statement that it is fair to compensate from the joint quiz earnings the party who did not get paid for the image-labeling task. The other item asked them how much they agree with the statement that how much someone earned from the image-labeling task should have a bearing on the fair division of the joint quiz earnings. In the second cluster, which was five items, players answered questions about their satisfaction with and feelings about their bargaining outcomes. Finally, they learned how much they earned in the experiment in total and were paid. Sample ~~~~~~ Subjects in the bargaining experiment and respondents to the judge survey were recruited from Corvinus University of Budapest, Hungary, spanning a wide variety of study fields such as marketing, sociology, international relations, economics, applied economics, and business administration. Those who participated in the preliminary judge survey were not allowed to participate in the bargaining experiment. There were no other exclusion criteria for participation. Five-hundred twenty-two subjects (261 pairs) participated in the bargaining experiment, but eight players failed to pass the image-labeling task (by chance, all in the independent symmetric history condition). The final sample consisted of 514 players (257 pairs), 91 pairs in the shared asymmetric history, 75 pairs in the independent asymmetric history, 45 pairs in the independent symmetric history loser-loser, and 46 pairs in the independent symmetric history winner-winner condition. Additionally, a total of 109 participants completed the preliminary judge survey.2 Specifically, 23 individuals made judgments about earning combinations occurring in the shared asymmetric history condition, 26 about the independent asymmetric history conditions, 32 about the independent symmetric history loser-loser, and 28 about the independent symmetric history winner–winner conditions. Comparing the demographics of the shared and independent history conditions, the only difference between the groups was that winners in the shared asymmetric history condition are younger (Mean (SD) = 20.75 (1.51)) than winners in the independent history treatment (Mean (SD) = 22.17 (4.59)), W(1, 87) = 6.66, p ≤ 0.05. Comparing survey participants’ (judges’) and players’ demographic characteristics, the mean age in years is significantly higher among judges (Mean (SD) = 23.07 (4.25)) than players (Mean (SD) = 21.39 (2.84)), F(1, 617) = 25.66, p ≤ 0.001. In other aspects, there are no differences between players’ and judges’ demographic characteristics. Detailed demographics are presented in Appendix Table B1 for players and in Tables B2 and B3 for judges.","First, we present descriptive results from all four pair-level treatments. Then, we continue with testing predictions on the restricted sample of shared asymmetric history and independent asymmetric history treatments. Descriptive results ~~~~~~~~~~~~~~~~~~~ The upper panel of Table 1 presents descriptive results of players in the bargaining experiment. From the first row one can see, consistent with the random assignment to winning or losing in stage one and to one of the history conditions, that there is no significant difference across individual-level conditions in players’ contributions to the stage two pooled quiz earnings. Players in a pair contributed, on average, 50% of the to- be-divided joint earnings, suggesting that the stage one manipulation did not impact their stage two effort levels. The marginal (i.e., overall, across treatments) mean (SD) is 1044.75 (223.54) HUF, which corresponds to a mean (SD) of 6.96 (1.49) correctly answered quiz questions. Note, that there is also no difference in the stage two pooled earnings between the four pair-level treatments, F(3, 253) = 0.29, ns; the mean (SD) for all players is 2089.49 (323.77) HUF. The mean experimental earnings in HUF (excluding the show-up fee) differs across individual-level treatments, mostly due to the stage one loser-winner manipulation. In general, winners earned more than losers in the experiment. Players’ beliefs about their fair share expressed proportional to their stage two quiz earnings (i.e., how much the player stated that s/he should get in HUF according to impartial judges divided by her/his individual stage two quiz earnings in HUF) are presented in the second row of Table 1. Results show that losers in both asymmetric history conditions believe they are entitled to a greater share than their winner partners believe they are entitled to. Losers in the shared asymmetric history condition believe they are entitled to a greater share than losers in the independent asymmetric history condition, F(1, 164) = 21.61, p ≤ 0.001. Losers in both asymmetric history conditions believe that they are entitled to more than their contributions (their 95% CIs are above 1), whereas winners, on average, believe that they are entitled to no more than their contribution (95% CIs include 1). The lower panel of Table 1 presents judges’ views about players’ fair share expressed as a proportion of stage two joint earnings. Two findings emerge here. First, in both asymmetric history conditions judges think it is fair to give a larger share to losers than to winners. Second, they believe that losers in both asymmetric history conditions are entitled to more than their contributions in the second stage, as the 95% CIs are above 1. This amount is, however, larger for losers in the shared than in the independent asymmetric history treatment, W(1, 24) = 49.83, p ≤ 0.001. Although judges’ views of fairness in the two symmetric conditions were mainly solicited to determine fair shares for subjects in these conditions, these judgments do address the question of whether judges believed that, absent asymmetries, players should be paid in proportion to their earnings. The answer is that they do hold this view: on average, judges believe that losers and winners in the two symmetric history treatments are entitled to their contributions. When comparing the 95% CIs of the means of players’ beliefs (second row of first panel) and judges’ views (first row of second panel) about players’ fair shares, we observe that these views overlap between players and judges in every individual cell except for winners in the shared asymmetric history treatments. These winners believe that fairness does not dictate granting losers beyond their contribution to the stage two joint earnings; but judges do believe that such additional compensation was fair. Further details about stage two joint quiz earnings in HUF, players’ stage two quiz and experimental earnings in HUF, and judges’ views in HUF are presented in Table B4 in Appendix B. The main results of the experiment are presented in Fig. 1. Looking first at settlement versus non-settlement (right-hand bar in each figure), the distribution of impasse differs between the four pair-level treatments, χ2(3, Npair = 257) = 14.68, p ≤ 0.01 in the predicted fashion. Specifically, impasse is more likely in the shared asymmetric history treatment (22%) than in the independent asymmetric history treatment, where it was only 7%, χ2(1, Npair = 166) = 7.53, p ≤ 0.01. Comparison of the independent asymmetric history treatment to those in the two symmetric history treatments shows that impasse rates are similar across these three treatments, χ2(2, Npair = 166) = 0.32, ns. However, the pattern of settlement does differ between these three treatments. Specifically, pairs seemed to settle more quickly in the two independent symmetric history treatments than in the independent asymmetric condition. Finally, we find no difference in the likelihood of impasse within the two independent symmetric treatments, χ2(1, Npair = 91) = 0.24, ns. Tests of predictions involving fairness perceptions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To specifically test our predictions about the differential impact of shared and independent histories on asymmetric history pairs, we restrict the sample to the shared and independent asymmetric history treatments (Nindividual = 332). To simplify comprehension and analysis, we express players’ beliefs and judges’ views about the fair solution in terms of loser's share, which facilitates interpreting the results from the perspective of compensation given to loser. The loser's share is defined as the difference between the stated beliefs (players’ or judges’) about the amount the loser should receive and the amount contributed by the loser, scaled by the joint stage two quiz earnings. Loser's share = (amount in HUF loser should get – amount in HUF loser contributed)/pooled quiz earnings in HUF This form is commensurable within and between pairs and between players and judges. Additionally, it highlights whether the loser's share that is perceived as fair is more than his/her contribution to the pooled stage two quiz earnings, indicating the presence of compensation for the loser at the winner's expense. Using this formula, we first calculate, and present in Fig. 2, the mean players’ belief about the loser's fair share in each individual-level condition. Three important findings emerge from the 95% CIs of the means. First, winners in the shared and independent asymmetric history conditions believe that splitting according to contribution is the fair settlement, as the 95% CIs of mean beliefs include zero. Second, losers in both pair-level conditions believe that the fair settlement includes giving more to losers than their contribution, as the 95% CIs include only positive numbers. In other words, losers in both conditions believe that their fair share includes compensation from their winner partners. Third, losers in the shared asymmetric history condition believe that a greater share is fair than losers in the independent condition or winners in any condition. The statistical significance of the third finding is confirmed by the history-by-loser interaction in Model I of Table 2, which regresses players’ beliefs about the loser’s fair share on history (shared asymmetric = 1 vs. independent asymmetric = 0) and role (loser = 1 vs. winner = 0). The psychology suggested by this pattern receives further support from responses to the second item from the short survey administered at the end of the experiment: “How much someone earned from the image-labeling task has a bearing on what is the fair division of the joint earnings.” Losers (Mean (SD) = 2.92 (0.92)) in the shared asymmetric history condition agree more than their winner partners (Mean (SD) = 2.56 (0.90)) with this statement, F(1, 180) = 7.23, p ≤ 0.01. They also agree more than losers in the independent asymmetric history condition (Mean (SD) = 2.55 (0.79)), F(1, 164) = 7.76, p ≤ 0.01. At the same time, there is no loser and winner difference within the independent history condition. These responses should be treated with caution, however, since they were not incentivized and could just be justifications of behavior displayed in the experiment rather than true motives. Next, we calculate the difference between negotiators’ beliefs about the loser's fair share (i.e., loser's–winner's beliefs about loser's fair share) and plot the distributions in the two pair-level asymmetric history conditions in Fig. 3. Following Gächter and Riedl (2005), we refer to this difference as ‘tension’ between the individuals. As evident from the figure, the distribution of tension differs between the two conditions, Kolmogorov–Smirnov Z = 2.55, p ≤ 0.001. Most notably, the bulk of the tension is centered around zero in the independent asymmetric history condition, while it is concentrated in the positive domain in the shared asymmetric history condition. Mean tension is greater among pairs in the shared asymmetric history (M (SD) = 0.16 (0.16), 95% CI [0.13, 0.19]) than in the independent symmetric history condition (M (SD) = 0.06 (0.17), 95% CI [0.02, 0.10]), F(1, 164) = 14.33, p ≤ 0.001. In fact, as one can see from the 95% CIs, tension is greater than zero in both pair-level treatments, indicating that players in both treatments hold incompatible beliefs about the loser's fair share, though beliefs are more discordant between players in the shared than in the independent asymmetric history condition. We perform a series of binary logistic regressions to test the effect of shared history on impasse (Model I), and then we add tension (Model II), see Table 3 for a summary of results. From Model I, we learn that impasse is approximately four times more likely in the shared than in the independent asymmetric history condition. Adding tension, in Model II, we find that an increase in tension from its mean value (0.16) to its maximum (0.60) in the shared asymmetric history condition results in a change in expected probability of impasse from 20.5% to 56.3%. The equivalent increase from the mean (0.06) to the maximum (0.58) in observed tension in the independent asymmetric history condition also results in an increase in the expected probability of impasse, but a smaller one, from 5.3% to 32.6%. In Table B5 in Appendix B we also show that the effects of history and tension are robust after controlling for the difference between players’ age and effort levels (i.e., stage two quiz earnings) within a pair and the gender composition of a pair. One may wonder if there is a difference between settled and not-settled pairs in the shared asymmetric history condition with respect to their stage two quiz effort levels. Restricting the sample to this condition (Npair = 91), we find no difference between non-settled and settled pairs in this respect. The mean (SD) and 95% CI of quiz correct score signed difference for settled-pairs are −0.03 (0.17) [−0.07, 0.01] and for non-settled pairs are 0.03 (0.19) [−0.05, 0.13], F(1, 89) = 1.89, ns. Similarly, the distribution of quiz correct score signed difference does not differ between settled and non-settled pairs in the shared history condition, Kolmogorov–Smirnov Z = 0.83, ns. Results of a causal mediation analysis using the method of Imai et al. (2010a, 2010b) help to reveal the mechanism via which shared history affects impasse. On average, shared asymmetric history increases the probability of impasse by 0.15, and we find that 28% of that total effect is due to history's positive effects on tension. In other words, at least part of the reason that the shared asymmetric history increases impasse is because it increases the divergence between players’ beliefs about the loser's fair share.3 Next, we examine judges’ views about the loser's fair share. Recall that judges' views are in reference to a pair and are from the preliminary survey of those loser–winner quiz earnings combinations that occurred in the two asymmetric history conditions of the bargaining experiment (23 in the shared and 26 in the independent asymmetric history condition). The distribution of judges’ views in the two pair-level treatments is plotted on Fig. 4, which reveals a difference between the two distributions, Kolmogorov–Smirnov Z = 3.02, p ≤ 0.001. Perhaps the most striking result is that, while in 80% of the independent asymmetric history cases judges would not compensate losers beyond their contribution to second stage joint earnings, in most of the cases in the shared asymmetric history treatment judges would grant shares to losers that surpass their contribution to the second stage task. In fact, on average, judges would grant 21% more to losers in the shared than in the independent history condition. The mean (SD) and 95% CI of loser's share is 0.23 (0.10) and [0.19, 0.27] in the shared and 0.02 (0.03) and [0.01, 0.04] in the independent history conditions, W(1, 26) = 95.24, p ≤ 0.001. Nonetheless, from these 95% CIs we also see that judges do grant compensation to losers in both conditions. Finally, we determine the difference between players’ and judges’ beliefs about the loser's fair share (i.e., each player's stated beliefs minus the mean judges' view for the corresponding joint earnings case) and regress these differences on history (shared asymmetric = 1 vs. independent asymmetric = 0) and role (loser = 1 vs. winner = 0). As can be seen from the significant history by role interaction in Table 4, which is evident also in Fig. 5, in absolute terms, winners’ beliefs in the shared asymmetric history condition deviate the most from judges’ views in a way that benefits them (i.e., giving beyond their contribution to losers is unnecessary). Losers in this condition, however, believe they should be entitled to slightly less than the amount that judges believe is fair. The pattern is almost flipped for players in the independent asymmetric history condition. On average, losers here believe that they are entitled to more than judges would give them, while their winner partners formed concordant beliefs with the judge view.","We introduced two types of prior wealth asymmetry between bargainers and demonstrated that when the asymmetry is due to negotiators’ shared history, it is more likely to cause impasse than when it is due to negotiators’ independent history. This result is consistent with key findings of Camerer and Loewenstein (1993) and Dezső et al. (2015) about the pernicious potential of asymmetric history in negotiations. Though the conclusions of these papers are similar to ours, in the former paper asymmetry was not experimentally manipulated, whereas in the latter paper, asymmetry and interdependent history were not disentangled. To remedy these issues, we experimentally manipulated asymmetry and isolated the effect of asymmetry from interdependent history. We showed that it is not the asymmetry per se but the allocation schema by which it occurs that leads to elevated rates of impasse. The finding that sharedness and not asymmetry is the key issue between negotiators is also supported by the lack of difference between the likelihood of settlement between independent asymmetric and symmetric treatments. We argue, and demonstrated empirically, that these stalemates are due, in large part, to negotiators’ divergent and self-serving beliefs about the fair settlement. Losers in both asymmetric history conditions believed that they deserve more than their contribution from the joint proceeds, indicating that they believed that the fair solution prescribes compensating them for their prior loss. Nonetheless, losers with shared histories believed that they were entitled to a greater share than losers with independent history. On the other hand, winners with shared asymmetric history did not believe that the loser should get more than his/her contribution, which caused a significant divergence between losers’ and winners’ beliefs about the fair settlement in the shared asymmetric history treatment. By contrast, divergence between losers’ and winners’ beliefs about the loser's share were much smaller in the independent asymmetric history treatments, even though winners here also did not believe that giving to the loser beyond his/her contribution was necessary from the vantage point of fairness. All in all, players in all conditions self-servingly selected the fair solution that was the most beneficial for them, which led to a discordance between them on the issue of compensating the loser. The divergence of these beliefs was, however, smaller for pairs in the independent asymmetric history condition and hence, hindered settlement to a lesser extent. This key role of beliefs confirms the speculations of Camerer and Loewenstein (1993) and Dezső et al. (2015) about the self-serving potential of asymmetric history, though beliefs about the fair settlement were not elicited in either of these papers. The finding that negotiators’ relatively extreme perceptions of a fair settlement drove impasse provides empirical support for the proposition of Birkeland and Tungodden (2014) regarding principled disagreement. When bargainers arrive at the negotiation with highly incompatible beliefs, they insist on them even at the cost of reaching impasse. The results of the judge survey demonstrate that, from the vantage point of a neutral judge, past outcomes do have implications for fairness in the present. Redistribution is seen as fair when it involved parties sharing an asymmetric history, but less so when they had independent histories. In other words, a winner who benefited at the expense of a loser should compensate the loser with a greater amount than when the winner did not benefit at the loser's expense, as in the independent history condition. All in all, unbiased judges believe that fairness dictates compensation for those who received the short end of the stick in a prior division. When we compared players’ beliefs about the losers’ fair share to those of unbiased judges’, an interesting pattern emerged. On the one hand, only independent history winners’ beliefs overlapped with the impartial fair solution. Losers in this condition believed they were entitled to more than the judge's judgments prescribed. On the other hand, shared asymmetric history winners’ views of fair compensation to losers fell below the judges’ unbiased judgments of the losers’ fair share, and, surprisingly, the losers also believed they were entitled to less than the amount deemed fair by the judges. An alternative interpretation ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A potential alternative interpretation of our key finding is that the shared asymmetric history may have ignited a motivation to compete in losers when the playing field was finally leveled. This interpretation does not seem incompatible with the one we propose; it is an alternative, related mechanism that could produce a desire for compensation on the part of the loser, with no commensurate willingness to provide such compensation on the part of the winner. Note, however, that this interpretation would not predict the difference in views on fairness that the self-serving interpretation of fairness account predicts, and that we in fact found.","We believe that these and prior related results help to shed light on the prevalence of conflict between parties with a shared, and almost inevitably asymmetric history, such as business partners seeking a dissolution of the partnership, divorcing spouses, and competing countries. Our experiment suggests that the effects of self-serving bias in these situations will be worst when one party's history is less favorable than the other's, and, more specifically, when one of those parties gained at the other party's expense. Research on climate change negotiations reveals similar tendencies in countries’ argumentations on the burden of emissions’ reductions. Due to the complexity of this issue, multiple fair solutions arise, each corresponding to different sharing rules and different costs being imposed on countries (Kriss et al., 2011; Ringius et al., 2002). One element of this complexity, which is self-servingly invoked, is the interpretation of the significance of historical emissions for the fair sharing of current costs (Cazorla and Toman, 2000). Decades of unrestricted greenhouse gas emissions strongly contributed to the economic prosperity of many developed countries. Current developing countries would now like their turn at industrialization, and restricting emissions seriously constrains their economic growth prospects. Due to their asymmetric histories of pollution and economic growth, rich and poor countries self-servingly invoke or ignore emission history. Poorer countries argue that richer countries must take a lead in accepting greater abatement costs because they are more responsible for the current high level of greenhouse gases and they have already benefited from polluting activities (Lange et al., 2010). By contrast, richer countries claim that the fair solution is “cleaning the slate” and splitting the burdens, independent of emission history. Given the subtlety of the shared/independent manipulation, it is remarkable that it had such a strong effect. After all, the two people were paired together, and could have viewed the gain of one and loss of the other as relevant. In the real world, such asymmetric independent histories might have a smaller impact. A potential threat, however, is that leaders will use the effects revealed in the experiment to manipulate constituencies. They may attempt to instill in people who are in inferior wealth positions the view that their misfortune stems from a shared asymmetric history with other groups, such as immigrants. Creating such a narrative, the current research suggests, may instill and invoke self-serving interpretations among those feeling they are in a disadvantaged position. The unfortunate consequence is, as demonstrated in our experiment, a further discord between these people and, perhaps, a breeding ground for a lingering desire to even the score with those who previously benefited. There is, however, a small silver lining to our research. An important property of self-serving perspectives on fairness is that people generally believe that impartial judges will share their own biased views. Since both parties genuinely believe that their perspective on fairness will be shared, they should both be open to arbitration. In conclusion, a shared, asymmetric history can lead to discord between parties, due to individuals’ self-serving interpretations of the implication of history. Whether one adopts the view “let bygones be bygones,” or sees the past as relevant to the future, depends on how one fared in the past, and on the degree to which the other party was responsible. Disputes between overworked employees are examples of relatively innocuous consequences of self-serving invocations of history, but these are also evident in high-stakes situations, such as disputes on fulfilling refugee quotas in Europe or bearing the costs of reducing greenhouse gas emissions all over the world. Divergent, self-serving invocations of history hamper agreement on how to share sacrifices for the greater good and may also further undermine international cooperation on issues from which we would all gain from settling. Identifying the motivations ignited after being wronged is an important link to understanding how interacting parties’ shared and asymmetric history could spill over into subsequent disputes. These conflicts may range from the mundane to crucial international debates, including the creation and maintenance of well-balanced power structures and, most importantly, saving our planet."],["We measure people's pro-social behavior, in terms of voluntary money and labor contributions to an archetypical public good, a bridge, and in terms of voluntary money contributions in a public good game, using the same non-student sample in rural Vietnam at four different points in time from 2005 to 2011. Two of the observed events are actual voluntary contributions (one in terms of money and one in terms of labor), one is from a natural field experiment, and one is from an artefactual field experiment. Despite large contextual variations, we find a strong positive and statistically significant correlation between voluntary contributions, whether correcting for other covariates or not. This suggests that pro-social preferences are fairly stable over long periods of time and contexts. © 2014. --------------------------------------------------------------------------------","The present paper investigates the stability of social preferences by utilizing unique data on people's voluntary contributions to an archetypical public good, a bridge, and contributions in a public good game. The analysis is based on two observations of voluntary contributions, one natural field experiment on contributions to a real bridge in rural Vietnam, and one artefactual field experiment on contributions in a public good experiment, conducted from 2005 to 2011 and using the same sample consisting of all (about 200) households in a village in rural Vietnam. Thus, we obtain repeated information on people's pro-social preferences over a long period of time. An overwhelming amount of psychological and behavioral economics research shows that the Homo Economicus characterization of human behavior, in terms of complete selfishness in a narrow material sense, is often importantly wrong; human behavior is in part pro-social. At the same time, a large heterogeneity in pro-social behavior is typically found. Several studies have consequently attempted to categorize people, based on their experimentally observed behavior, in terms of different types of social preferences, e.g., as free riders, conditional cooperators, and unconditional cooperators (Fischbacher et al., 2001); as selfish versus inequity-averse individuals (Fehr and Schmidt, 1999); and as non-sharers, reluctant sharers, and willing sharers (Lazear et al., 2012). Yet, from these studies one cannot conclude that people are inherently of different types. An alternative explanation is that people simply act differently at different points in time, and that people's degrees of cooperativeness, or non-selfishness, are approximately constant on average. Indeed, that people's pro-social actions vary over time is obvious since most of us sometimes contribute to a certain charity and sometimes not. Yet, how much of the observed heterogeneity in social preferences that can be explained by within-people variations is not clear, nor is it clear whether it is significantly more likely that an individual who acted cooperatively at one point in time is more likely to act cooperatively in a similar task several years later. Moreover, even if people are of different types with respect to pro-social preferences, it is an important research issue to find out whether these types are stable over longer periods of time. The present paper is, to our knowledge, the first in economics to systematically investigate whether pro-social preferences, manifested in terms of cooperative behavior, are fairly stable over several years. In contrast, the extent to which preferences, and in particular social preferences, are stable over a short period of time, and also across decision environments, has been studied in a number of papers with different methods. Some studies have looked at the differences in pro-social behavior between similar experiments conducted at different points in time. For example, Brosig et al. (2007) conducted dictator and public good games with the same subjects at several points in time over the course of one week. Other-regarding behavior was found to decrease over time, and in the final experiments the subjects' behavior was close to that predicted by conventional economic theory. Subjects who behaved selfishly were found to be the only ones who behaved stably over time. This pattern is similar to the one typically obtained with repeated public good games.3 Volk et al. (2012) had subjects participate in a series of three identically designed public good experiments over a period of five months. At the aggregate level, cooperation preferences were stable. However, at the individual level there was considerable variation. Classifying subjects into three categories – conditional cooperators, free riders, and others – it was found that half of the subjects remained in the same category in all three experiments. Other studies have looked at differences in pro-social behavior between different experiments (with the same subjects). Blanco et al. (2011) ran four different experimental games: dictator, ultimatum, sequential-move prisoners' dilemma, and public good games, and tested whether the Fehr and Schmidt (1999) inequity aversion model can explain the results. They found that the model could explain the results reasonably well at the aggregate level, but that it performs considerably less well at the individual level. De Oliveira et al. (2012) found that preferences for contributing to public goods are positively related across different experimental decision contexts, and also positively related to self-reported donations and volunteering outside the laboratory. Yet other studies have compared contributions in the lab and in the field. Benz and Meier (2008) conducted a dictator game with two social funds as external recipients and found a positive, albeit relatively weak, correlation between subjects' behavior in a lab experiment and actual charitable giving. Laury and Taylor (2008) found mixed evidence regarding the correlation between non-selfish behavior in laboratory experiments and contribution to a charitable organization. While they found that some measures of altruistic behavior in the lab could be predictive of contributions to the charity, the relationships were generally weak, and some measures of altruism were even negatively correlated with contribution to the charity. Based on a trust game in Peru, Karlan (2005) found that subjects identified as trustworthy, i.e., receivers who returned a relatively large share of what they received from the senders, tend to repay their microcredit loans to a larger extent than those who were not identified as trustworthy in the experiment. No significant correlation between those identified as trusting, i.e., senders who sent a relatively large share to the receivers, and repayment of the loans was obtained. Fehr and Leibbrandt (2011) found that fishermen in Brazil who are more cooperative and patient in lab experiments are also less likely to exploit the common pool resource in the sense that they use shrimp traps with bigger holes (where small shrimp can escape) and fishnets with larger mesh sizes (where only bigger fish are caught). Algan et al. (2013) combine an online experiment with field contribution data for 850 Wikipedia contributors. They found that actual contributions to Wikipedia are strongly related to levels of reciprocity as revealed by both a conditional public good game and trust game. However, they found only weak links with altruism as revealed by dictator game contributions. Cesarini et al. (2009) used a different approach based on twin studies combined with modified dictator experiments to determine the extent to which giving is heritable. Their best point estimate suggests that genes explain about 20% of the variation in behavior among subjects and hence that social preferences, as manifested in giving behavior in dictator experiments, are in part explained genetically. Yet, this is not necessarily a good measure of the degree to which social preferences are constant over time. First, a certain genetic set-up may in principle induce variation in behavior over time. Second, there are many environmental factors that may work in the direction of stabilizing social preferences, e.g., the development of close relations and social norms.4 In summary, most of the studies discussed above point in the direction that social preferences are partly stable over different domains and over time, although the extent of the stability varies between studies to a rather large extent. The main contribution of the present paper is that it investigates the stability of social preferences over long periods of time, where also the contexts differ; we observe behavior in four different periods: in 2005, 2009, 2010, and 2011. As a result of the relatively long time between our events, we argue that we can also more or less rule out the potentially confounding effects of compensatory behavior due to moral licensing and moral cleansing.5 The rest of the paper is organized as follows. Section 2 describes the four observations of voluntary contribution and the design of the two experiments, and provides the corresponding background statistics. Section 3 presents the results. We find strong positive and statistically significant correlations between voluntary contributions in these events, whether correcting for other covariates or not, suggesting that pro-social preferences seem to be quite stable over long periods of time.6 Section 4 concludes the paper.","We use observations of subjects' pro-social behavior in four related events in 2005, 2009, 2010, and 2011, i.e., implying rather long time periods between events and a total study period of about six years. Two of the events (in 2005 and 2010) are naturally occurring ones where we simply observed the behavior, and the other two (in 2009 and 2011) were designed by the authors. Two of the events (in 2005 and 2009) concern monetary contributions to a local public good in terms of the construction of a crucial bridge in the middle of the village, one of the events (in 2010) concerns labor contributions to construction of the same bridge, while the last experiment (in 2011) was a public good experiment not related to the bridge at all. The first three observations, i.e., in 2005, 2009, and 2010, deal with a real event related to the construction of a bridge. The event in 2009 can moreover be classified as a natural field experiment, whereas the last experiment is an artefactual field experiment, i.e., a lab experiment conducted in the field with a non-student sample. All four events focus on voluntary contribution mechanisms, and although there are a number of contextual differences, in each instance we observe the behavior of the same (approximately 200) subjects, representing all households in the village. The events were undertaken in the Giong Trom hamlet,7 in the Mekong river delta of Vietnam, where about 200 households live and use the bridge (if/when it is in sufficiently good shape). Most households in the hamlet are engaged in rice cultivating activities.8 The hamlet suffers from a problem that is common in the Mekong river delta: lack of basic infrastructures such as rural roads, bridges, and irrigation canals. The government only provides larger public goods such as roads between villages, whereas small-scale infrastructures within a hamlet are considered to be the responsibility of the hamlet. The bridge and the voluntary contributions ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The bridge is important for the village because villagers use it to go to the rice fields, to the market, to school, and to visit friends, given that the bridge is in sufficiently good condition. If they do not use the bridge, they have to choose one of two other routes, each located parallel to and about 1200 m from the bridge's pathway; see the map in Fig. 1. There are people living on both sides of the canal, and there is also important infrastructure on both sides of the canal. For example, the main market is located on the north side of the canal, while the school is located on the south side. For all but a few households, the travel distance to reach the destinations on the other side is considerably longer if they cannot use the bridge. The first three observations of pro- social behavior, to be described next, concern funding of a bridge for the hamlet. The 2005 event ~~~~~~~~~~~~~~ In the first event in 2005, the hamlet council had decided to try to build a bridge, which was to be funded by voluntary contributions. A group of three delegated individuals visited every household in the hamlet to present the plan to build the bridge and to ask for voluntary contributions. Probably in order to persuade villagers and increase contributions, the delegates showed a list of names, contribution amounts, and signatures of those who had already contributed. The hamlet council did not set an upper contribution limit, and the highest contributed amount was 300,000 dong.9 Since the total contribution was not sufficient to build a concrete bridge, the hamlet council decided to build a wooden bridge. Yet, the bridge became degraded relatively quickly, and in 2009 it looked like in the picture in Fig. 2 and could obviously not be used for tractors and motorbikes. The full list of contribution amounts from each household was kept by the hamlet council and made available to the authors of this paper. The 2009 experiment ~~~~~~~~~~~~~~~~~~~ In collaboration with an NGO, we conducted a field experiment using a threshold public good game that concerned the funding of a new bridge for the hamlet in 2009. Fifteen experimenters were recruited via advertisements and a careful selection procedure at the University of Economics Ho Chi Minh City. The selected experimenters received extensive training and spent nearly one week practicing the experiment in role-play pairs and for pilot interviews with farmers. Moreover, the selected experimenters also had experience from similar fieldwork in a neighboring rural area. A list of likely questions and answers related to the project was provided and the experimenters were repeatedly told about the importance of using the exact wordings prescribed in the script. For a detailed description of the experiment and the results, see Carlsson et al. (2013). The main objective of the experiment was to investigate the role of social influence for voluntary contributions to public goods. We devised a threshold public good game, in which each of the 200 households received a 400,000 dong endowment from the NGO and had the option to either keep the money or contribute some or everything to the funding of the bridge. The threshold level was set at 40 million dong, meaning that if all villagers together would contribute at least 40 million dong, the bridge would be built; otherwise it would not. The threshold was explained as follows: The concrete bridge will be established if all families together contribute at least 40 million dong. This means that if the total contribution is equal to or above 40 million dong, the project will use these 40 million dong, add more funding in order to meet the costs of the bridge, and take the responsibility to build the bridge. If the total amount of money collected exceeds 40 million dong, the excess amount will be returned to your family according to the proportion you contributed. If the families are unable to contribute a total of 40 million dong, your contribution will be returned to you, and the concrete bridge will not be built. The experiment involved five randomly distributed treatments as follows: 1. Baseline case with no reference contribution level and no default option. 2. High reference contribution level (300,000 dong) and no default option. 3. Low reference contribution level (100,000 dong) and no default option. 4. No reference contribution level and a default option at zero contribution, and 5. No reference contribution level and a default option at full contribution (400,000 dong) of the endowment. The reference contribution levels were conducted by providing the subjects with information about a typical previous contribution of others. These numbers, in turn, were obtained from an initial treatment where we did not tell the subjects anything about others' contributions. The default options were conducted using a metal card with 9 different contribution levels. A magnetic token was initially put at the 0 dong level or at the 400,000 dong level. The subjects were then asked to move the token to the amount that they wanted to contribute to the public good. In all treatments, the contributions were anonymous to everybody except the solicitors, i.e., the contributions were not revealed to any parties. Since the households contributed enough to reach the threshold, the new bridge was built in early 2010; see the picture in Fig. 3. The 2010 event ~~~~~~~~~~~~~~ During the planning of the construction, we had a meeting with the head of the hamlet and representatives from the farmers' association. At the meeting, we were informed that they planned to ask the villagers to contribute labor to connect the road with the new bridge. We took this opportunity to collect another naturally occurring contribution data set. As the construction work required joint efforts in a short period of time, three and a half day were scheduled for the joint work. Two persons from the hamlet council visited the households in the hamlet to invite villagers to contribute labor. Hence, an important difference compared with the previous two events is that instead of being asked for monetary contributions, they were asked for labor contributions. Since some households were not expected to be able to contribute any labor at all, mainly because the household members were too old, not all households were asked to make contributions. In total 19% of the households were not asked to make any labor contributions.10 At this time, households were not told anything about what others were contributing, and there was no provision point. We hired an external supervisor who monitored the construction progress and quality, and recorded villagers' labor contributions. Thus, what we observe in this event is the actual amount of labor contributions, and not what they promised when asked to contribute. The 2011 experiment ~~~~~~~~~~~~~~~~~~~ While the bridge is presumably useful to all households, the usefulness varies with, e.g., the distance to the bridge and ownership of different vehicles. Although we have information on the use of the bridge, it is possible that we still cannot perfectly correct for it in our analysis. In order to avoid such potentially confounding effects, it makes sense to also include an experiment that by design is not at all related to the bridge. The experiment conducted in 2011 was therefore not in any way related to the bridge or the 2009 experiment, and no reference to either the bridge or the earlier experiment was made. While the 2009 experiment was funded by an NGO, the 2011 experiment was conducted and funded by the University of Economics Ho Chi Minh City. There were 15 experimenters also in this experiment. They were recruited and trained in a similar manner as for the 2009 experiment, but none of them were the same as in the 2009 experiment. The experiment was conducted in the subjects' homes and the first part of the experimental script read: Hello, my name is […] I'm working for University of Economics Ho Chi Minh City, which is implementing an economic study in this village. I would like to invite you to participate in an economic experiment, which will be conducted now. You will be taking part in an experiment on decision-making and will earn some money by choosing among alternatives. The experiment is designed so that your earnings will depend on your choices, as well as on the choices of others. Your earnings will be paid to you in cash tomorrow afternoon privately to preserve the confidentiality of your earnings. The results from this study will only be used for academic purposes, and no other people in the village will have information about your choices. We chose to conduct a public good game, much like a standard linear public good game conducted in laboratory experiments. Yet, in order to fit the setting of the village, and in order to be able to easily compare the contribution behavior in this experiment with the other experiments, the group consisted of all households in the village; thus, the size of the group was approximately 200 subjects. Each of these households received 200,000 dong, which was clearly a substantial amount for them. Just as in a standard laboratory public good experiment, they had to decide how much of the endowment to keep and how much to put into a group account. In order to make the experiment simple to understand, each subject was told that any money that was put into the group account would be doubled by the experimenter, and that the total amount in the group account would be distributed to the group members, and hence to the households. The following sentences were part of the script that the experimenters read to the subjects. This experiment is organized with a group of 200 households including yours. We will give 200,000 dong to each of the 200 households participating in this experiment. Here is the agreement, which has a signature and stamp from the University of Economics Ho Chi Minh City stating that 200,000 dong belongs to your household. Your household and each of the other households have to decide how much of the 200,000 dong to keep and how much to put into a group account. The amount of money you put into a group account will first be doubled by the experimenter, meaning that if you transfer 1 dong into the group account, it will become 2 dong. The total amount will then be divided equally between all households. The instructions included several examples, and the subjects were allowed enough time to understand everything that was said. To inform the subjects that their contributions would be kept completely confidential, the experimenters told them before they made their contribution decision: “No one in the commune, not even the officials, will know about your decision. We will keep your contribution information completely secret.” Summary of the four events and household characteristics ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As mentioned, the first three events were related to the bridge in the hamlet. The first event concerned monetary contributions to build a small wooden bridge in 2005; the second concerned monetary contributions to build a new and better concrete bridge in 2009; and the third event concerned voluntary labor contributions to connect the road to the new concrete bridge in 2010. The fourth event was instead a public good experiment that was not related to the use of the bridge. The settings of the four events are summarized in Table 1. While we designed only two experiments (the 2009 and 2011 experiments), we have data for four different points in time, i.e., 2005, 2009, 2010, and 2011, for essentially the same subjects. Table 2 reports background statistics, as of 2009, for the households. There were a total of 200 households from the village that participated in the events. However, as explained above, not all households were asked to contribute labor in the 2010 event, and in the last experiment four of the households were unable to participate.11 The mean monthly household income is around 1.8 million dong per month. This amount corresponds to about 95 USD per month, which is less than one USD per household member and day. Thus, the households in the study are poor, and the average level of education is very low. The average size of land on which a family is currently cultivating rice is also rather small, approximately half a hectare. Average contributions in the four events ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Before looking at the correlations between the contributions, let us briefly look at Fig. 4, which displays the histograms of contribution in each of the events. As can be seen, the contribution patterns are strikingly different, in particular between the 2005 and 2010 events, on the one hand, and the 2009 and 2011 events on the other. The 2009 and 2011 events have in common that they are based on windfall endowments, which may contribute to the substantially higher contribution levels in these events.12 Comparing the 2005 and 2009 contributions, which were both in terms of monetary contributions to a new bridge, there are thus strikingly large differences. The mean contribution in 2009 was almost seven times as large as in 2005 (270,000 dong compared with 40,000 dong), and while almost everyone contributed something in 2009, almost half of the households chose to free ride in 2005. While there may be many different explanations for this, two clearly stand out: First, as mentioned, the contributions in 2009 were based on a windfall endowment provided by the NGO, while the contributions in 2005 were paid out of the households' existing wealth. Second, the 2009 experiment involved a matching contribution by the involved NGO. Such matching contributions, or seed money, have been shown to increase voluntary contributions substantially (e.g., List and Lucking-Reiley, 2002; Karlan and List, 2007). Moving on to the 2010 event, we can observe that even fewer chose to contribute than in 2005, in this case in terms of labor contributions. In 2010, the mean contribution of labor was 0.5 labor days per household, which corresponds to about 40,000 dong based on an average daily labor wage of 80,000 dong. Finally, in the 2011 experiment, there is again a small share of people contributing nothing, where the mean contribution is substantial, 125,000 dong. The mean contribution as a fraction of the maximum contribution is actually quite similar to the 2009 experiment; these events also share the features that they are concerned with voluntary financial contributions to a public good and that the contributions are based on windfall endowments. The contribution levels in the 2011 experiment are clearly unusually high compared with what is typically found in public good lab experiments, in particular given the small marginal per capita return. There are several potential reasons for this. Contributions in one-shot public good experiments are typically larger than average contributions in multi-round public good experiments (e.g., Fischbacher and Gächter, 2010). Also, there is some evidence that non-student samples might be more cooperative and show higher levels of reciprocation than student samples (e.g., Gächter et al., 2004; Falk et al., 2013). Moreover, the participants in this particular non-student sample may know each other better than most students participating in lab experiments, and there is (not surprisingly) evidence that subjects tend to behave more pro-socially when the social distance to the other subjects is small (e.g., Hoffman et al., 1996). Finally, it is also possible that the fact that the experiment took place in the subjects' homes may have induced stronger psychological pressure to contribute than if it had been undertaken in a common venue, which appears natural given the evidence that subjects' contributions tend to be affected by whether or not the experimenters directly observe the contributions (List et al., 2004; Alpizar et al., 2008). Yet, we are not primarily interested in the various contribution levels, or whether these levels are similar across decisions. Instead we want to find out to what extents decisions are correlated, i.e., whether or not those households that contribute more in one event also tend to contribute more in another. Raw contribution correlations between the events ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As described above, we observe the contributions in each of the events at the household level. As a first step, we therefore analyze the simple pair-wise correlations between the four events. Remember that, for each household, we have three observations of contributions to the bridge and one observation of contribution in a public good experiment. We present correlation coefficients for the whole sample and for the restricted sample of households that had the possibility to contribute in 2010. For those that were not asked to make labor contributions, we set the contribution to zero when calculating the correlations for the whole sample. Table 3 presents the pair-wise correlation coefficients. Despite the large differences in contribution levels between the events, including in the fraction that did not contribute anything, the correlation coefficients between the four events are substantial and in most cases highly significant. For the events related to the bridge, the largest correlation coefficients are found between the 2005 and 2010 events. This may seem surprising since the 2005 event concerns monetary contributions, while the 2010 event concerns labor days. Also, it seems likely that some people have a comparative advantage in labor contributions, implying that there is scope for a degree of specialization in contributions, which should reduce the correlation coefficient. Our favorite explanations for the large correlation coefficient are as follows: First, in both cases the subjects had to pay with their own resources (money and time, respectively), and hence there were no windfall resources obtained for the individual decision. Second, and perhaps even more importantly, both of these events were non-anonymous, and if some people are more sensitive to the peer pressure to contribute, then they should contribute more than others in both instances, implying a positive effect on the correlation coefficient. Indeed, Alpizar et al. (2008) demonstrate in a field setting that anonymity matters for charitable contributions and Andreoni and Rao (2011) found more recently that communication per se seems to dramatically affect altruistic behavior, primarily through increased empathy. The correlation coefficients between the experiment in 2011 and the other three events are also substantial, with the exception of the event involving labor contributions in 2010. Here, the coefficient is not statistically significant at conventional levels when based on the restricted sample. There may be several reasons for this, in addition to the fact that only one of the two is related to the bridge: one concerns labor contributions while the other concerns monetary contributions, one is anonymous while the other is not, and one was conducted based on windfall money while the other was not. Yet, it is interesting to see that the correlation coefficients between the contributions in 2005 and 2011 and between the contributions in 2009 and 2011 are large and statistically significant, despite the large differences in set-up. Together, this clearly shows (i) that the strong correlations between the events cannot only be due to the fact that they concern contributions to a similar good, i.e., the bridge, and (ii) that there is clear support for the idea that social preferences reflect traits that to a large extent are constant over time and domains. Simple comparisons of Low, Intermediate and High contributors' behavior over time ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Correlation coefficients are not always straightforward to interpret, in particular as in our case where the contribution patterns are very different between the events and where large fractions of the subjects chose the extreme alternatives of contributing zero or the full amount. We therefore also illustrate the stability of social preferences by investigating how subjects who initially contributed small or large amounts in 2005 contributed in the three subsequent events. Based on the histograms above it appears natural to classify those who contributed strictly more than 100,000 dong as High contributors, implying that 9% of the subjects are characterized as High contributors in 2005. Since as many as 47% of the subjects contributed zero in 2005, this constitutes a natural (albeit quite large) group of Low contributors (implying that the remaining 44% are classified as Intermediate contributors). The subjects for the three subsequent events were similarly classified as Low, Intermediate or High contributors. For the event in 2009, Low contributors are those contributing less than or equal to 100,000 dong (20% of the subjects) whereas High contributors are those who contributed the maximum amount 400,000 dong (43% of the subjects). For the event in 2010, Low contributors are those who did not provide any labor days (77% of the subjects) and High contributors are those who provided two or more labor days (11% of the subjects). Finally, for the event in 2011, Low contributors are defined as those who contributed less than or equal to 50,000 dong (22% of the subjects) and High contributions are those who provided the maximum contribution 200,000 dong (40% of the subjects). Table 4 reports the shares of Low, Intermediate and High contributors in the events 2009 to 2011 based on three different samples, those who in 2005 are classified as Low, Intermediate, or High contributors, respectively. In general, we find strong links between contributing behavior in the 2005 event and the behavior in the subsequent events. Since the overall fractions of Low and High contributors are different from each other in each of the experiments the relevant comparisons to make are between the different samples, i.e., to compare the different shares horizontally.13 That is, we want to compare whether those who in 2005 are classified as a High contributors are more likely to be classified as a High contributors in the three subsequent events than what those who are classified as a Low contributor in 2005 are. Correspondingly, we want to compare whether those who in 2005 are classified as a Low contributor are more likely to be classified as a Low contributor in the subsequent events than what those who are classified as a High contributor in 2005 are. As can be seen in Table 4, the pattern follows expectation for each of the six cases. For example, it is striking to compare the non-anonymous voluntary contributions to the bridge in 2005 with the contributions in the anonymous public good game (not related to the bridge) six years later. Of those classified as Low contributors 2005, 32% are classified as High contributors 2011. This can be compared with those classified as High contributors in 2005, where as many as 82% are classified as a High contributor 2011. Similarly, of those classified as Low contributors 2005, 30% are also classified as Low contributors 2011. This can be compared with those classified as High contributors in 2005 where as few as 6% are classified as a Low contributor 2011. In order to formally test whether these differences are statistically significant, we focus on the two samples of Low and High contributors in 2005. We can then use simple Chi-square tests when comparing the distribution of Low, Intermediate and High contributors for each of the 2009, 2010 and 2011 events. For example, we can test if there is no difference in the distribution of Low, Intermediate and High contributors in the 2009 event for those that were Low contributors in 2005 and those that were High contributors in 2005. In all three cases, i.e. for each of the events in 2009, 2010 and 2011, we can reject the null hypothesis that the distribution is independent of whether a subject was classified as a Low or High contributor in 2005, at the 1% level.14 Econometric analysis ~~~~~~~~~~~~~~~~~~~~ While the strong positive correlation coefficients, as well as contribution patterns more generally, between contributions in the first three events (i.e., those related to the bridge) are interesting per se, one should be hesitant to interpret these coefficients as clear evidence of stability of social preferences. Indeed, there are several possible interpretations of these positive correlations. For example, if the households using the bridge the most are also willing to contribute the most (e.g., for partly selfish reasons), we should obtain a positive correlation between contributions in the first three events even if there are actually no differences between the households in terms of underlying social preferences. Still, the use of the bridge can of course not explain the correlation coefficients between contribution in the public good experiment in 2011 (which was not related to the bridge) and contributions in any of the other events. We deal with this problem in two ways: First, as mentioned, we run a public good experiment without any reference to the bridge, based on the same sample. Second, we use regression techniques correcting for possible explanatory variables that can be assumed to vary across the households but at the same time are independent of underlying differences in social preferences. The most obvious variable here is the use of the bridge. More specifically, we use multivariate tobit regressions since we have non-negligible shares of subjects who either contribute the full amount or do not contribute at all; hence, we use truncations at both zero and the full amount for the experiments in 2009 and 2011, and at zero for the events in 2005 and 2010. Using a multivariate model, we estimate the correlation coefficients of the error terms for each event. These error terms are assumed to reflect the part of social preferences that cannot be explained in terms of our explanatory variables used in the regressions. Moreover, simple correlations do not take into account that there were different treatments in the experiment in 2009. In order to deal with these issues, we estimate multivariate models where the four equations are estimated simultaneously, allowing for a correlation between the error terms of each of the equations, and where the dependent variables are censored. We present three sets of regressions: In the first set, we use no explanatory variables (except for an intercept). In the second set, we use only variables reflecting the use of the bridge in the first three events, since these variables presumably vary across the households and at the same time are independent of underlying differences in social preferences. For the second and fourth event we also include treatment dummy variables and experimentalist dummy variables. Finally, we present a third set of regressions including all relevant explanatory variables. In this last set, we thus face the risk of “over-compensation” in the sense that there may exist variables, such as age and income, that are correlated with true underlying social preferences. For example, suppose that all actual variations in social preferences are determined by gender. If we then correct for gender in the regressions, the result would indicate that there is no stability of social preferences over time even though the actual stability through gender may be large. Yet, as is the case when not including any explanatory variables, it constitutes a natural benchmark case. We focus mainly on the sample of households that had the possibility to contribute labor in 2010. However, we also report the results based on the full sample, where we have hence set the contribution in labor to zero in 2010 for those who were not asked to contribute. We also use the full dataset for the pairwise correlations that do not include the 2010 event. Yet, as can be observed, the results turned out fairly similar.15 The estimated correlation coefficients for our three sets of multivariate regressions are presented in Table 5, for each separate event. Starting with the first three events, we can observe that the pairwise correlation coefficients are consistently positive, substantial, and statistically significant. Consequently, even when controlling for a number of observable differences among households and the treatment effects, there are strong correlations in behavior between the three events. The relative sizes of these coefficients follow expectations in that they are generally largest when we do not correct for any variables, and smallest when we include the full set of variables. Yet, the differences between when we correct for the use-of-the-bridge variables and when we do not are rather small. Again, the highest correlation coefficients are found between the 2005 and 2010 events, probably largely due to the fact that these events were not anonymous, as discussed previously. Yet, it could still be the case that we did not manage to completely correct for the usefulness of the bridge to different households, and that the remaining part could explain the positive correlations. For this reason, we have the fourth artefactual field public good experiment, which is completely unrelated to the bridge. Here we again find a weak effect between 2010 and 2011 (we speculate about possible reasons above), while the large and highly significant correlations between contributions in 2005 and 2011 and between contributions in 2009 and 2011 largely prevail after including various covariates. Consequently, we can again conclude that it is not the fact that three of the four events concerned the same underlying good, a bridge that explains the significant correlation between the contribution decisions. Consider next the estimated coefficients for the covariates. Table 6 presents the results for the restricted sample. Few of the household characteristics have a significant impact on the contributions in any of the events and there are not any consistent patterns across the different events (except for the use of the bridge). The contributions in 3 out of the 4 events, including the public good experiment in 2011, are positively related to the amount of rice land owned. Monthly income, in contrast, is only positively correlated with contribution in the public good experiment. One may speculate about the reason for the difference. Perhaps it is more natural to relate to own income when making the contribution decision in a non-framed public goods experiment than when making decisions that are linked with the bridge. For the 2010 event, when people made contributions in terms of labor time, it is not surprising that there is a negative effect of age. It is less expected to find a negative effect of the household head being male. As suggested by a referee, a rough, indicative story can be given as follows: One does not have to use the bridge very much to give out of the windfall, one needs to use it a little more to give labor (free), and one needs to use it a lot to give money if poor. Are the obtained correlation coefficients large? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Yet we also find, in line with several previous studies, that contributions are highly context dependent. Related to this, we find that some correlation coefficients are substantially larger than others, whether corrected for other explanatory variables or not. Perhaps most strikingly, we obtain a correlation coefficient between voluntary monetary contribution in 2005 and voluntary labor contribution in 2010 in the order of magnitude of 0.4 or larger, despite the fact that almost 50% contributed nothing in 2005 and over 70% (of the restricted sample that were asked) contributed nothing in 2010. Our conjecture is that this finding may not only reflect stability of social preferences, but also to some extent stability of what may be called sensitivity to social pressure, since none of these events were anonymous. This is an important observation in its own right, and calls for further research. Overall, we conclude that social preferences seem to be quite constant over long periods of time.","In this paper we have compared voluntary contributions to a public good, in terms of a bridge in rural Vietnam, for the same complete sample of about 200 households in a village, at four different events spanning over a time period of 6 years. By doing so, we have been able to avoid the potentially confounding factor related to moral licensing and moral cleansing when measuring the extent of pro-social stability over time. Overall, we find substantial and highly significant correlation coefficients, suggesting that pro- social preferences are quite constant over long periods of time. Although not our main research task, the substantial and positive correlation between the unframed economic experiment and naturally occurring events supports the idea that social preferences obtained in economic experiments have validity also outside the somewhat artificial experimental context. Although our events took place in a village that is typical for this part of the world, it is an open question whether there are large cultural differences in the extent to which social preferences are constant over long periods of time. Previous findings have concluded that there are non-negligible differences in the strengths of social preferences, as measured by economic experiments, in different cultural contexts (e.g., Henrich et al., 2005, 2010). Our conjecture is nevertheless that the extent to which social preferences vary over time does not vary much culturally, although this is of course an open question. For this and other reasons, we encourage further studies in the field in order to test the robustness of our findings."],["What kinds of tariff reform are likely to raise welfare in situations where tariff revenue is important? General conditions for welfare to rise without reducing tariff revenue are opaque. We show that they can be greatly simplified using a small number of sufficient statistics, primarily the generalized mean and variance of tariffs. We present sufficient conditions for a class of linear tariff reform rules that guarantee higher welfare without a loss in revenue. The rules consist of convex combinations of (i) trade-weighted-average-tariff-preserving cuts in dispersion; and (ii) uniform tariff cuts that preserve domestic relative prices among tariff-ridden goods. --------------------------------------------------------------------------------","What kinds of tariff reform are likely to raise welfare in situations where tariff revenue is important? The question is an important one: despite steady reductions in average tariffs, tariff revenue is still a significant component of total tax revenue, especially in low-income countries. Baunsgaard and Keen (2010) review the empirical evidence on the revenue effects of trade liberalization in recent decades, and conclude that, while middle-income countries have managed to offset reductions in trade tax revenues by increasing their domestic tax revenues, many low-income countries have not. Even in rich countries, the revenue effects of changes in tariffs can be substantial in absolute if not in relative terms, and can be a factor influencing the decision to liberalize trade. The implications of trade reform for revenue have featured prominently in discussions of the EU's association agreements with countries in the Southern Mediterranean region (see Abed, 1998), and even in official discussions of the case for the U.S. joining NAFTA (see Congressional Budget Office, 1993).1 Unfortunately, as we shall see, general conditions for welfare to rise without reducing tariff revenue are opaque, and provide little guidance to practical policy-making. Our main contribution is to show that they can be greatly simplified using a small number of sufficient statistics, primarily the generalized mean and variance of tariffs. Reexpressing the general conditions in terms of these sufficient statistics leads to new operational guidelines for tariff reform that guarantee higher welfare without a loss in revenue. The rules consist of convex combinations of cuts in tariff dispersion that preserve the trade-weighted-average-tariff, on the one hand, and uniform tariff cuts that preserve domestic relative prices among tariff-ridden goods, on the other. These guidelines provide a theoretical foundation for the standard World Bank advice to developing country clients that they should reduce dispersion of tariffs while maintaining average tariffs to preserve revenue. In plausible special cases, the rules require only observable data and a small number of aggregate elasticities. Our approach builds on the sizeable literature on trade policy reform in open economies, stemming in particular from Hatta (1977). Much of this literature provides guidelines for welfare-improving tariff reform when government revenue is not a concern, which amounts to assuming that the government has lump-sum tax/transfer power. This approach has been extended to study the interplay of revenue and efficiency considerations in trade policy reform by a number of authors, including Falvey (1994), Emran and Stiglitz (2005), Hatta and Ogawa (2007), and Raimondos-Møller and Woodland (2015). However, these papers either use relatively special low-dimensional models, or do not provide rules that can be easily implementable. An exception is a branch of the literature which advocates replacing border taxes with domestic consumption taxation. (See for example, Hatzipanayotou et al. (1994), Keen and Ligthart (2002), and Kreickemeier and Raimondos- Møller (2008).) The intuitive argument that the base is broader can be supplemented with optimality considerations. Diamond and Mirrlees (1971) demonstrated that it is inefficient to distort productive efficiency when raising revenue with distortionary taxation. Trade taxes, by subsidizing production, drive a wedge between domestic and international marginal rates of transformation. However, Anderson (1999) shows that gradual reform of this type need not improve welfare when uniform radial reductions are used to lower tariffs. The present paper admits a much broader class of trade reforms when wage taxation is the alternative revenue source and provides more optimistic prospects for tariff reforms which reduce dispersion. The present paper draws on Anderson and Neary (2007), where the approach using generalized moments of the tariff structure was introduced and applied to devising rules for trade policy reform in the conventional setting of no revenue constraint. That paper derived linear welfare-improving reform rules as implications of reform that reduced either or both of two sufficient statistics, the generalized mean and generalized variance of the tariff structure. Here we extend these methods to the case where lump-sum taxes and transfers are not feasible and so the government faces a binding revenue constraint. All government tax changes become costly at the margin because they involve distortions. The same sufficient statistics prove useful in the case of an active revenue constraint, supplemented by some additional aggregate elasticity terms. In a big step toward applicability with very limited information, Anderson and Neary (2007) also showed that, in a special CES case, the generalized mean and variance reduced to the readily observable trade-weighted version of these statistics. A second contribution of the present paper is to demonstrate that observability of generalized moments obtains with weak separability, nesting not only the CES but most other widely-used preference and technology structures. A group of goods such as clothing under separability can contain pairs that are complements (shirts and trousers) and other pairs that are substitutes (cotton and silk shirts). The separable setting also permits a further realistic extension which replaces the representative agent with heterogeneous agents while maintaining feasible operational rules that yield Pareto improvement. Section 2 sets up the model and derives the general expressions for tariff reform in the presence of revenue constraints. Though insightful, these do not easily lend themselves to practical implementation. The remainder of the paper shows how they can be operationalized using the tariff moments approach introduced in Anderson and Neary (2007). Section 3 reviews and extends that approach, while Sections 4 and 5 use it to analyze trade reform and to derive the main results of the paper. Section 6 extends the results to the case of many households, while Section 7 concludes. The setting ~~~~~~~~~~~ The tariff reform problem is to advise on directions of change of tariffs from initial values. Full optimization is not feasible by assumption. The setting is a competitive small open economy which raises its revenue with a set of tariffs and with a wage tax. For simplicity we will present the results in terms of a perfectly competitive economy though, as shown in Anderson and Neary (2005), the results also apply to a variety of monopolistically competitive models with fixed entry costs and firm heterogeneity. The wage tax is distortionary because labor supply is variable (due to household choice in an economy where immigration is shut down) and leisure cannot be taxed. Tariffs and the wage tax are initially set sub-optimally. The objective of the reform is to move the taxes gradually toward their optimal (Ramsey) values. This section first describes the economy and then derives general expressions which show how tariff changes affect welfare and tariff revenue. These results are the key building blocks for our results in later sections that are expressed in terms of tariff aggregates. The representative consumer's net expenditure function is given by e(π, w, u). It gives the minimum spending needed to sustain a level of utility or real income u when the representative consumer faces a vector of prices of traded goods subject to tariffs, denoted by π, and a net-of-tax wage rate denoted by w. The domestic prices π differ from world prices π* by a vector of specific tariffs t. Since the economy is small, the world prices are exogenous, so changes in domestic prices dπ are equal to changes in tariffs dt throughout. Implicit in the list of arguments of e is the price of a composite export good, which we take as numeraire so its price can be set equal to one. By Shephard's Lemma, eπ gives the vector of final demand for traded goods, while − ew gives labor supply. As for the supply side of the economy, the maximum value of GDP which can be produced given its technology and facing goods prices π and a gross wage w + τ, where τ is the tax on labor income, is given by the GDP function g(π, w + τ). By Hotelling's Lemma, the vector of supply of traded goods (or where appropriate, minus the demand for traded inputs) is given by gπ while − gw gives labor demand. E gives the net transfer to the private sector needed to support utility u when domestic prices of traded goods are set at π and the wage tax is set at τ. Its derivative with respect to π, Eπ, is the vector of excess demand for traded goods, which equals the vector of net imports m; while its derivative with respect to τ, Eτ = − gw, is equilibrium employment, where the maximization with respect to w ensures that the labor market clears: ew = gw. Since e-g is concave in (π, w, τ), E is concave in (π, τ): compensated net import demand functions are downward-sloping, and a higher wage tax reduces employment. Here, s is the transfer from the government to the private sector. If the government has lump-sum power, s is an active policy instrument. Otherwise, it is an exogenous transfer, which also serves as a useful analytic link between the private-sector and government budget constraints. Here, R0 represents the government's revenue requirement, to fund public goods, repay foreign loans, or finance some other goal which does not directly affect private-sector decisions. Extending the model to endogenize this revenue requirement is straightforward, along the lines of Atkinson and Stern (1974), or, in a trade context, Abe (1992). We do not pursue this approach here, since it distracts from our primary focus, using summary statistics to simplify the guidelines for tariff and tax reform. In this case, the government's revenue requirement is not an independent constraint on policy-making because the transfer s adjusts endogenously. This makes a crucial difference for evaluation of tariff reform: Eq. (4) leads to the standard results of piecemeal trade policy reform, augmented to allow for an exogenous wage tax. (See Appendix A for details.) In the setting considered in the rest of the paper, by contrast, lump-sum transfers are infeasible, so the gradual reform problem is to determine welfare- improving directions of change in the set of reformable tariffs t, equivalent to changes in π, while at the same time not decreasing revenue. One class of reforms takes the wage tax as given and looks for tariff reforms that raise both welfare and revenue. A more ambitious class of tariff reforms permits the wage tax to vary endogenously in order to maintain government revenue. We consider each of these classes in turn. The problem ~~~~~~~~~~~ A diagram illustrating the three-good case, two of them subject to tariffs, aids intuition. In Fig. 1, initial tariffs are such that domestic prices equal πA, corresponding to point A. Point F represents free trade, but of course that yields zero revenue. Optimal revenue-raising tariffs imply Ramsey-optimal prices πR, corresponding to point R. The tariff reform problem is to devise rules that will improve welfare locally; directions of change for π starting from A that bring the economy closer to R in the sense of attaining a higher iso-welfare contour.4 However, the shape of the iso-welfare contours depends on the policy instruments available to the government. The line though point A labeled dTa = 0 shows combinations of domestic prices that keep the trade-weighted average tariff Ta constant. As we will see below, it is also an iso-welfare locus for the case where the wage tax is given: i.e., it is implicitly defined by the private-sector budget constraint (2) for given u, τ, and s. By contrast, the elliptical locus drawn through point A shows combinations of domestic prices that keep welfare constant when the wage tax adjusts to maintain revenue: it satisfies both the private-sector budget constraint (2) and the government budget constraint (3) for given u (s is now irrelevant) and with τ adjusting endogenously. This locus is one of a family of iso-welfare curves centered around R: each raises the same amount of revenue but yields successively higher levels of welfare as R is approached. As drawn, the locus encloses a convex set of π's and is upward-sloping at A, but these properties are not guaranteed. To determine more precisely the location of these iso-welfare contours and hence the desired direction of tax reform away from A, we need to develop a more formal analysis. Tariff changes only ~~~~~~~~~~~~~~~~~~~ We will assume that this term lies in the unit interval: RI is positive and less then one. A host of arguments has been given in the literature on piecemeal policy reform to defend this assumption. Normality suffices, as does homotheticity or a standard stability condition.5 Violation of the assumption would be perverse indeed, since it would imply that a gift of foreign exchange to the private sector, enabling a rise in real income, would at constant prices π either reduce government revenue or raise it by more than the value of the gift. In the presence of lump-sum redistribution, moreover, a negative value of RI would imply that gifts make the economy worse off. The left-hand side suggests a clear presumption that more for the government means less for the private sector. However a positive value for the term on the right-hand side can offset this presumption, permitting a rise in both real income and revenue. This possibility arises from reforms that remove inefficiency in the tariff structure. Below, we characterize such possibilities in terms of tariff moments. The marginal cost of funds ~~~~~~~~~~~~~~~~~~~~~~~~~~ When we turn to consider choices between different forms of taxation, it is useful to express our results in terms of the marginal cost of funds (MCF) of different instruments. Consider the cost to the government of supporting the representative agent's real income u with a hypothetical subsidy ds when the wage tax τ changes to raise revenue R by one dollar. Recalling that E is concave in τ, the direct substitution effect of a wage tax on labor supply Eττ tends to reduce Rτ below Eτ, and so encourages a value for the social cost of funds greater than one. This could be offset by the cross effect: if leisure is a complement for imports, so Eπτ is positive, a rise in the wage tax τ increases tariff revenue, encouraging a value for the social cost of funds less than one. However, values greater than one are typically found in applied studies and must be considered the norm. Similar considerations, mutatis mutandis, apply to the magnitude of the marginal cost of funds of any other tax instrument. Tariff changes compensated by wage tax changes ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Although we will hold revenue fixed, it is insightful to include the change in the revenue requirement R0 in Eq. (12), since it allows us to interpret the coefficient of the welfare change on the left-hand side. Consider the effect of a hypothetical transfer to the government which reduces the revenue requirement R0 by one unit. The impact effect of this on welfare is equal to the marginal cost of funds μτ. In an economy with tariff and wage tax distortions this impact effect has multiplier repercussions, which are measured by the inverse of the coefficient of Eudu. This shows that the term 1 − μτRI is the inverse of the shadow price of foreign exchange modified for the endogeneity of the wage tax. Moreover, as with the corresponding term in Section 2.3, we are justified in assuming that it is positive: a negative value would imply that a gift to the economy, which relaxes the revenue requirement R0, would lower real income.6 The intuitive implication of Eq. (13) is that reducing tariffs on all goods for which μiπ > μτ and increasing tariffs on all goods for which the inequality is reversed will produce a surplus. This in turn causes an increase in real income, provided the shadow price of foreign exchange is positive. This provides an insightful contrast with the usual expression for welfare change in the theory of piecemeal tariff reform when lump-sum taxes are available (see Eq. (34) in Appendix A), and it reduces to it when labor supply is fixed so a wage tax is effectively lump-sum (μτ = 1 and Eτπ = 0). However, saying more about the tariff reform problem using Eq. (14) as it stands is challenging. To make progress with this problem, we turn in the next section to extend the method of tariff moments developed in Anderson and Neary (2007) to the present context. Summary statistics for the structure of tariffs ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The key intermediate step in the analysis of trade reform is a decomposition of the effect of tariff changes into their effect on various moments of the distribution of tariffs. Anderson and Neary (2007) examine welfare-improving directions of tariff reform in the case where revenue considerations are unimportant, so μτ = 1. Here we extend their moments decomposition technique to the revenue tariff problem. Table 1 summarizes the notation. Except for the trade-weighted average tariff Ta, these generalized moments and their changes are complicated functions of consumer and producer behavior. Nonetheless, they summarize the implications of the full matrices of aggregate demand and supply responses in an intuitive and parsimonious way. In this respect, they are analogous to the shadow price of foreign exchange introduced by Hatta (1977): it allows all income effects to be summarized in terms of a single convenient aggregate statistic, whereas earlier work typically made strong assumptions about income effects on a commodity-by-commodity basis, such as requiring all goods to be normal.7 In the same way, the generalized tariff moments provide a set of sufficient statistics for the substitution effects in the economy. As we will show, analytic expressions in changes in generalized means and variances help formulate linear tariff change rules that guarantee welfare improvement even in the absence of detailed information about substitution effects.","Under separability as in Eq. (18), cuts in tariff dispersion which leave the trade- weighted average tariff unchanged raise revenue while not harming welfare. As discussed in Section 2.4, there is a presumption that the marginal cost of funds is greater than one for each individual good, so it must be considered highly unlikely that this marginal cost of funds of a composite group could be less then one. We also expect RI, the effect of a unit gift of foreign exchange on government revenue, to lie between zero and one, as discussed in Section 2.3. Given this, the sign of the right-hand side of Eq. (23) is ambiguous, although there is a presumption that the direct price effect 1/μT outweighs the income effect RI: uniform absolute reductions ordinarily imply that revenue falls: 14 Uniform absolute reductions in T raise both welfare and market access, but raise revenue if and only if the inverse of the marginal cost of funds for all tariff-ridden goods, μT, is less than the income responsiveness of revenue, RI. Summarizing this section, given that very large dispersion is common in tariff structures, even in countries that raise a substantial portion of government revenue from tariffs, Proposition 2 implies considerable scope for efficiency improvement from dispersion cuts. On the other hand, Proposition 3 implies that absolute tariff cuts with constant dispersion decrease revenue, which creates a presumption against average tariff reductions as part of a reform package when tariff revenue is important (i.e., when tariffs are the only instrument). Tariff reform rules in terms of generalized moments ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The square bracket term is smaller than the inverse of μT under the conditions of (i), and hence the entire expression is positive under the condition of (ii). □ Trade-weighted mean-preserving dispersion cuts within separable goods classes; Any convex combination of such dispersion cuts and a uniform absolute tariff change across as well as within classes that decreases tariffs when they are over-utilized or increases them when they are under- utilized. Proposition 5 is proved in the Appendix. The key element is that, from Proposition 1, the condition of Proposition 4 is met under separability. The proposition is quite useful because separability is a ubiquitous assumption in applied work. Faced with some ten thousand tariff lines, aggregation is inevitable for any econometric or simulation work. The proposition guarantees that trade-weighted average preserving dispersion cuts within classes are welfare-improving without detailed knowledge of substitution effects (either parameter values or specification) within goods classes. National tariff schedules are full of dispersion in detailed product classes, so there is a lot of room in practice for beneficial cuts. It is worth noting that, under separability, a trade-weighted mean-preserving tariff dispersion cut improves welfare strictly by raising government revenue; trade expenditure remains constant under this reform. The CES special case ~~~~~~~~~~~~~~~~~~~~ The CES expression (31) for the marginal cost of funds reveals that the focus of Propositions 4 and 5 on convex combinations of mean-preserving tariff cuts and dispersion- preserving mean cuts does indeed capture all the relevant characteristics of welfare- improving revenue tariff reform which can be guaranteed without full knowledge of substitution effects. If exact values of η and σ are assumed to be known, it is of course possible to improve welfare with tariff reforms outside the cones based on Eq. (31).20 As substitution possibilities range more widely beyond the CES, more welfare-improving revenue tariff reforms can be found which are not within the cones of Propositions 2, 3, and 4. But again, showing that these reforms raise welfare depends on information that this paper assumes, realistically, that the analyst is unlikely ever to have with any certainty. Note that the CES expression sheds light on the esoteric possibility that some tariffs may actually have a marginal cost of funds less than one. From Eq. (31), the necessary and sufficient condition for μiπ < 1 is (1 − η/σ)Ta > Ti. The sufficient condition requires either that η/σ < 1, substitution elasticities within the separable group exceed substitution elasticities between that group and all other goods, or that good i is subject to an import subsidy, so Ti < 0. Normally neither condition would be met. The desirability of dispersion cuts ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Returning to Fig. 1, the ray OR through the Ramsey optimal tariff point R divides the domestic price space into half spaces. Starting at point R, consider a mean-tariff- preserving line (not drawn) to the uniform tariff ray OF. For points on this line between the uniform tariff ray OF and the optimal tariff ray OR, trade-weighted mean-preserving dispersion increases are welfare-improving. For points in the space below ray OR, dispersion increases are welfare-decreasing. If the cone FOR is small, the World Bank intuition about the desirability of dispersion reduction holds for most of the tariff space.","The preceding expressions extend with appropriate modification to the case of many households. The government budget constraint continues to hold using E for the aggregate trade expenditure function and its derivatives, while Eh denotes the individual household h trade expenditure function. Summing over households h, the first two terms cancel out. The condition that reform be beneficial in the aggregate (representative agent) case is that the third term be negative. Propositions 4 and 5 apply. The potential for individual loss is confined to the deviation due to the balance of the first two terms. Agents can differ in their tastes for work vs. consumption, generating differences in the aggregate weights attached to the average tariff differentials, and they can differ in their consumption patterns within the tariffed goods bundle when faced with the same price vectors. The latter results in dTa,h ≠ dTa while the former results in Eπh ⋅ π ≠ (Eτh/Eτ)Eπ ⋅ π. What minimal information is needed to specify welfare-improving rules for each household (Pareto superior rules)? Tariffs are widely levied on intermediate goods. In this case there is no household-specific weighting, Tai = Ta, so dispersion cuts are Pareto-superior. As for final goods, assume that imported goods in a separable goods class have no domestic perfect substitute, and that household expenditure patterns Eπh are observable. The former is a widely used empirical assumption because the perfect substitutes assumption yields implications wildly at variance with the trade data. The observability of household expenditure patterns is a more problematic assumption but it is satisfied for a number of countries. Under these assumptions, the βh parameters can be set equal to the household level trade-weighted average tariff Ta,h to implement the mean preserving dispersion cut: dTh = (T − Ta,h)dα. The mechanism is a uniform deviation from the common tariff cut rule for each household: dTh − dT = (Ta,h − Ta)ιdα. All tariffs are changed according to dT = (T − Taι)dα. Implementation of the household specific deviations could presumably take place at the retail level (as with food stamps or senior citizen discounts), supplemented by some governmental identification system. Doing so, for example, all clothing tariffs change according to the common rule, then each household receives or pays its household specific deviation (Ta,h − Ta)dα. Alternatively, the implementation could be done through income tax credits. To avoid shirking, the common rule could be set around the highest Ta,h, so that all households with lower average tariffs receive a rebate. In this scheme of tariffs, the real income of each household is maintained, the individual variation of βh is revenue neutral since ∑h(Ta,h − Ta)π′Eπh = 0, and the government revenue will rise due to the revenue-increasing cut in dispersion. Thus dispersion cuts are a Pareto-superior reform. As for uniform absolute cuts in tariffs, the requirement of Propositions 2 and 3 that ‘tariffs are over (under) utilized’ becomes extremely stringent because it requires that the marginal cost of funds of the alternative revenue source be less (more) than each individual agent's marginal cost of funds of tariffs. This is seldom likely to appear plausible to analysts evaluating potential reforms. The implication is that the Pareto-superiority of dispersion cuts holds in the many household case under the separability assumption, understanding that trade- weighted average tariffs must be calculated and applied at the household level. The separability assumption is plausible for some goods classes and not for others. Still, this discussion suggests the surprisingly wide desirability of dispersion cuts.","This paper has derived rules for welfare-improving trade reform that permit confident policy advice despite the (assumed partial) ignorance of analysts about the true structure of the economy. Dispersion-reducing trade reform is surprisingly widely beneficial: whenever households have implicitly separable preferences with respect to the same partitions of goods, dispersion of tariffs within separable groups is inefficient. Cuts in average tariffs are efficient when the marginal cost of funds of such tariffs is greater than the marginal cost of funds from alternative revenue sources. Convex combinations of uniform absolute cuts and mean-preserving dispersion cuts are beneficial under these conditions. Over and above these specific results, we have shown the value of reexpressing complex results for multi-dimensional policy reform in terms of a small number of summary statistics. This approach should prove useful both in guiding empirical work and in allowing an intuitive focus on the key parameters that matter for policy evaluation."],["The impact of teacher pay on school productivity is a central concern for governments worldwide, yet evidence is mixed. In this paper we exploit a feature of teacher labour markets to determine the impact of teacher wages. Teacher wages are commonly set in a manner that results in flat wages across heterogeneous labour markets. This creates an exogenous gap between the outside labour market and inside (regulated) wage for teachers. We use the centralised wage regulation of teachers in England to examine the effect of pay on school performance. We use data on over 3000 schools containing around 200,000 teachers who educate around half a million children per year. We find that teachers respond to pay. A ten percent shock to the wage gap between local labour market and teacher wages results in an average loss of around 2% in average school performance in the key exams taken at the end of compulsory schooling in England. --------------------------------------------------------------------------------","Education in England is compulsory between the ages of five and sixteen. While children can be educated privately, the public (state) system dominates. State sector pupils attend primary school from age five to eleven and secondary school from age eleven to sixteen. Pupils can then stay on for a further 2 years to get qualifications that allow them to undertake university level education. In 2007 approximately three million young people (around 84% of eleven to sixteen year olds) were attending public secondary schools. In each secondary school there are five (or seven if the school provides education up to 18) separate age cohorts within the school at any one time. Pupils take nationally set exams at four points during their ages of compulsory school attendance. At primary school these are Key Stage 1 (KS1) at age 7 and Key Stage 2 (KS2) at age 11 (the year of exit), in Mathematics, English and Science. In secondary schools these are Key Stage 3 (KS3) exams at age 14 in Maths, English and Science and Key Stage 4 (KS4) examinations in multiple subjects (typically between eight and twelve) at the end of compulsory schooling at age 16. We focus on KS4 (GCSE) examinations as our measure of school performance as these are high stake examinations. For pupils they determine progress into education after age of 16, as a minimum of five pass grades required to continue on to further education, are used by parents to choose secondary schools for their children, by the media to rank school performance to create school ‘league tables’ and by local and central government to identify ‘failing schools’. Schools in England are heavily regulated by central and local government. Summary statistics on school performance have been published annually since the early 1990's. The key measure used to compare schools has been the number of pupils attaining at least 5 good grades in the KS4 exams (known as 5 A*-C GCSEs), though the number of metrics published increased during the mid- and late 2000's. In addition, in- depth assessments of the quality of the school are undertaken by the schools regulator, OFSTED. Each school is inspected roughly every five years. Inspections often last several days. On the basis of these site visits, OFSTED publishes a report rating the school's performance on numerous dimensions, including an overall rating of the school and an assessment of teaching and learning at the school. More details on these metrics are provided below in Section 3.1. Teacher pay in England ~~~~~~~~~~~~~~~~~~~~~~ Teacher wages are set by Local Education Authorities (LEAs) based on guidelines issued by the national Government Department for Education.2 Despite the existence of four pay bands (‘Inner London’, ‘Outer London’, ‘The Fringe’ and ‘The rest of England’), teacher wages have exhibited very little regional variation relative to private sector wages since the early 1970's. For example, the average teacher wage differential outside the South East of England and Inner London is approximately 15%, while the equivalent private sector wage differential can be as large than 45%.3 Since its formation in the 1990's, the School Teacher Review Body (STRB), an advisory board which comments on teacher conditions and pay, has frequently argued that the Department for Education should be doing more to encourage locally flexible wages.4 Although an increasing amount of discretion over wages has been granted to LEAs, and more latterly schools, they have not utilised the option (Sibieta, 2015). This is possibly due to the fact that local authorities have faced strong national teaching unions for many years, making the costs of local negotiations high for a single school or local authority (Zabalza et al., 1979).5 Further these costs are incurred before any gains are realised, so schools and the elected local authorities may place greater weight on current costs versus longer-term potential gains. In addition, there has been very little scope for schools to provide differential non-pecuniary benefits for teachers - there is no variation in holidays and contact hours are generally fixed. Recent changes to the schooling system in England which have come into play in the second decade of this century have given schools potentially more power over wages and teacher conditions. The UK government has encouraged the setting up of ‘free schools’ and academies in England. Such schools are free of LEA control and able to choose their own curriculum. While the programme was started under the Labour administrations which operated upto 2010, before the school year 2008–9 there were less than 100 such schools. The programme really accelerated after 2010 (Sibieta, 2015). To avoid contamination we confine our analysis to before 2008 (we discuss the data in detail in Section 3.1). How centralised pay may affect school performance ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We follow Hall et al. (2008) and propose a simple dual-region model of the English market for teachers. In this market, ‘The North’ has lower living costs and fewer outside options relative to ‘The South’. Even when controlling for worker composition, the local private sector wage is therefore lower in the North. Because of these factors, for each given wage, teacher supply is higher in the North than in the South. An ideal pay structure would therefore allow differential wages in each region to equalise supply and demand. As shown in Fig. 1, by setting the centrally regulated wage to be constant across the two regions at WC, even if on average the regulated wage is at equilibrium, a wedge exists between it and the equilibrium wage at the regional level. In this model the regulated wage acts as a pay ceiling in the South. This model presents the case of an invariant regulated wage across regions. In England there is some wage variation across four large geographical regions but, as can be seen from Fig. 1, unless the regional variation is such that the teacher wage in the South is set equal to WS and the teacher wage in the North is set equal to WN, the nature of the problem persists; disequilibria in local markets will remain, affecting teacher supply in certain regions. Based on the lack of variation that we observe in teacher wages compared with private sector wages (and indeed the focus of the STRB on the issue), it is highly unlikely that the regional variation in England goes far enough. This model highlights the possibility of insufficient supply in high wage areas, but Fig. 1 would be unchanged if it were referring to the supply of quality teachers, or indeed the supply of effort of teachers; in either case there would remain a shortage in the South as a consequence of the invariant wage.6 This highlights that the effect of an invariant wage on school performance in high wage areas could work through a number of mechanisms which relate to both lower effort and the sorting of lower quality teachers to relatively lower paid areas (Lezear, 2000). In terms of sorting, problems may arise in recruitment and retention. First, in England, Dolton (1990) finds that wages are an important factor in recruiting good teachers and Ma et al. (2009) find a negative relationship between relative teacher wages and posted Local Authority level teacher vacancies. Second, public sector wage increases in the UK have been shown to improve the qualifications of new public sector workers (Nickell and Quintini, 2002), suggesting the negative effect may not be seen just through vacancies, but also through reduced teacher quality. In terms of effort, teacher quality has been shown to be important for school performance (Barrow and Rouse, 2005; Rockoff, 2004; Benton et al., 2003; Rivkin et al., 2005; and for England, Slater et al., 2012) and teachers have been shown to be adversely affected by lower quality colleagues and by high turnover rates (Ronfeldt et al., 2011). There is scope for reductions in effort in response to lower relative wages as the nature of teaching in England means a large proportion of the work is discretionary (time spent lesson planning, engagement in after-school programmes, time invested worrying about particular children). Two early English studies of school performance suggest that relative pay is important but neither test this hypothesis. Gordon and Monastiriotis (2007) investigate neighbourhood and regional effects on education performance and conclude that schools from some of the most affluent areas perform worst relative to expectation. They attribute this to ‘crowding out’ of public sector activity in affluent areas. Zabalza et al. (1979) examine English secondary schools in the 1960s and find fewer qualified teachers and higher turnover rates in London compared to the rest of the country and attribute this to the poor relative wages in London.","We examine the relationship between local wages and school level productivity. Our main measure of school productivity is value added by the school in key national exams at the end of compulsory schooling. We exploit the fact that there are national exams taken by all students immediately prior to secondary school entry to control for initial ability of pupils. In a set of extensions we also explore other measures of school productivity. In our case, the inside wage is the regulated wage which is set over a large region (there are four in the whole of England, a country with a population of around 53 million). The outside wage is estimated from wages in the local labour market. The use of a one period lag in wages in Eq. (1) is problematic in the context of school production in England, where pupils normally attend the same secondary school for five years. This has two consequences. First, since education is cumulative and final examination results will depend on the education a pupil received in all of the years they attended the school, it is likely that there will be long lags in the effect of the outside wage.7 Thus a regression with only one lag in wages is likely to suffer from omitted variable bias. While in principle we could estimate Eq. (1) with 5 lags of wages, in practise the outside wages in year t are likely to be not dissimilar to those in year t-1, so identification of the separate effect of each year's wages will be difficult. Second, since teachers teach children across year groups, different year groups within one school will be subjected to the same shocks. This is likely to create high levels of serial correlation in the outcome data. In Appendix B we examine the impact of these two problems. These estimates show the problem of estimating a full dynamic model, but they also confirm that in our data there appears to be a negative relationship between outside wages and value added for almost all lags of the wages, and that this increases over time. The explanatory variable of interest is the gap between the school outside wage averaged over five years from t-1 to t-5 and the regulated inside wage at time t-1. The dependent variable is school level mean KS4 points obtained at time t for school i. KS2i,t-5 is the intake performance (measured in the last year of primary school) of the pupils who take their KS4 in year t. This boils down to regressing changes in exam scores on changes in outside wages, keeping constant any relevant Xit and baseline exam scores. Conditioning on the fixed effects at school level, this is like a difference-in-difference analysis. A disadvantage is that as school performance data which contains both KS4 and KS2 scores is only available from 2002 onwards we only have two observations per school, in 2007 and 2002, matched to wage data starting in 1997. However, we have a large sample of schools, and the two observations per school allow us to control for time invariant heterogeneity at school level. The outside wages are estimated from local area data, so to deal with this we bootstrap standard errors. Finally, as outside wages are labour market variables and not school level variables, we cluster standard errors at the (larger) local labour market.","We use several sources of administrative data. At the core of our analysis is data on school performance matched to data on the gap between local outside labour market wages and the regulated inside wage. To test robustness and to examine the potential pathways by which wage regulation may affect school output we augment this with other data on the local labour market and teacher tenure. This section describes our data: details of sources and the years we use are provided in Appendix A, Table A1. School performance data ~~~~~~~~~~~~~~~~~~~~~~~ Our main measure of pupil performance at school level is taken from the Pupil Level Annual School Census (PLASC). This dataset records the performance of all pupils in national exams in all 3285 public (state) secondary schools in England. PLASC began in 2002. We use data from the inception of PLASC to 2007. We confine ourselves to this window to avoid any potential contamination from schools which operated as academies (there were only 85 in 2007).8 Pupils can take up any number of KS4 exams, with a minimum of one and a conventional maximum of around fourteen, in a range of subjects including Mathematics, English language, English literature, Science subjects, and History.9 KS4 exams are graded from A* to G, and these grades are translated into points, such that an A* is worth eight points, an A is worth seven, a B is worth six, and so on. We use the total number of points that a student obtains (i.e. points from each exam summed across all the exams they take), averaged at school level, as our key measure of output of the school. To control for the effect of pupil type (including effects of peers in the school), we control for the attainment of the same cohort of pupils immediately before they entered the secondary school.10 This is the pupils' average point score in the KS2 exams, also from PLASC. KS2 exams are graded from 2 to 5, and are taken in Mathematics, English and Science.11 In analyses of other measures of performance we examine the proportion of pupils in the school who achieved 5 GSCEs at grades A*-C or better; the average number of KS4 exams taken by pupils (all from the PLASC dataset); and the number of pupils excluded from school (from the Dept. for Education). We also use measures of school performance derived from the in-depth inspections of schools undertaken by the government schools regulator (OFSTED). OFSTED undertakes expert in-depth inspection of each school to provide a published assessment of the quality of each school. These assessments take place over a number of days and are intended to ‘net out’ any pupil or parental effect. The school's performance is rated on numerous dimensions, including an overall rating of the school and an assessment of teaching and learning at the school.12 Both are rated on a four-point scale from 1 (outstanding) to 4 (inadequate). Given the frequency of inspections, these data are available for only a sub-sample of the schools in our main analysis. We use inspection data from 2002 to 2008. In our analysis of pathways, we use data on teacher tenure at school level from the School Workforce Census. This is a recently released administrative dataset from which the proportion of teachers in a school who are in post for less than a year, and the proportion that have been in post for over 10 years can be extracted. It has only, to date, been released for 2010. Wages and employment data ~~~~~~~~~~~~~~~~~~~~~~~~~ Our key measure is the ‘wage gap’, specifically the difference between ‘outside wages’ and ‘inside wages’. Our ‘outside wage’ measure is intended to measure the alternative private sector wage which teachers could command. We define the outside wage for each school as the average wage of all Local Authorities (LA) whose headquarters is within a 30 km radius of the school. This circle around the school represents a ‘travel to work’ area (TTWA), in which teachers at the school could seek alternative employment.13 In some areas, there are as many as 45 LAs within this radius, whilst in many others there is just one as LAs differ in geographic size. For schools where there is not a headquarters of a LA within the 30 km radius (or the wage data is missing for the LAs within that range) the nearest LA with wage data is allocated to a school, provided that LA is within 60 km of the school. If the nearest LA with wage data is outside that distance, the school is excluded from our analysis.14 The local wage data are from the Annual Survey of Hours and Earnings (ASHE) dataset, a 1% sample of all employees in Great Britain, covering approximately 300,000 workers per year, sampled in April of each year which provides wage data at Local Authority level. The publicly available dataset contains average wages, split into manual and non-manual, part and full time. As our measure of outside wages we use (the log of) the full time male non-manual hourly earnings.15 To construct the lagged five-year average wage for each school we use the average of outside wages from t-5 to t-1 for all LAs that fall into the TTWA of the school. The five year gaps between our outcome observations in Eq. (2) mean we use average wages from 1997 to 2001 and 2002–2006. In robustness checks we use a more complex measure of outside wages which corrects wages for local labour market composition. The correction uses Labour Force Survey (LFS) as ASHE does not contain the necessary data to make this correction.16 Details are provided in Appendix B. Inside wages are from the Department for Education Teacher Pay and Conditions Handbooks. For our ‘inside wages’, we use teacher wage data at the payband level to avoid concerns over endogeneity.17 We use the (log of the) inside wage lagged for one year, so inside wages are for 2006 and 2001. The results are robust to using contemporaneous inside wages (for 2007 and 2002). In analyses of heterogeneity we define large ‘outside wage regions’ on the basis of the long run level of outside wages, as measured in ASHE. We group the 10 Government Office Regions (GORs) in England into three (details in Appendix Table A2, Panel (A)) and assign each school to a single ‘outside wage region’.18 In extensions to our main analyses, we examine the impact of outside wages and employment prospects for different groups (the youth labour market, manual workers) in the outside labour market on school performance. These wage and employment data are from ASHE. Controls ~~~~~~~~ The PLASC data contain information on the final-year (age 16) composition of pupils. But our identification approach means that results with and without controls should give the same estimates and pupil characteristics may be endogenous. So we initially present results with no controls other than prior attainment. We then use a small set of controls (those that we find are associated with value added controlling for outside wages) to allow for the effect on the standard errors of heterogeneity across pupil type. The controls we include are Key Stage 2 (KS2) scores, percentage of male students, percentage of students eligible for free school meals (FSM), and the percentage of severe special educational needs students (SEN) We also use this set of pupil controls to test the key identification assumption in our design. We also include school expenditure per pupil (EPP) to address the concern that wages might differ between schools across different paybands, but that school resources might not.19 Data description ~~~~~~~~~~~~~~~~ Summary statistics for all variables we use are in Table A3. The table shows the range of KS4 points across school is large, with a minimum score of zero and a maximum of 99. The mean is just under 44. KS2 scores are in different units to KS4 and have a mean of just under 27. The log of the outside wage, averaged for each school over a five-year period, has a school level mean of 7.134, which equates to an average salary of £29,524 per annum.20 The log of the inside wage has a mean of 6.43 and on average the wage gap is 0.704. All variables, particularly the wage gap, exhibit considerable within group variation, indicating change over time within schools. Table A2, Panel (B) shows the absolute growth in the (nominal) level of wages by the three outside wage regions for the whole period covered by our data, 1997–2006. For the whole period, absolute growth is highest in the high outside region at £12.2 K and lowest in the low wage region at £8.4 K. In the first sub-period, the growth is again monotonic across the three outside wage regions. In the second sub-period, the growth is highest again in the high outside wage region and very similar in the medium and low wage regions. Given this, there is a close mapping between the definition of broad outside wage region in terms of long run levels of wages and a definition of outside wage areas in terms of growth rates in wages.21 Table 1 presents cross sectional estimates year by year between the (log of) the wage gap (lagged one year) and school average GCSE performance. We include controls for mean KS2 scores at intake and local authority fixed effects. The estimates are the change in average GCSE points per pupil associated with a 10% increase in the outside wage. All the cross- sectional associations are negative, though not all are statistically significant. The association is largest in 2005 at just over half a GCSE point lost per student. Test of assumption that school fixed effects can be differenced out2222We are grateful to an anonymous referee for making this issue clear and for suggesting tests to address this. This section draws heavily on their help. ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The model we estimate is a model of value added i.e. we condition on the prior attainment of the pupils who are in the school at age 16. We assume that the school fixed effects in Eq. (2) are fixed over time and so can be differenced out. This assumption may be violated for several reasons. The most important of these is that the cohort of students whose KS4 achievement is measured in 2007 is different from the cohort whose KS4 achievement is measured in 2002. This is particularly a problem in our context if within school differences in cohorts are correlated with changes in wages and different types of student different in expected growth in achievement. We seek to overcome this, in part at least, by defining the outside wage as covering a large area compared to a school. The travel to work area in our analysis is 30 km round the school. Thus the wage we use is not the wage simply of workers who live locally to the school. This means that the outside wage is less likely to be a measure of the wage of parents of the children in the school. We also control for a number of observed characteristics of the children in the each of the two cohorts. However, we may be concerned that our design is not robust and here we implement a direct test of our assumption that changes in the outside are not correlated with observable changes in the school cohorts. We regress, at school level, the initial level of the achievement of the school cohort KS2i,t-5 plus the other covariates pertaining to the school cohort, on the outside wage, controlling for school fixed effects and time dummies.23 We present results for all years 2002–2007 and also for 2002 and 2007 only. In our main estimates of Eq. (2), we use two observations per school of value added. The dates of these are 2007 and 2002 and therefore the school level covariates we use in our regression are 2006 and 2001. The results are in Table A4. The table presents the estimates of KS2i,t-5 and the other student characteristics we use in the main results on the outside wage along with the joint F-test of all the covariates. Column (1) presents results for the whole period and column (2) for 2002 and 2007. The results show that the F-tests on the covariates in both columns are small and insignificant. In addition, none of the student characteristics are significant associated with the outside wage across the two regressions. It is possible that areas with high wage changes experience differential changes in school composition over the period than areas that had lower wage changes. To examine this we repeat the analysis above, but this time splitting the sample split into three groups according to the growth in outside wages over the period.24 This split into three regions according to wage growth rates has very close correspondence to a split in terms of long run levels of wages. The results are presented in columns 3–8 of Table A4. This shows that there is a little more association within ‘growth in outside wages’ regions between the pupil characteristics and the level of wages, but even here only two of the F-tests are statistically significant at the 5% level. Further, the patterns of association between changes in pupil characteristics and changes in wages is not consistent across, or within, regions. For the highest ‘growth in outside wages’ region, wage growth is associated with an increase in the proportion of children with parental low incomes, as measured by eligibility for free school meals. In the lowest ‘growth in outside wage’ region, an increase in wages is associated with a fall of pupils with special needs, which might also be taken as a marker of low parental income. But in this region a growth in wages is also associated with a fall in performance in the exams taken immediately prior to school entry (Key Stage 2).25 Tests of common trends ~~~~~~~~~~~~~~~~~~~~~~ Our analysis is essentially a difference-in-difference (DiD) estimation in which we have multiple areas and two years per school. To check our common trends assumptions we regress the growth in outside wages between 2002 and 2006 at the LEA level on the average characteristics of school pupils (jointly) in 2002 and the level of total KS4 points per pupil in 2002 (the earliest year for which these data are available). Differences in outside wage growth that are associated with the initial level of covariates or school output could indicate that areas that differ in terms of wage growth also differ in terms of unobservables and threaten our identification strategy. The results are presented in Appendix A, Table A5 columns [1] and [2]. The results show that there are no significant associations between the 2002 characteristics and the subsequent wage growth at LEA level. While the baseline school covariates and school level average GSCE points are not available before the release of PLASC in 2002, there is available data on the performance at school level on the 5 A*-C metric. As a further test, we examine the association between this metric in 1997 with subsequent wage growth 1997–2006. This therefore spans the full period covered by our wage data (as our analysis uses wages averaged over 5 years, lagged once). The results are presented in Appendix A, Table A4, column [3] and show no association between baseline performance and wage growth. The subsequent columns of Table A4 repeat these analyses at the level of the three outside wage regions. Again, we find no association between baseline covariates or baseline school performance and subsequent wage growth. We conclude that our design assumptions are likely to be satisfied. Baseline results ~~~~~~~~~~~~~~~~ Table 2 presents estimates of the effect of wages on school productivity as measured by value added on KS4 test scores. Standard errors are robust, clustered at the TTWA level and bootstrapped to allow for the estimation of outside wages.26 Columns [1] – [4] presents the results for value-added (total exam points per pupil controlling for initial intake scores). Columns [5] and [6] present results for the percentage of pupils who achieved at least 5 A*-C GCSEs (also with controls for initial intake scores). Columns [1] and [2] present OLS estimates. The remaining columns present fixed effects estimates. The first of each pair of estimates has no controls and the second includes those pupils and school level covariates which are associated with test scores after controlling for school fixed effects and prior attainment. The table shows all coefficients on the wage gap are negative and significant. The increase in size between the OLS and fixed effects estimates in Columns [1] and [2] indicates that controlling for omitted school level factors is important.27 On the other hand, the effect of controlling for time varying school level covariates is small, supporting the appropriateness of our identification strategy, as if this is correct the covariates should only affect the standard errors. The coefficients represent the estimated change in the outcome associated with a 10% increase in the gap between the five-year outside average wage and the inside wage. Column [4] indicates a loss of approximately 1 GCSE point per pupil in value added in response to a 10% increase in the wage gap, which is equivalent to dropping one GCSE grade in one subject or around a 2% average fall at the mean of 44 points. Column [6] indicates a fall of around 2.5%age points in the proportion of pupils achieving 5 or more A*- Cs, which equates to a 5% fall at the mean of 56.5 percent. Thus the estimates suggest that a positive shock to outside wages leads to a small, but non-trivial, fall in both measures of school productivity. We now subject our results to a battery of tests. We begin by testing the robustness of the results to the definition of our key variable, the outside wage. We then allow for correlation of errors across schools within labour market areas. We then test the salience of our design by examining whether the effects of regulation are larger in various settings in which the wage ceiling should have greater bite. Finally we examine robustness of the results to potential gaming by schools that may be correlated with the outside labour market shocks. Alternative specification of the outside labour market ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our primary specification uses the average of the one year lagged gap between the (five year averaged) non-manual male outside wage for a TTWA defined as 30 km round the school and the regulated (inside) wage. We subject these key measures to a number of robustness tests. First, in the spirit of a placebo test, we check that there is less response to a less relevant measure of outside wage. As teachers are graduates, if the outside wage has an effect on their performance, they should be less likely to respond to shocks to the wages of less skilled workers. In row [2] of Table 3 we replace the non-manual wage with the wages of less skilled workers (manual workers). We find the coefficient on the wage gap more or less halves and is insignificant at conventional levels. Second, we check robustness to the definition of the TTWA. Our main specification uses a radius of 30 km round each school to define each school's unique TTWA. We change this radius in 10 km steps from a minimum of 10 km to a maximum of 120 km. The estimated wage coefficient for each radius and the associated 95% confidence intervals are plotted in Fig. 2. This figure clearly shows the results are insensitive to the precise choices of the radius for distances between 20 and 60 km. Larger areas cannot really be considered to be a TTWA for a school and we also find no effect at the very small radius of 10 km. It is possible at this small radius wages are endogenous; we return to this below in Section 5.3.28 Third, outside employment prospects may matter as well as wages. At the very least, if employment falls, average wages rise due to composition effects. To check whether this impacts on our results we add the employment rates of 25–49 year old at the local authority level over the 5-year period for which the cohort is in school as an additional control. Row [3] shows that the coefficient on the wage gap falls a little but remains significant at the 10% level. The coefficient on the employment rate is insignificant. Fourth, our model is driven by the difference between inside and outside wages. The inside wage at school level may be endogenous if schools try to circumvent wage setting. To test this we omit the regulated wage and estimate the effect of only the outside wage. The results in row [4] indicate that the coefficient on the outside wage is actually very little different from that on the wage gap in our baseline estimates in row [1], though inclusion of the regulated wage improves the precision of our estimates. This suggests that the average school was not able to the circumvent wage regulation.29 Fifth, we use a more sophisticated measure of outside wages that creates an area- and time-specific outside wage for teachers using the observed characteristics (age, gender, years of schooling, etc.) of teachers in a particular area-year cell (we do not observe these characteristics at the school level).30 The results in Row [5] show that this has little effect on the estimated wage coefficient. Examination of the adjusted wage series shows that the difference in characteristics between teachers and those working in other sectors does not vary greatly over time in an area (available from authors). Thus the main cause of area- specific time-series changes in the wage gap is simply the growth in the non-manual wage in an area rather than changes in observables or the price of these observables over time. This supports our use of a measure that does not adjust for composition. Sixth, there may be common unobserved shocks to wages at the local geographical level (the TTWA). To allow for this, we use spatially correlated standard errors as in Conley (1999). Row [6] shows the estimates remain significant at the 5% level. Finally, wage changes in an area may lead to school composition changes, both at intake at age 11 and during the 5 years in the run-up to Key stage 4 exams if children change school between the ages of 11 and 16. We go some way to picking this up by controlling for KS2 and other characteristics of the school cohort who sit the KS4 exams. But there may remain unobservable time-varying changes in the school cohort which we cannot control for with school fixed effects (as they are not time varying) or may not control for with the cohort level time-varying measures of initial ability (KS2 mean) and other pupil characteristics. Estimates without controls for KS2 provide some indication of whether unobservable changes in ability might affect our results. In Row [7] of Table 3 we present results without controls for KS2. These are about 20% higher than our baseline estimate. This suggests that if an increase in wages is negatively correlated with changes in unobserved post-intake ability of the school population, our estimated effect may be a biased upwards. On the other hand, if changes in post-intake unobserved ability are positively correlated with shocks to wages, then our estimates will be biased towards zero. Arguments can be made for both a positive correlation (parents can buy more goods to complement schooling) and a negative one (parents who have an income shock substitute towards paid work and have less time to supervise their children). Tests of the salience of our research design ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our argument is that pay regulation acts as a ceiling and we exploit this to identify the impact of pay on school productivity. However, we do not observe what the unregulated wage for teachers would be. While it is difficult to estimate the exact counterfactual wage for a teacher in the outside labour market (see Ma et al., 2009, for one approach) even without observing this counterfactual we should expect to find more effect where the ceiling is more likely to bite. This is in labour markets where outside wages are highest and where schools have least power over wage setting and conditions of employment. To compare the effect of an outside wage shock across heterogeneous outside labour markets we estimate the response to an outside wage shock separately for schools in each of the three ‘outside wage regions’. The results presented in Table 4 show a monotonic relationship between the wage gap and reduction in value added. Column [1] of Table 4 shows a negative interaction term for the highest wage region. Column [2] shows a positive interaction term for the lowest wage region. Column [3] presents estimates for each outside wage region separately. These confirm the results of the previous two columns: the effect of pay is to reduce school productivity most where the pay regulation bites hardest and to not affect it where the wage gap is small.31 These results fit with the pattern in regional wage changes that were discussed in Section 3.4. While London and the South East may have the highest long run level of outside wages, our estimates recover a response to shocks to outside wages. These have been largest in the high ‘outside wage’ region but they also have arisen in several labour markets located in the middle ‘outside wage’ region. For example, of the LAs in the top quartile of highest wage growth in our period, 35% are in the middle ‘outside wage’ region. For these LAs our estimates of the counterfactual wage (what would be paid in the absence of regulation) is £1750 higher than the regulated wage.32 Thus wage regulation bites in the middle, as well as the highest wage area, and we find a response to wages in both.33 We also check that our results are not driven by London schools. Column [4] of Table 4 omits all London schools and shows are results are not ‘a London effect’. The point estimate of a wage shock for such schools is around 30% higher, though the difference with the full sample is not statistically significant. We would expect a larger effect for schools which have least power over their terms and conditions of employment. Other schools might try to circumvent the regulation in a way that would lead to a smaller impact of the wage gap for such schools. In England, secondary schools are classified into a number of types, the most common being Community Schools. These schools are not permitted to select pupils and LEAs have control over their curriculum and teacher wages. In the other types of state schools pupil selection is sometimes an option (for example, publicly run religious schools) and there is, in theory at least, more flexibility in terms of teacher wage setting. In columns [5] and [6] we present the estimates for Community Schools only. Column [5] is for KS4 results and column [6] for the 5A*-C proportion. The magnitude of the coefficient increases in each case by around 30% compared to the estimates for all schools. These results provide support for our argument: there is a stronger effect amongst schools with the least (no) power over their wage setting which is where the gap will bite hardest.34 School gaming ~~~~~~~~~~~~~ KS4 exams are high stake exams, not just for children, but also for schools. This may lead to school gaming and if this is associated with higher wage growth it could bias our results. By looking at both a measure of average performance (adjusted for intake) and performance on the key 5 A*-C metric, which matters most for lower ability students, we look at performance at different parts of the ability distribution. So if gaming only affects the performance of low or high ability students, our approach should deal with this. But to further investigate this we first consider whether schools subject to wage shocks try to prevent children from sitting exams by excluding them. We can examine this only at regional (the 10 GORs) level due to lack of published data on exclusions at the school or LEA level.35 The results of a regression of wage shocks on exclusions at the regional level are presented in Table 3, row [8]. This shows no relationship between the outside wage and the number of exclusions, suggesting schools do not react to wage shocks by baring pupils from exams. Second, it is possible that schools react to outside wage pressure by limiting the number of exams that pupils take in order to get better average performance. To examine this we re-estimate our baseline model using the average number of exams taken as the dependent variable. The results in row [9] show that schools do respond at this margin: a 10% increase in the outside wage causes schools to reduce the average number of exams taken by each pupil by just over 0.16. They may be doing this to hit the key 5 A*-C metric. However, as Table 2 shows, this strategy does not appear to prevent them from having lower performance on this metric. But it may mean that our results are under-estimates of the impact of the outside wage: schools subject to shocks may divert more of their effort to hitting the targets at the expenses of higher scores above the target.36 We conclude that while wage shocks may induce schools to game, gaming does not drive our results. Schools subject to wage shocks do not exclude pupils to avoid poor performance and while they reduce the number of exams their pupils take, wage shocks still negatively affect performance at both the high end (students who take many exams) and at the lower end (those students aiming for the 5 A*-C minimum) of the ability distribution. On the basis of this battery of tests, we conclude that our identification strategy is robust. School performance is decreased by positive shocks to wages in the local labour market, there is an effect for both high and low ability pupils, and the effect is stronger where the wage ceiling has more bite.","Our argument is that shocks to outside wages can drive teacher effort (through an efficiency wage effort) and labour supply (the loss of good teachers) and that this lowers the quality of teaching. But it is possible that the results are not due to responses of teachers but are driven by responses of pupils and their parents to outside market conditions. While we cannot rule this out completely, we provide descriptive (as it is mainly cross-sectional) evidence on this by examining a measure of teacher quality that should not be affected by pupil type, so shutting down a pupil effect; we directly examine the relationship between wages and teacher tenure to see if schools subject to wage shocks experience problems with teacher retention; and we examine whether the plausible responses of pupils to outside labour market conditions drive our results. Finally, we discuss evidence relevant to parental behaviour. The effect of wage regulation on quality of teaching ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ One way to examine the effect between wages and quality of teaching would be to examine the qualifications of teachers. However, teacher qualifications are not likely to be a useful margin. First, research on teacher qualifications suggests there is little correlation between teacher qualifications and teacher effectiveness (e.g. Rockoff, 2004, Rivkin et al., 2005, Aaronson et al., 2007, for the USA and Slater et al., 2012, for England). Second, in the LFS data, most teachers are graduates and this figure does not vary systematically across regions. Instead we examine a direct measure of the effectiveness of teaching in the school: the ratings the school received in their OFSTED inspection of school quality between 2003 and 2008. This measure should be less contaminated by unobservable (to us) attributes of the student body as the regulator ratings are based on in-depth inspections of the school carried out by experienced educators. School performance is rated on a number of dimensions, each using a four-point scale which ranges from 1 (outstanding) to 4 (inadequate).37 In Table 5, column [1] we examine the relationship between overall performance of the school and outside wages and in column [2] we examine the relationship between the ‘quality of teaching’ score and outside wages.38 We use the same lagged five-year average definition of outside wages and include pupil characteristics for the year of observation as in our baseline model. As there is only one observation per school we do not include school fixed effects but instead include LEA fixed effects. The results show that a larger wage gap is associated with a poorer overall rating of the school (a positive coefficient indicates an increase in wages is associated with poorer performance). More importantly, it is also associated with poorer teaching quality. A 10% increase in the local labour market wage decreases the quality of teaching by 1.4 points.39 This is a large effect: the mean teaching quality score is 2.6 so this is a fall in the teaching quality score of over 50% at the mean. The effect of outside wages on teacher tenure ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In our model, one effect of wage regulation is that teachers leave schools in high wage areas. This is likely to affect student performance through a variety of routes: time will need to be spent by the senior management team on recruitment rather than other activities, less experienced teachers may be less effective (Dolton and Newson, 2003, provide UK evidence on this) and there may be spillovers on the morale of the remaining teachers that may dampen effort (Ronfeldt et al., 2011). To look at this channel (which is essentially selection) we examine the relationship between outside wages and teacher tenure in each school.40 We exploit a recently released data set, which presently covers only 2010. We cannot, using these data, examine hires and separations, but we can examine the association between average tenure in schools in time bands and outside wages. We present the association between the proportion of the teachers who have very short tenure (less than a year in the school) and long tenure (over 10 years in the school) and lagged outside wages, averaged over five years to replicate the modelling of wages in our baseline specification.41 We only have one observation per school so estimate Eq. (2) using LEA rather than school fixed effects, and report the results in Table 5, columns [3] and [4]. Column [3] presents estimates of wages on short tenure while column [4] examines long tenure. The results show that where outside wages are high, schools have a higher proportion of teachers who have been in tenure less than one year and a lower proportion of teachers who have been in tenure for ten years only. The wage effect is little altered by controls for pupil type.42 A £1000 increase in the outside wage results in a 0.25 percentage point increase in the proportion of teachers who have been in the school less than a year, and a 0.43 percentage point decrease in the proportion of teachers who have been in the school more than 10 years. While the magnitude of these effects is not large (the coefficients represent a 2% change for both short and long tenure), schools in high wage areas lose teachers faster and also have less experienced teachers. A pupil or parental effect? ~~~~~~~~~~~~~~~~~~~~~~~~~~~ An alternative hypothesis is that our results are due to the responses of pupils and/or their parents to outside wages. The relationship we find could be driven by pupils responding to better labour market opportunities by decreasing their effort at school because they know there is an employment alternative. If this is the case, we would expect to find a negative relationship between school performance and higher outside wages and/or the demand for youth labour in the local labour market. To examine this we estimate the relationship between school performance and the demand for youth labour, as measured by local authority wages and, separately, the unemployment rate, of 16–25 year olds, lagged one year.43 Regressions are at school level and include the same controls as our baseline specification and school fixed effects. Table 5, column [5], presents the results for the association of school performance and youth wage rates. Column [6] presents the association with youth unemployment. Neither association is large or statistically significant. Further, for the negative relationship we find between outside wages and school performance to be driven by a pupil response, pupils would have to be responding negatively to positive outside wage or employment shocks. Whilst this is plausible, it seems equally plausible that at least some pupils respond to positive wage shocks by putting in more effort at school, on the grounds that if they get better exam grades they are more likely to get a (better) job. These plausibly heterogeneous responses do not fit with our finding of a negative outside wage effect for both low ability students (those at the 5 A*-C margin) and all students (total exam points achieved). An outside wage shock would also be a positive shock to parental income, which could result in worse performance if greater parental income means less supervision of children or more leisure time for children. While we cannot rule this out, the lack of significance of the full time male manual wages in the robustness checks in Section 4 brings into doubt whether the effect is working through parental income. If parental income was important, it is not clear why a parental income effect should operate only for parents in non-manual occupations. In addition, the large (though often correlational) literature shows a positive rather than a negative association between parental income and child attainment. It thus seems less likely that the negative relationship we find between outside wages and child attainment is driven by a parental effect. School performance could affect outside wages, which would bias our results. The most obvious mechanism by which school performance may affect outside wages is through sorting: good schools attract high income parents to move into the area surrounding a school.44 This would give a positive shock to the average outside wage and would bias our estimated coefficients upwards. However, the 30 km radius TTWA we use weakens this argument. If we had used the catchment area of a school to determine the outside wage this would be problematic, as parents try to buy houses in the catchment areas of ‘good’ schools. But much of that gaming is within area. Individuals are likely to choose areas based on their job and general lifestyle choice and then select their specific within-area locations based on the schools available. Fig. 2 shows a smaller relationship between outside wages and school performance at radii of 10 km and at 20 km. The TTWA radii at these distances give more weight to the local catchment area round each school, which may indicate endogeneity at this smaller spatial distance. Our analysis uses 30 km distance to avoid this problem.45 In summary, whilst these results are primarily descriptive as data limitations mean they rely on more restrictive assumptions than our main analyses, they suggest that the effect of shocks in outside wage on school performance is, at least in part, through lower teaching quality and labour supply and not from responses of parents and children to the local labour market. In fact, whilst pupils (and their parents) might respond, they probably do so in a way which biases our estimated coefficients towards zero.","This paper exploits the national regulation of teacher wages, national exams at entry into and exit from secondary (middle/high) schooling and a national school inspection system in England to estimate the effect of teacher pay on school productivity. We find that a larger gap between regulated pay and the outside labour market remuneration reduces school performance as measured by student performance in key exams and that the effect is larger where the ceiling imposed by regulation bites harder and for schools that have no control over pay and conditions at school level. At the average a 10% increase in the local labour market wage would result in an average increase of 2% in the scores attained in the high stake exams taken by pupils at the end of compulsory schooling in England. But the effect in areas or schools where the ceiling bites harder is around 30% higher.46 Lazear (2000) emphasises that incentives can affect performance through both effort and sorting. The national set up of the wage regulation in England means that both channels are likely to operate. Wage regulation which keeps teacher relative wages low in one (large) area of the country and high in other (large) areas will encourage both effort reduction and mobility of teachers to area where they get better relative remuneration for the same job. Data constraints mean that we cannot trace through all the pathways through which the pay effect operates but it seems likely that both channels operate. We have shown that school performance and direct measures of the quality, which are important to schools under the ‘name and shame’ rating system used at national level in England, are lower in schools where regulation bites harder. This may reflect reduced effort of the teachers in the schools but we also show that schools subject to high outside wages relative to regulated pay also have higher staff turnover, which may reflect movement of teachers away from these areas. Our findings support the view that teacher pay is important for school performance. The recent focus of many governments has been on using pay for performance for teachers (e.g. Lavy, 2009). However, centralised pay setting affects teachers in many more countries than are using pay for performance in the classroom. Our findings suggest that policy effort could be usefully directed towards increasing flexibility in these centralised wage setting processes."],["One of the most important developments in international finance and resource economics in the past twenty years is the rapid and widespread emergence of the $6 trillion sovereign wealth fund industry. Oil exporters typically ignore below-ground assets when allocating these funds, and ignore above-ground assets when extracting oil. We present a unified stylized framework for considering both. Subsoil oil should alter a fund's portfolio through additional leverage and hedging. First-best spending should be a share of total wealth, and any unhedgeable volatility must be managed by precautionary savings. If oil prices are pro-cyclical, oil should be extracted faster than the Hotelling rule to generate a risk premium on oil wealth. Finally, we discuss how our analysis could improve the management of Norway's fund in practice. --------------------------------------------------------------------------------","Since 1994 the number of sovereign wealth funds has nearly quadrupled to 73 (SWF institute, 2013). These funds hold some of the largest portfolios in the world and globally account for over $6 trillion in assets (ibid.). Two thirds of the sovereign wealth fund industry (by size) has been funded by selling below-ground assets such as oil, natural gas, copper and diamonds (“oil” for short). These funds often comprise a large part of commodity exporters’ wealth. Azerbaijan’s US$ 34 billion fund accounts for almost half its GDP, Qatar’s US$ 170 billion fund accounts for almost two thirds of GDP, Saudi Arabia’s US$ 740 billion funds are approximately four-fifths of GDP, Norway’s US$ 840 billion fund is nearly one and a half times GDP, and the United Arab Emirates’ US$ 1 trillion funds are over two and a half times its GDP (SWF Institute, 2013; IMF, 2013). The purpose of these funds is to smooth consumption of oil income: across generations because oil reserves are finite, and between periods because oil and asset prices are volatile. While such funds are professionally managed and often allocate their assets using modern portfolio theory, we argue that their investment strategies do not take due account of oil price volatility and subsoil reserves. Similarly, existing theories of optimal oil extraction do not take into account volatile financial markets. These are important issues for resource exporters, since commodity prices are notoriously volatile and below-ground assets can be worth much more than the above-ground fund. Our aim is therefore to answer four questions about how below-ground resources should influence above-ground portfolios, and vice-versa. Firstly, how should one allocate above-ground assets given a volatile stock of below-ground assets? Secondly, how quickly should financial and oil wealth be consumed? Thirdly, how does this change if financial markets are incomplete, so that oil shocks cannot be completely hedged in the portfolio? Finally, how should the optimal extraction rate of below-ground assets be affected by risky above-ground assets? We will show that policy-makers should adjust their above-ground portfolios to accommodate the volatility and erosion of below-ground oil stocks (hedging and leverage effects respectively); consume a fixed share of total wealth; manage shocks that cannot be hedged with precautionary savings; and, if the marginal rent from extracting an additional barrel of oil, namely the oil price minus marginal extraction costs, co-varies positively with average equity market returns, then oil should be extracted faster. Our analysis combines three large and previously unrelated strands of literature. Firstly, the allocation of financial assets is described by CAPM equations modified for subsoil oil wealth. This extends the continuous-time analysis of optimal consumption-saving and portfolio choice (Merton, 1990).1 Secondly, consumption is described by a stochastic Euler equation,2 extending the literature on prudence and precautionary savings to the case when both financial assets and oil extraction can be chosen.3 Thirdly, the optimal rate of oil extraction is described by a stochastic Hotelling rule modified if the proceeds of extraction of below-ground wealth are invested in a risky above-ground financial portfolio.4 Our intended contribution is to introduce a stylized framework that combines canonical insights from all three of these fields. These insights would be modified by including transaction costs and illiquidity premiums, which would help to explain why in practice fund managers do not adjust their portfolios too frequently by introducing some mean reversion into the portfolio decisions (Constantinides, 1986; Acharya and Pedersen, 2005; Garleanu and Pedersen, 2013; Jong and Driessen, 2015). This paper is laid out as follows. Section 2 introduces our model for portfolio choice, saving and oil revenues. Section 3 shows how to allow for below-ground oil wealth with a predetermined path for oil production when the oil price is completely spanned by returns in asset markets. Section 4 deals with the case of investment restrictions which prevent the oil price being fully spanned. Section 5 derives the optimal path for oil extraction. Section 6 discusses the implications of our results and compares these with the policies adopted by the Norwegian fund. Finally, Section 7 concludes and qualifies our results. Additional precautionary saving ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The wealth effect describes the change in the expected return on total wealth from not investing in a particular asset (see (15)). If an asset cannot be held by the fund (cf. asset h in (15)), there is still some exposure to it embodied in the oil price. With complete markets this exposure is offset inside the fund, so the net exposure is a constant share of total wealth. With incomplete markets this net exposure cannot be fully offset and will earn a rate of return, changing the expected return on total wealth. Its importance will diminish as oil reserves are depleted. Stylized illustration of oil-CAPM model ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We now illustrate how a sovereign wealth fund is affected by the presence of subsoil oil, depending on whether or not it has access to hedging assets. We suppose that there is a risk-free asset, r, and two risky assets: 1 uncorrelated with the oil price (the market asset) and 2 perfectly negatively correlated with the oil price (the hedging asset).23 To ensure the latter asset is used for hedging only, we assume it has a zero excess return. This focuses our attention on the precautionary effect (and sets the wealth effect to zero). Fig. 1 first gives the declining expected paths of oil revenues and oil wealth and their 95% confidence bounds. Fig. 3 indicates that the consumption path is smoothed in face of declining and volatile oil revenues and grows in line with total above- and below- ground wealth to reflect precautionary saving. As oil wealth is run down (red dotted line in panel (b)), the fund is built up (blue dotted line) reflecting the basic insight that total wealth should grow at the same constant rate, if the oil price is completely spanned. Investment restrictions: incomplete markets Now consider the situation where the fund is prevented from investing in the risky hedging asset 2 or, equivalently, going short in an asset that correlates positively with the oil price. The dashed lines in Fig. 2 describe the case with investment restrictions, and indicate that the portfolio weight of the uncorrelated asset 1 is unaffected by restrictions on investing in the hedging asset (see (8)). The difference arises merely from the change in the drift of the fund F due to the precautionary effect discussed below. By restricting investment in the hedging asset (or, equivalently, preventing short positions in an asset that is positively correlated with the oil price), there is less need to borrow the safe asset (assuming pure hedging assets with zero excess return as in the numerical illustration, thus avoiding wealth effects). Residual volatility will then be managed by additional precautionary savings. The effect of incomplete markets on consumption is illustrated by Fig. 4 for the case CRP =3 (CRRA = 2). Although not having access to the hedging asset (with zero excess return) does not have a direct effect on the expected evolution of total wealth, it leaves the consumer subject to additional now unhedgeable risk calling for additional precautionary savings. It is clear from panel (a) that initial consumption has to drop in favor of consumption at later times. This effect is larger for larger degrees of prudence, from Eq. (14). Panel (b) shows optimal consumption as a share of total wealth. If oil price risk cannot be hedged due to incomplete markets or investment prohibitions, the share of consumption in total wealth is no longer constant.","The optimal speed of extracting oil may be understood using the Hotelling rule. This states that the expected capital gains from keeping an additional barrel of oil in the ground must equal the return from extracting, selling and earning interest on it (Hotelling, 1931). We now extend this rule for volatile oil and financial asset prices. Optimal rates of oil extraction ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Eq. (20) indicates that the optimal rate of oil extraction is positively correlated with the oil price, so that a sudden jump in the oil price requires a jump in the extraction rate to make the most of it. Oil price shocks affect the rate of extraction most when reserves (and in turn O) are highest, since this is when the majority of oil remains exposed to volatile prices. As the date of exhaustion approaches, the rate of oil extraction gets closer to what it would be without volatile oil and asset prices. Note that the size of the fund does not matter for the optimal rate of oil extraction, only the properties of the assets in the background. Sovereign wealth funds with endogenous rates of oil extraction ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As the asset allocation and consumption problems can be expressed in terms of total wealth (22), propositions 2 and 3 apply. Judicious management of the fund allows consumption to be smoothed in line with the permanent income hypothesis and to buffer consumption from oil price volatility by hedging it with traded financial assets.","The policies of Norway’s Government Pension Fund Global (GPFG) 28 closely follow standard CAPM recommendations ignoring oil wealth. Firstly, the GPFG uses the FTSE Global All Cap Index as the equity benchmark (with around 7,400 individual stocks, a close approximation of the market).29 This is consistent with holding the optimal risky (or market) portfolio in (8) if W = F instead of W = F + V. Secondly, the Ministry of Finance chooses the equity/bond mix, and in 2007 moved from 40/60% to 60/40%, as it was willing to accept more risk for a higher return. This is consistent with choosing the size of the risky portfolio based on preferences and the overall risk and return of the market, as in (9) with W=F. Thirdly, a fixed share of the fund (4% according to Norway’s handlingsregelen) is consumed each year, as in (12) with W=F. GPFG’s management mandate does not mention oil wealth at all (NBIM, 2013), thus leaving Norway exposed to its large and volatile stock of oil wealth: the “elephant in the ground”.30 Norway, and other oil-rich countries with similar funds, would benefit by letting the asset allocation and the consumption rule in the GPFG vary over time. Norway’s asset allocation should vary over time to hedge as much of the volatility of remaining subsoil oil as possible.31 In the first-best case described in Section 3 this would involve taking large long positions in some industries, and large short positions in others (that may exceed the size of the fund), and reversing these positions as oil is extracted. Such highly leveraged positions expose the country to substantial risk if there are systematic shocks (Das and Uppal, 2004). They may also become illiquid, which invalidates the assumption of exogenous prices. Furthermore, the short positions assume that the covariance matrix is stable over time. In practice correlations between oil and each sector vary depending on the type of shock hitting the world economy (Kilian, 2009). As these correlations can only be estimated using past data and the size of the hedging positions are so large, there is the potential for large basis risk between oil and the hedging portfolio. Finally, as oil is extracted the highly leveraged positions must be reversed which will incur substantial transaction costs for a large fund.32 Therefore the target index should not be rebalanced too frequently and portfolios should only be adjusted gradually. A more pragmatic, second-best approach to asset allocation might be to only vary the equity/bonds mix.33 This would be transparent and easy to explain to investors and the public. It would also notrequire short positions, have lower transaction costs, and would not rely on a large, time-varying correlation matrix covering all market assets. In this approach, the only risky asset is the overall equity market (e.g., the FTSE Global All Cap Index). If oil is sufficiently positively correlated with this market, the hedging demand to offset oil risk will exceed the leverage demand.34 In this case, the GPFG should hedge the exposure of subsoil reserves to oil price risk by holding fewer equities and more safe assets while there is oil in the ground. Over time the oil reserves will be depleted and the exposure to equities embodied in subsoil oil will fall. This allows the above ground fund’s equity exposure to rise, so that equities make up a greater share of the portfolio as oil is extracted. The consumption rule should be a constant share of total assets, and thus should fall as a share of the fund as oil is extracted. If oil price risk is perfectly hedged as described in Section 3, this rule should hold exactly. If hedging is imperfect, as would happen by only varying the equity/bond mix, slightly more precautionary savings would be needed. More precautionary saving is also needed if the fund faces a short-sales constraint. Recently, the fund has stopped investing in coal and oil stocks. If the aim is to hedge subsoil oil, it should go further by taking short positions in oil, gas and other stocks that are positively correlated with oil prices. If the aim is to protect the environment, spending should be curtailed to build up a buffer against less diversified risks. In general though, spending as a share of the fund should fall over time as above-ground assets account for an increasing share of total wealth. These recommendations are relevant for the current debate in Norway. The fund excludes investments in certain assets for social and political reasons, such as tobacco and defense firms, and early 2015 also in assets affected by climate change and other environmental concerns such as coal, oil sands, cement and gold mining. In late 2014 Norway also established a government commission to assess its 4% spending rule due to concerns about excessive fiscal stimulus (Ministry of Finance 2014b). This follows declining spending as a share of GPFG assets, from nearly 6% in 2010 to below 3% in 2014, and there have been calls to limit spending to 3% in the future (Olsen, 2014).","Commodity exporters have two major types of national assets: natural resources below the ground and a sovereign wealth fund above it. Although some attempts to hedge commodity price volatility have been made, from long-term forward agreements in iron ore until 2010 to the purchase of oil options by Mexico in 2008, there is no evidence of systematic coordination of below- and above-ground assets. We have made the case for coordinating the management of these two types of asset by integrating the theories of portfolio allocation, precautionary saving, and optimal oil extraction under oil and asset price volatility. Our main findings are as follows. Firstly, commodity exporters should change the allocation of their sovereign wealth fund by leveraging all risky assets and hedging subsoil oil risk. These effects are proportional to the ratio of oil and fund wealth, so unwind as resource reserves are depleted. Secondly, consumption should be a constant share of total oil and fund wealth. This stabilizes the mean and variance of spending as total wealth evolves steadily whilst oil reserves are replaced by financial assets, but relies on the degree to which the oil price can be hedged by components of the above-ground portfolio. Thirdly, if oil wealth cannot be adequately hedged, less should be consumed initially in the interests of precautionary savings in the face of the additional unhedgeable risk that remains. Fourthly, the rate of oil extraction should be faster than predicted by the standard Hotelling rule if oil prices are volatile and positively correlated with financial markets. This generates a risk premium on subsoil oil, as convex extraction costs will fall faster than the rate of extraction. The size of the premium will depend on oil’s correlation with the market, and disappears to zero if their returns are independent. Our analysis attempts to offer a first step towards an integrated approach to managing sovereign wealth funds and natural resources under uncertainty. To do this we combine canonical models of asset allocation, precautionary savings and oil extraction. These models, while widely used and theoretically appealing, have received empirical criticism (Griffin, 1985; Jones, 1990; Fama and French, 2004; Anderson at al., 2014). Future work can address this along three dimensions. The first is to analyze the effect of financial assets on natural resources in more detail, allowing for the exploration and discovery of new reserves35, and extraction decisions at the discrete well level (Kellogg, 2014; Anderson, et al., 2014; Venables, 2014). The second is to extend the analysis to include other non-financial assets such as domestic non-traded capital, human capital and pension liabilities, absorption constraints, general equilibrium effects of spending resource revenues,36 and the benefits from structural reform to make the economy less vulnerable to commodity price volatility. Finally, there is scope for modelling oil and asset prices in more detail. In practice prices exhibit mean reversion (Wachter, 2002), stochastic volatility (Chacko and Viceira, 2005; Fouque et al., 2013), large jumps (Ngwira and Gerrard, 2007) and time-varying correlations (Bollerslev et al., 1988; Longin and Solnik, 1995). Although these extensions allow a better empirical testing of our results, we conjecture that the qualitative nature of our policy insights will be unaffected."],["In many countries large parts of the population do not have access to health insurance. Peru has made an effort to change this in the early 2000s. The institutional setup gives rise to the rare opportunity to study the effects of health insurance coverage exploiting a sharp regression discontinuity design. We find large effects on utilization that are most pronounced for the provision of curative care. Individuals seeing a doctor leads to increased awareness about health problems and generates a potentially desirable form of supplier-induced demand: they decide to pay themselves for services that are in short supply. --------------------------------------------------------------------------------","In developing countries, a large number of individuals is not covered by health insurance (Banerjee et al., 2004; Banerjee and Duflo, 2007). The reasons for this are manifold. On the one hand, individuals are often used to relying on informal forms of risk-sharing instead of being covered by formal health insurance and therefore do not demand insurance.2 On the other hand, it has in the past not been seen as the role of the government to provide health insurance. Moreover, the World Health Organization and the World Bank stress that, even when there is public health insurance, it often does not reach large parts of the population and especially not the poorest families because it is only provided to the minority of employees in the formal sector (WHO, 2010; Hsiao and Shaw, 2007). For instance, until the late 1990s, only 23% of the individuals in Peru had health insurance (CEPLAN, 2011). This may be a cause of concern, because health insurance does not only protect individuals against catastrophically high health expenditures (Wagstaff and Doorslaer, 2003). It also encourages them to see a doctor instead of simply buying medication, and thereby promotes appropriate treatment of illnesses that is often argued to be absent (Commission on Macroeconomics and Health, 2001; International Labour Office et al., 2006). In reaction, many low and middle income countries have recently introduced Social Health Insurance (SHI) targeted to the poor, with the goal to improve their health and also to provide them with protection against the financial consequences of health shocks. Coverage by SHI may or may not be free and typically means that individuals receive medical attention from a service provider. The costs are usually paid out of a designated government budget that is completely or partially funded by taxes. However, to date, it is not well understood through which channels health insurance coverage contributes to the well-being of individuals and how this relates to the incentives provided to health care providers and patients and, more generally, to the institutional environment.3 Important questions in this context are to what extent it is possible to encourage individuals to seek medical attention rather than simply buying medication in a pharmacy, how they can be motivated to invest into preventive care, and what the effects of medical attention are on care consumption and out-of-pocket spending.4 Answering those questions is challenging for at least two reasons. First, we lack detailed data on health care utilization and out-of-pocket expenditures, and second, it is challenging to control for selection into insurance. The second problem means that a regression of utilization or expenditure measures on insurance coverage will yield biased results and will not estimate the causal effects of health insurance. In this paper, we make progress in both directions. We use unusually rich data from the National Household Survey of Peru (“Encuesta Nacional de Hogares”, ENAHO) to evaluate the impact of access to the Peruvian SHI called “Seguro Integral de Salud” (SIS) for individuals outside the formal labor market on a variety of measures for health care utilization and out-of-pocket expenditures. We account for selection by exploiting a sharp regression discontinuity design. The Peruvian case is interesting because SIS resembles Western public health insurance systems and private insurance products in that it covers health care expenditures related to both curative use and preventive care, but does not provide extra incentives to invest in preventive care. Coverage is free for eligible individuals, and those who are not covered by SIS typically lack insurance coverage.5 SIS was created in 2001 and subsequently reformed. Prima facie, these reforms have been successful, as coverage by SIS is comprehensive and the fraction of the population making use of it has increased from 20% in 2006 to 45% of the total population in 2011, reaching a relatively high rate among the SHI programs in low and middle income countries (Acharya et al., 2013). Yet, even though aggregate data suggest that some health outcomes improved since the program has been implemented—between 2000 and 2010 total maternal mortality rates decreased from 185 to 93 per 100,000 children born and child mortality rates decreased from 33 to 17 per 1000 children born6—to date there is little evidence on the effects of insurance coverage that is based on micro level data. A notable exception is the paper by Neelsen and O’Donnell (2017) who study the effects of introducing SIS by means of a differences-in-differences analysis.7 In this paper, we instead study the effects of an upgraded version of the program and make use of the opportunity to control for selection by exploiting the institutional setup in Peru that gives rise to a sharp regression discontinuity design (RDD). It originates in a reform that was agreed upon in 2009. Since the end of 2010, an individual who is not formally employed is eligible for free public health insurance if a welfare index called Household Targeting Index (“Índice de Focalización de Hogares”, IFH) that is calculated by Peruvian authorities from a number of variables is below a specific threshold. We have access to this information and use it to re-calculate the composite index of economic welfare. Variation in this index around the threshold provides the natural experiment that we exploit. Two aspects of the institutional background in combination with our research design are particularly appealing. First, all individuals who are eligible for the program can be considered covered by it. The reason for this is that enrollment is easy and quick, as it can take place at the facility at which individuals seek treatment. It does not involve any fees, and as individuals can usually receive free treatment within a few days, often on the next day. Second, for our population of interest, crossing the eligibility threshold implies coverage. This means that we will estimate the local average treatment effects of insurance coverage for those individuals whose welfare index has a value close to the eligibility threshold. The evidence we provide is policy-relevant, as it addresses the question what would happen to individuals who are just not covered if eligibility would be expanded by increasing the threshold. Making use of the rich data from the ENAHO of Peru and the discontinuity generated by the institutional rules, we find large effects on several measures of curative care use in combination with increases in out-of-pocket spending. Individuals are more likely to receive medicines, it is more likely that a medical analysis is performed, they are more likely to visit a hospital, and it is more likely that they receive surgery. We shed light on the underlying mechanisms by characterizing who pays for each of these forms of care. We show that the increased access to doctors who perform medical analysis and prescribe medicines is usually fully financed by SIS. At the same time, we find that medicines, hospital visits and/or surgeries are financed by households themselves. In line with this, insurance coverage leads to increases in out-of-pocket spending for medicines, hospital visits and/or surgery that are likely driven by limitations faced by the health care suppliers. Using an estimator of quantile treatment effects, we find that the effects on spending are particularly pronounced in the top end of the distribution. We interpret these findings in more detail by looking at them through the lens of a simple conceptual framework that we present in Section 3 below. The main contribution of our paper is that we provide evidence in favor of two arguments that are less common in economics.8 First, insurance coverage leads to increased awareness about health problems, because it increases the likelihood that individuals see a doctor; and second, this even generates a willingness to pay for services that are not covered or not available, which in the context of Peru is a potentially desirable form of supplier-induced demand. There is a huge literature on the effects of health insurance coverage in the developed world. It is beyond the scope of this paper to summarize this literature. Cutler and Zeckhauser (2000) provide an excellent survey. One of the most important general findings is that more generous insurance coverage leads to increases in health care utilization. This has been convincingly documented in the context of the RAND Health Insurance Experiment (Newhouse, 1974, 1993; Aron-Dine et al., 2013) and more recently in the context of the Oregon Health Insurance Experiment (Finkelstein et al., 2012). In comparison, the literature on the effects of SHI in low and middle income countries is scarce, but growing.9 The evidence points towards large effects on utilization. At the same time, the picture is still blurred when it comes to out-of-pocket spending. In part, this is because earlier contributions have not focused on linking evidence on the effects on spending by type of care to the institutional environment. Our data allow us to make progress in this direction and thereby shed light on the underlying pathways. For Peru, Neelsen and O’Donnell (2017) find positive effects of an earlier form of SIS on receipt of ambulatory care and medication, but no impact on inpatient care and average out-of-pocket expenditures. Thornton et al. (2010) find that initial take-up of subsidized, but for-pay insurance “Seguro Facultativo de Salud” among informally employed individuals in Nicaragua was as low as 20%. Moreover, after the subsidy expired most individuals who previously signed up cancelled their insurance. The results for the few who did sign up and kept their insurance suggest that insurance could have a positive effect in the sense that average health care expenditures, which are generally seen as too low, increased. This could, however, also be the case because those who bought insurance and kept it constitute a negative selection of risks for whom the effect of insurance is particularly high. Next to this, there are a number of studies on Mexico, including Barros (2008), King et al. (2009), Sosa-Rubi et al. (2009), and Galárraga et al. (2010). All of them investigate the effects of the “Seguro Popular” program, whose aim is—as the SIS's in Peru—to improve access to health insurance for the poor. Unlike in the Peruvian, but like in the Nicaraguan program, coverage in the Mexican program is not for free. The findings in all four papers consistently suggest that the demand for medical care has shifted to providers that are part of the system, and in line with this, individual health care expenditures have been reduced, including catastrophic health expenditures. In that sense, the program was successful in being a transfer program, but less so in encouraging individuals to seek care when ill. The findings do not suggest that utilization has increased for types of care other than obstetric utilization. Turning to China, Lu (2014) shows supplier-induced demand to be potentially important. She finds that when doctors expect to obtain a proportion of patients' drug expenditures, then they write more expensive prescriptions to insured patients. Wagstaff et al. (2009) find that the launch of a heavily subsidized voluntary health insurance program in the rural parts of the country led to increased outpatient and inpatient utilization, but has not reduced out-of-pocket expenditures. In contrast, Wagstaff (2010) finds for Vietnam that insurance coverage led to a reduction of out-of-pocket spending and no impact on utilization. The design of the program in Georgia is very similar to the one in Peru. However, and in contrast to our findings, Bauhoff et al. (2011) find no effect of insurance coverage on utilization. They argue that the reason for this is that individuals were not aware of being covered, or that there were administrative problems that caused them to indeed not be covered, that they did not make use of the services because the program did not cover drugs, and because the perceived quality of the services was low. Therefore, it is not surprising that their findings are different from ours for Peru. Next, turning to Colombia, and comparing the results to the ones in this paper for Peru, it becomes clear that the effects of insurance coverage depend on the design of the system. In Colombia, private insurers mainly receive a capitation fee and therefore have incentives to increase preventive services on the one hand and to limit total medical expenditures on the other. And indeed, Miller et al. (2013) mainly find effects on preventive care. In Peru, SIS covers both preventive and curative services and doctors are reimbursed on the basis of the treatments they provide. Hence, participating hospitals and health care facilities do not have an incentive to discourage curative treatments or medical procedures in favor of preventive services. This explains why in Peru most of the effects are on curative use. Finally, two recent papers, Gruber et al. (2014) and Limwattananon et al. (2015), investigate the effect of a large-scale increase in health insurance coverage for the poor in Thailand. They find that the program had positive effects on health care utilization, negative effects on out-of-pocket expenditures, and negative effects on child mortality rates. These findings are similar to ours for Peru, except that we find positive effects on health expenditures at the top end of the distribution. Our explanation for this is that individuals, once covered, became aware of additional health care needs and payed for some of them out-of-pocket. We proceed as follows. Section 2 discusses the institutional background and provides details on the SIS program. We present a conceptual framework—an informal sketch of a model of demand for health insurance and health care utilization—in Section 3. In Section 4 we provide information on our data and in Section 5 we describe the econometric approach. Results are presented in Section 6. A number of robustness checks are conducted in Section 7. Section 8 concludes. Additional results are presented in an Online Appendix. Seguro Integral de Salud ~~~~~~~~~~~~~~~~~~~~~~~~ The public health insurance program “Seguro Integral de Salud” (SIS), whose effect we study in this paper, was introduced in 2001. Its overarching goal is to improve access to health care services for individuals who lack health insurance, giving priority to vulnerable groups of the population who live in extreme poverty and are not formally employed (Arróspide et al., 2009). The creation of SIS and subsequent reforms led to a substantial increase of health insurance coverage over time. Bitrán and Asociados (2009) and Francke (2013) provide interesting descriptive analyses of this increase and its relevance within the Peruvian health system in general. Between 2006 and 2011 the fraction of the population making use of services provided by SIS increased from 20 to 45%, which means that by then SIS was the main health insurance provider in Peru.10 Eligibility and benefit package ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The aim of the government was to target poor groups in the population. For this, ideally, eligibility should be based on accurate information on income at the level of the individual or family. However, such information is typically not available in developing countries because a large part of the population works outside the formal sector and therefore does not pay income taxes and social security contributions. Eligibility for SIS is therefore based on the so-called Household Targeting System (“Sistema de Focalización de Hogares”, SISFOH). A unified household registry is maintained and is used to calculate targeting indicators at the level of the family.11 Data are collected by government officials on a continuous basis and using a standardized form. There are questions on, among other things, housing characteristics, asset possessions, human capital endowments and other factors. The IFH index is the main eligibility criterion for the sample of non- formally employed individuals we consider in this paper. It is a linear combination of the variables in the household registry that takes on lower values for households that are poorer. Eligibility for SIS is based on SISFOH in the capital Lima from 2011 on, and in the rest of the country from 2012 on. Online Appendix F explains in detail how the IFH is constructed, including the complete list of variables and their weights. Individuals are eligible if it is below a regional-specific threshold.12 Importantly, whereas potential beneficiaries intuit the importance of their answers to the questions of the government official, they do not know how exactly the IFH index is calculated and what their cutoff value for eligibility is. SISFOH does not inform households about the value of their index and only provides the result of the eligibility evaluation. If eligible, individuals have the possibility to enroll into SIS at a number of places, including the Ministry of Health (“Ministerio de Salud”, MINSA) facilities. They are covered as soon as eligibility is confirmed, which is usually a matter of days, often only one. Then, they receive the health services that are offered at MINSA facilities and that are part of the benefit package. In this sense, eligibility also means coverage.13 In urban areas, the eligibility evaluation is valid for a period of 3 years (4 years in rural ones). This means in practice that re-enrollment after a year is automatic provided that individuals are not covered by another health insurance, individuals do not ask to be un-enrolled in the meantime, an individual changes address, and provided that there is no evidence for fraud. It is not related to the IFH index and in practice, exclusion of individuals after a period of enrollment is very uncommon. SIS offers a comprehensive package of health care benefits, composed of a basic plan of health benefits called PEAS and two supplementary plans. The PEAS plan is based on a wide-ranging list of needs that any public and private insurance plan (including SIS) must address, grouped in the following six categories: healthy population (preventive care), obstetric and gynecological care, as well as care related to pediatric conditions, neoplasm conditions, transmittable conditions and non- transmittable conditions.14 There is an extensive list of benefits related to each listed need. The plan covers ambulatory patient services, hospitalization and emergency care. Table A.1 in Online Appendix B shows that PEAS covers 994 out of the 12,421 possible needs listed in International Statistical Classification of Diseases and Related Health Problems (ICD-10), that is, up to 8.0% of all possible needs.15 It is estimated that PEAS covers 65% of the total disease burden (Francke, 2013). Importantly, for individuals covered by SIS, benefits included in PEAS cannot be subject to exclusions, waiting times or latent periods.16 Moreover, there are no co-payments, coinsurance, deductibles, or similar fees.17 PEAS includes theoretical limits to the number of times an individual can receive each listed benefit. However, these limits are effectively annulled by the two SIS supplementary plans. Individuals enrolled to SIS are automatically covered by these plans, that is, they do not need to sign up to them nor to fulfill new requirements. The Regular Supplementary Plan18 adds 1640 more needs to the PEAS list, that is, an additional 13.2% of needs included in ICD-10 (see Table A.1). This plan includes a monetary limit by event close to US$ 1875. 19 Next to this, the Extraordinary Coverage Supplementary Plan20 is particularly generous as it allows SIS to go beyond the list of needs established by the previous two plans in a discretionary way, and also beyond the established limits. This supplementary plan states a new extremely generous monetary limit: costs for an individual must not exceed 2.5% of the annual SIS budget. There is an application procedure to access benefits using standardized forms. In sum, SIS covers more than 2634 needs (21.2% of needs included in ICD-10) and, through supplementary plans, the stated limits are offset so that SIS offers a very generous package of health care benefits.21 Supply side ~~~~~~~~~~~ The Ministry of Health (MINSA) runs a network of health care centers and hospitals that provides services to individuals covered by SIS. They also serve individuals who are not covered by SIS, who pay for this themselves. Patients usually first visit a health care center and are referred to a hospital when the health care center cannot provide a proper diagnosis or treatment. Health care centers do however not act as gatekeepers. That is, individuals can also directly visit a hospital. Each MINSA facility is reimbursed by SIS on the basis of the treatments it provides. Reimbursement rates are based on estimates of the variable costs plus a markup.22 There is no capitation fee and there are no extra incentives to either limit curative care use or encourage investments in preventive care, as it is the case in Colombia, for example. Importantly for our findings, some MINSA facilities suffered from a number of substantial supply limitations.23 They originate in a cut that SIS experienced in its budget, which in turn resulted in a failure to transfer resources for reimbursement to MINSA facilities during the entire year of 2011. The effects of those supply limitations are not systematically documented, but there is evidence that in response, especially hospitals charged money for medicines and treatments to insured patients.24 To counteract this, “SIS agents” were put in place, who are supposed to ensure that hospitals stop this practice in later years.25 Besides, there has been a shortage of dentists and ophthalmologists. The rate of odontologists per ten thousand inhabitants is one of the lowest among all medical professionals (Giovanella et al., 2012) and it is even lower when they work as providers for SIS (Defensoría del Pueblo, 2013). In addition to that, at that time, only a small number of ophthalmologists provides services to SIS participants, which in turn limits the use of ophthalmological care. Only recently, after our study period, the National Ophthalmological Institute, the largest provider in Peru, joined the list of SIS providers. In sum, even though SIS offers a comprehensive package of health care benefits, supply limitations and institutional problems related to the transfer of resources from SIS to MINSA facilities made some medicines and treatments not fully available and free for the insured. It is an empirical question how this affects care consumption and out-of-pocket expenditures. We turn to this question in our analysis below.","It is instructive to interpret our empirical results through the lens of a conceptual framework. We provide a formal model complementing this framework in Online Appendix A. The primary purpose of this framework is to discuss the implications the institutional setup has on health care demand, with a focus on individuals not being aware of some of their health care needs when they are not covered by health insurance. Without them seeing a doctor individuals are aware of some health care needs, but not all. Sen (2002) distinguishes in this context between “internal” and “external” views of health and stresses that “the patient's internal assessment may be seriously limited by his or her social experience”, such as seeing a doctor or not.26 Importantly, and in contrast to what is common in developed countries, in Peru the individual can buy all drugs at the pharmacy. That is, there are no prescription drugs. Therefore, the baseline case is that she buys drugs at the pharmacy to treat the health care needs she is aware of and pays for this herself. This has been common practice for a long time and may of course ultimately have adverse effects on health. However, evidence on this is scarce (Laing et al., 2001). Suppose now that the individual considers seeking professional care at some cost. Because of the cost, she will only do so if the health care needs she is aware of are important enough to her. She will be more inclined to do so when insured, because she will expect that at least some treatments will be covered by the insurance once she visits a doctor. Suppose she decides to visit the doctor and the treatment is indeed covered by insurance. Then it could be that because of this she will consume more care than she would if she would have to pay for it herself, which could be due to the fact that individuals are liquidity-constrained and health insurance helps them pay for health care, or because of moral hazard.27 The value associated to this generates an additional incentive to see a doctor in the first place. Conversely, one reason not to visit a health care center is the perception that—even though health insurance gives individuals access to doctors—this is not valuable because advice obtained from them is often of low quality, and therefore making use of insurance coverage is not worth its (opportunity) cost, including the time it takes to register.28 An important additional effect is that once individuals visit a doctor, he may make them aware of additional health care needs. This is a form of what has been termed supplier-induced demand (McGuire, 2000). Strauss and Thomas (1998) argue that this is an important potential determinant of health care expenditures in developing countries. The additional care consumption may or may not be provided to them for free, even though there are effectively no formal coverage limits. As we explain in Section 2.3 above, the reason for this is that some facilities suffered from severe supply limitations. This could lead to patients paying for some treatments or buying drugs elsewhere, for instance at a pharmacy. If this is the effect of insurance coverage, then, as long as one thinks of doctors as not providing misinformation to patients, one can make the argument that spending money reveals the preference of the individual for these increased expenditures and that the supplier-induced demand is therefore beneficial to the individual. To summarize the empirical predictions, we expect utilization to increase once individuals are covered by health insurance, and out-of-pocket health care expenditures to either increase (when individuals are made aware of many useful expenditures) or decrease (when the majority of the treatments are provided by the health care facilities and the overall out-of-pocket expenditures decrease because less money is spent at the pharmacy and in health care centers together).","This paper uses cross-sectional data from the ENAHO household survey for the year 2011, which is representative at the level of each of the 24 departments in Peru. It is the only data set that provides the information needed to re-compute the IFH index and, at the same time, on health care utilization, its financing and out-of-pocket expenditures.29 Online Appendix D contains details on the way we define our outcome variables. Data are collected using face-to-face interviews with one or more respondents per household, who are also asked to provide information on the other household members. Online Appendix C describes the interview procedure related to the health questions. In brief, the part of the survey related to health has two branches. In the first branch, individuals are first asked whether they experienced health problems and then what they did in response. In the second branch, individuals are asked which health care services they used and then who paid for it. This means that individuals may be asked twice whether they visited a doctor, for instance. Importantly, the set of outcomes in the first branch is finer than in the second branch, which is why information on the financing source is not available for all variables we use in our analysis. Our data also contain information on the level of out- of-pocket expenditures by financing source. We construct three mutually exclusive categories for the financing source: 1) fully insured, if the individual indicates only a governmental program or other insurance program as the financial source; 2) out-of-pocket, if the individual points only at a member of her household or a member of other household as the financial source; 3) partially insured, if the individual reports the financial source by combining alternatives of the first and second categories. SIS is targeted to individuals who work in the informal sector. For these individuals, the IFH index is the most important criterion to determine eligibility. Therefore, for our analysis, we select individuals that belong to a household in which no member is formally employed.30 This group comprises approximately 60% of the entire sample. In 2011, almost one third of the population lived in the Lima Province and half of Peru's Gross Domestic Product (GDP) was generated there. For two reasons, we focus on individuals from that province. First, in 2011 the IFH targeting rule was only applied in this area, before this was gradually extended to the rest of the country (Ministerio de Salud, 2011). In other parts of the country and in later years it was less strictly applied. Second, the Lima Province is very densely populated and therefore there are enough health care centers and medical professionals so that we can exclude that either a large distance or absence of the staff explain that individuals do not demand health care.31 This means, however, that our results do not necessarily apply to the rest of the country. Our sample contains information on 4161 individuals after the two exclusion criteria are applied. Tables A.4 through A.6 in Online Appendix E provide summary statistics for the full sample. In our main analysis below we use a more local sample to carry out regressions and report our estimate of the baseline level for each outcome along with our estimate of the effect of insurance coverage. As described in Section 5 below, this baseline is the expected outcome for individuals who are just not eligible for SIS. It is a more meaningful statistic than the raw mean in our sample, because it is for the same group of individuals for whom we estimate the effects by means of exploiting the RDD.","In this paper, we estimate the impact of SIS coverage on a host of variables characterizing health care utilization and out-of-pocket expenditures. Based on the institutional setup described in Section 2.2 we do this by means of a RDD using the IFH index as the continuous forcing variable.32 An individual is eligible for public insurance if she lives under poor conditions, which is measured at the household level. In the Lima Province, the condition for this is that the IFH index is below or equal to a value of 55, provided that both, water and electricity expenditures do not exceed 20 and 25 Soles, respectively. Hence, provided that the condition on water and electricity expenditures holds, we have a sharp RDD. The first assumption we need to make for our analysis is that if no insurance or insurance would be assigned to everybody around the threshold, then the respective distribution of the outcome conditional on the index would be smooth in the index zi around zero.34 Then, β2 is indeed the effect of coverage. This assumption cannot be tested directly and is therefore the main assumption we will make. As we have argued before, the institutional rules suggest that it holds, as no other programs or rules are based on this eligibility threshold. Moreover, this assumption is supported by further evidence that we present in Section 7 below.35 The second assumption is that insurance status is monotone in eligibility. This holds by construction, as we are facing a sharp regression discontinuity design and therefore, changing from a value of the index slightly higher than the threshold to a value lower than the threshold will directly make an individual eligible for insurance coverage.36 The final, third assumption is an exclusion restriction. It is that in a small neighborhood around the eligibility threshold, the value of the index, zi, is independent of the outcomes, and in particular εi.37 It would be violated if households would manipulate their answers to the government official in order to influence the value of the IFH index. As discussed in Section 2 this is unlikely to be the case. We nevertheless test for manipulation in Section 7.1. Under the same assumptions, it is also possible to exploit the RDD and estimate quantile treatment effects, as described in Frandsen et al. (2012). The underlying idea is straightforward. Instead of an average, the quantile treatment effect is the change in, say, the median of the distribution of an outcome that results from being covered by public health insurance. Results are presented in Section 6.2. Before presenting the results, it is worth noting that our econometric approach does not involve a “first stage”, as it is usually the case in similar studies exploiting a regression discontinuity design. It is easiest to see this by inspecting the estimation equation above. If individuals are anyway not eligible and hence not covered by health insurance because of their water or electricity consumption, then we will control for this.38 Consequently, β2 is the effect of becoming eligible due to crossing the IFH eligibility threshold for all other individuals. As we have explained above, given the institutional rules eligibility essentially implies coverage. Hence, this is not only the effect of eligibility, but also the effect of coverage. This is as if the first stage is one, which is always the case in a sharp regression discontinuity design (Hahn et al., 2001). Alternatively, as we discuss in Section 7.7 below, our estimates can be interpreted as intent-to-treat effects, or lower bounds of effect sizes. Health care utilization ~~~~~~~~~~~~~~~~~~~~~~~ We start by showing the relationship between the probability to receive curative care and the IFH index in Fig. 1.39 Recall that higher values of the index indicate a higher level of welfare. Individuals are covered by SIS when the index is below the eligibility threshold. In the figure, we plot estimates of the probability to receive curative care against the IFH index minus the eligibility threshold, which is why we expect the downward jump of utilization at zero that we also observe in the figure. The interpretation is that insurance coverage has a positive effect on the probability to consume curative care. Next, Table 1 shows estimates of the effect of SIS on the utilization of a number of health services, including the one in Fig. 1.40 These are local in the sense that we select individuals with an IFH index that is at most 20 points away from the eligibility threshold and, as described in Section 5 above, we control for the value of the index separately to the left and to the right of the eligibility threshold. We also control for age, gender, whether the head of the household is female, the number of household members, and years of education. We also report respective baselines in the second column, which are estimates of the mean outcome conditional on the IFH index being just above the threshold so that individuals are just not covered by health insurance. The last three columns use interactions between the financing source and the outcome variable. Here, we make use of the fact that individuals report on both jointly. For instance, individuals are asked whether they visited a doctor and then, if they say yes, whether the service they received was fully covered by insurance, whether they paid part of it out-of-pocket, or whether they paid everything out-of-pocket. This allows us to for instance construct the joint outcome “went to the doctor and services received were fully insured”. Table A.7 in Online Appendix E shows the respective numbers of instances in our data. We start by looking at more general forms of care, typically provided by easily accessible health care centers. The first row shows that health insurance coverage has a positive effect on the probability of visiting a doctor in the four weeks prior to the interview. It increases by 9 percentage points from a baseline of 30%(significant at the 10% level). This is driven by fully insured doctor visits (6 percentage points at the 5% level of significance), while the effect on doctor visits that individuals have to at least partially pay for is not significantly different from zero. The second row shows that coverage also increases the probability to receive medicines in the four weeks prior to the interview by 15 percentage points, from a baseline of 42%. In contrast to the effect on doctor visits, most of the effect—10 percentage points—is related to medicines that individuals pay for out-of-pocket. The third row shows that coverage increases the probability that medical analysis is performed in the last four weeks, by 5 percentage points from a baseline that is not significantly different from zero. More than half of this effect—3 percentage points—is explained by fully covered access. As for doctor visits, the effect on the probability of receiving medical analysis and at least partly paying for it is not significantly different from zero. As described in Section 2.3, MINSA health care centers provide only basic services. As for care provided by hospitals, health insurance coverage leads to an 8 percentage points increase in the probability to be hospitalized or to receive surgery, from a baseline of about 4%. The survey does not contain information on the financing source for both separately, but for both together. The last column suggests that households paid at least for part of this themselves. So far, these results suggest that coverage lead to increased access to doctors who perform medical analysis and prescribe drugs. While the increases in doctor visits and medical analysis are usually fully financed by SIS, drugs, hospitalization and/or surgery are at least partly financed by households themselves—even though these can actually be considered covered by insurance (see Section 2.2). The discrepancy can be explained by the supply limitations we describe in Section 2.3. The results presented here shed some more light on the underlying mechanism. They suggest that individuals are not getting all the drugs they need at the MINSA facilities and therefore, they go elsewhere to buy them, for instance at private pharmacies.41 Table 1 also shows that health insurance coverage has no significant effects on utilization of dental and ophthalmological care during the previous three months. This can be explained by yet another supply limitation described in Section 2.3, namely the shortage of dentists and ophthalmologists. Turning to care that is more likely of a preventive nature, we generally find no significant effects. As we explain in Section 2, the system does not provide extra incentives for that. In fact, the only significant effect we find is on the probability to see a doctor to receive information on the prevention of a sickness (preventive campaign), and this effect is actually negative. This could be an indication of moral hazard, in the sense that patients invest less in their health in case they are covered by health insurance. It is interesting to contrast these results to the ones by Miller et al. (2013) for Colombia, where the system provides larger incentives to invest in preventive care, as already discussed in Section 1. And indeed, Miller et al. (2013) find stronger positive effects on preventive care. Overall, the picture that emerges is that the effects of health insurance coverage are positive for forms of care that are of a more general nature and can be provided by MINSA health care centers at relatively low cost, such as doctor visits and medical analysis. We provide evidence that these services are indeed free to the patient. We also find positive effects on receiving medicines, hospitalization and surgery, but here it turns out that individuals pay for these services themselves. This suggests that insurance coverage may have positive effects on out-of-pocket expenditures. In Section 6.2 below we follow up on this by estimating the effects on out-of-pocket spending by type of care. Peru is a country in which poor individuals are accustomed to not receiving any professional diagnosis and where drugs can also be bought in a pharmacy without a prescription. Therefore, taken together, our findings point towards the expansion of the program being a success in the sense that it had a positive effect on health care utilization, even if this is at least partly paid for by the individuals themselves. Expenditures ~~~~~~~~~~~~ We argue in Section 3 that individual out-of-pocket spending could be positively affected by health insurance coverage if receiving medical attention motivates individuals to actually spend more on their health themselves, because they become aware of additional health care needs. This positive effect could originate in, or be reinforced by the supply side limitations described in Section 2.3. We have shown above that health insurance coverage has positive effects on the likelihood that individuals receive medicines that they pay for out-of-pocket, and also on the probability that they visit a hospital and/or surgery is performed and that patients pay for out-of-pocket. Obviously, health insurance coverage could also reduce out-of-pocket spending because individuals do not have to pay for certain treatments anymore, or pay less. So, whether the overall effect is positive or negative is an empirical question. In this section, we characterize the effect of health insurance coverage on the full distribution of out-of-pocket spending and also perform an analysis by spending category. Starting with mean spending, Fig. 2 suggests that total out-of-pocket expenditures actually increased with insurance coverage. In Table 2, we use a variety of outcome measures related to moments of the distribution of health expenditures that have also been used in other studies.42 They either attempt to measure expected health expenditures, their cross-sectional variability—interpreted as health risk—, or the likelihood to incur catastrophic health expenditures.43 Table 2 presents the results. The first dependent variable is an indicator for incurring at least some health expenditures. We find no significant effect on this outcome, suggesting that health expenditures are, if at all, mainly affected at the intensive margin. Turning to the intensive margin, the second outcome is the level of annual health care expenditures. We find that insurance coverage leads to an increase of annual spending by about 282 Soles on average, which corresponds to 102 U.S. dollars, —in line with the idea that individuals are motivated to spend more on their health when using medical services more often (which they do according to Table 1 in the Online Appendix). This is about 1.5% of the average household income among the insured (18,800 Soles according to Table A.4). The effect is not significantly different from zero when we use log expenditures. With our next outcome measure, we examine a possible effect on the variability of medical spending in the cross section. It is the mean absolute deviation of health expenditures, calculated separately by insurance status, and similar to the one used by Miller et al. (2013). Health insurance has no significant effect on it. The fifth and sixth measures are constructed from residuals obtained from a regression of health expenditures on the value of the index, insurance status and the interaction of these two variables. Our aim is to measure the variation of health care expenditures in a different way, and therefore we use the absolute value and square of the residual instead of the more commonly used absolute value of the expenditures and their square, respectively. Effects are significant at the 5 and 10% level, respectively. Looking at the results for the next two outcome measures, we find significant effects on the probability that health expenditures exceed the median or the 75th percentile of the distribution of health expenditures in the entire population. In order to control for possible income differences we also analyze the effect of insurance on the share of annual health expenditures spent out-of-pocket at the individual level, relative to the annual per capita household income. We find significant effects on the expenditure shares and also on the absolute deviation of the share and the absolute value of the residual of the share. For the last two outcomes in the second panel, we calculate the 50th and 75th percentile of the distribution of the share and find that health insurance increases the probability that this share exceeds the 50th and 75th percentile by respectively 14 and 13 percentage points. Finally, we look into whether SIS changes the probability of an individual incurring catastrophic health expenditures. Health expenditures are defined to be catastrophic if the share of the expenditures relative to per capita household income exceeds pre-defined threshold values. We follow Wagstaff and Lindelow (2008) and use the thresholds 5, 10, 15, 20 and 25% for this. We find that the probability that individual health expenditures exceed 5% of the per capita household income increases by 12 percentage points, from a baseline of 20%. We also find highly significant effects for higher cutoffs. Overall, the evidence presented in this section and the previous one suggests that health insurance coverage has positive effects on the level and the variability of out-of-pocket spending and that this is partly driven by supply limitations that led to individuals paying for medicines, hospital visits and/or receiving surgery. Fig. 3 complements this evidence with estimates of the quantiles of the distribution of health expenditures with and without health insurance coverage. As explained in Section 5, also these estimates are obtained exploiting the discontinuity at the eligibility threshold.44 Interestingly, we find that insurance has only a positive effect on the higher end of the distribution. It is remarkable that we never find a significant negative effect on either expected health expenditures or measures of variability or risk of high expenditures. Miller et al. (2013), in contrast, find for Colombia that insurance lowers both mean inpatient medical spending and its variability. Likewise, Limwattananon et al. (2015) find that health insurance coverage leads to a decrease in out-of-pocket spending in Thailand. To see what drives this and whether some types of spending were negatively affected, Table 3 reports estimates using the log of spending plus 1 in different categories as the dependent variable so that the reported effects are (approximately) average percentage changes in spending. Consistently with our results shown in Table 1, we find that spending on medicines increased by approximately 55% on average and that spending on care provided by the hospital and/or receiving a surgery increased by approximately 41%.45 In light of our discussion in Sections 2.2, 2.3 and 6.1 these results suggest that the out-of-pocket expenditures are driven by the fact that individuals did not receive all medicines at the facilities and went elsewhere to buy them, and had to pay for some services at the hospitals even though they were covered by insurance. This seems to be mainly explained by supply limitations, in particular the short provision of medicines and budget constraints faced for some hospitals. Rises in health care expenditures as we find them are usually seen in a critical way, especially if some treatments are formally covered but individuals have to nevertheless pay for them. However, one may question whether this is justified here. On the one hand, this increase in expenditures can be seen as an additional burden to the individuals, possibly also increasing the variability of expenditures in the cross-section. On the other hand, the alternative could also be that individuals are not treated at all because they are not aware of their health care needs. In that sense the increase in spending for medicines as well as hospital care and/or surgery could also be seen as a desirable consequence of insurance coverage, which leads to increased accessibility, and thereby gives individuals the idea of using medical services in response to being insured.46 Moreover, looking at it in yet another way, some treatments are at least partly covered by SIS, or a complement to it such as a doctor visit or a medical analysis is, so that the overall price of being treated is generally lower, which means that the law of demand (that lower prices mean more demand) would also predict an increase in usage. So also in that sense our findings could be less worrying than they may first seem.","In this section, after having presented the main results, we assess whether they are sensitive to the particular specifications we have used, and whether the identifying assumptions we have made can be supported by additional empirical evidence. We start by examining whether households may have manipulated the IFH index in order to become eligible for public insurance. Thereafter, second, we perform the analysis for a bigger sample. Throughout the analysis, we have controlled for covariates. Therefore, third, we also conduct the analysis without controlling for covariates and we assess whether there were jumps in the expectations of covariates at the eligibility threshold. In passing, we confirm that also the expectation of water and electricity consumption, respectively, do not exhibit such a jump. Fourth, we assess whether there were discontinuities at other values of the welfare index. This would raise concerns, as our approach builds on the premise that there is only one discontinuity of the expected outcomes in the welfare index, at least locally. After that, fifth, we conduct a non-parametric analysis. Sixth, we assess whether the existence of other programs could challenge the validity of our results. And finally, seventh, we discuss when our results could be considered lower bounds of the effects of interest. All corresponding tables and figures can be found in Online Appendix I. Manipulation of the running variable ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ A common threat to studies based on a RDD is the incentive to manipulate the running variable. For this, individuals need information on how the IFH is calculated. Then, they need to use this knowledge to manipulate their answers to the questions posed by the government official in order to qualify for SIS. This is unlikely to be the case for two reasons. First, even though the information on how the index is computed is, technically speaking, public, it is not easy to obtain and process it. Second, the set of variables included in the IFH construction are verified by the government officials and therefore difficult to manipulate. We nevertheless analyze this potential thread using the McCrary (2008) test. The idea is that if manipulation takes place, then the density of the running variable will be discontinuous at the cutoff. In our context, the density function would show many households barely qualifying for SIS, that is, to the left of the cutoff, and fewer failing to qualify, that is, to the right of the cutoff. The formal procedure is twofold: first, a finely gridded histogram is obtained and then this histogram is smoothed with a local linear regression on each side of the cutoff.47 Fig. A.10 presents the results. There is no evidence for a jump of the density at the eligibility threshold, supporting the assumptions made in the main analysis. Analysis for the full sample ~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our main results have been obtained under the assumption that expected outcomes are approximately linear in the welfare index, separately to the left and the right of the eligibility threshold. In order to alleviate the concern that this assumption is strong, we have performed the analysis locally, selecting a sample of individuals for whom the index is at most 20 points away from the threshold. In general, there is a tradeoff between precision and bias, and using a bigger sample has the advantage that the precision of our estimates may increase. Therefore it is interesting to also perform the analysis for the full sample and to compare the results to the main ones reported above. The first column of, respectively, Tables A.12 and A.13, shows the results. Comparing Table A.12 to Table 1 we see that the magnitudes of the estimated effects are slightly lower when we use the full sample while the precision increases. But qualitatively, the results are very similar. Comparing Table A.13 to Tables 2 and 3 we find a similar pattern, with the exception of the results on health expenditures as a share of income and catastrophic health expenditures. For the full sample, the magnitude of the estimated effects decreases by more than the standard errors and hence some of them are not found to be significantly different from zero. Overall, the picture remains qualitatively and also quantitatively the same: the effects on utilization are strongest, in particular for curative use, health expenditures increase in terms of levels and variability, driven by increases in out-of- pocket spending for medicines. Testing for discontinuities in household characteristics ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ For our approach to be valid it is necessary that the covered and non-covered individuals who have a value of the IFH index close to the eligibility threshold are similar to one another (Section 5). It is standard practice to test whether the expectation of covariates such as age or gender is a continuous function in the welfare index around the eligibility threshold. When it is found not to be, then one may be concerned that the assumptions underlying our analysis do not hold and one may want to conduct the analysis without controlling for covariates. We first conduct both a graphical and a formal analysis in which we replace the dependent health variables by the observed covariates gender, age, years of education, the number of household members, and whether the woman is the head of the household. These are the variables that we use as controls in order to be able to obtain more precise estimates. We also tested for discontinuities in household income, total household expenditure, and also water and electricity consumption. We would be concerned if water and electricity consumption would exhibit a discontinuity because eligibility is only based on the index if both of them are not big enough. Fig. A.8 and Table A.14 summarize the results. The latter reports estimates of the effect of insurance on these variables, conducted either at the household or at the individual level, depending on the variable. We do not find evidence for discontinuities. We also conducted the main analysis without controlling for covariates. Tables A.12 and A.13 show that results are actually very similar and that the main conclusions we have drawn remain the same. Jumps at non-discontinuity points ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our analysis implicitly assumes that the only discontinuities occur at the eligibility threshold, 55, as we have specified conditional expectations to be linear in the forcing variable, separately to the left and to the right of this threshold. A first way of assessing this is to conduct a graphical inspection. Figs. A.2 through A.7 suggest that the discontinuities do indeed mainly arise at the eligibility threshold. Besides, following Imbens and Lemieux (2008) we conduct separate additional RDD analysis for the samples of covered and non-covered individuals and use the midpoints of the index in the respective samples as the threshold values. That is, we test for a discontinuity at values of the index other than the actual threshold. Recall that in our main analysis we use only individuals whose index is at most 20 points away from the eligibility threshold. Here, we now use a sample of individuals with an index that is between 40 points lower than the threshold and the threshold, and another sample of individuals for whom the index lies between the eligibility threshold and 40 points above that. Results are presented in Tables A.15 and A.16. In general, we observe no significant effects on health outcome variables when we run the regressions using those hypothetical thresholds, with the exception of a few cases.48 Figs. A.2 through A.7 suggest that what is picked up by this robustness check is that the linearity assumption may be too strong for some outcomes. At the same time, we find many more effects to be significantly different from zero when we use the actual thresholds instead of the hypothetical ones. Next, we at least partially address a related concern by carrying out a nonparametric analysis. Non-parametric analysis ~~~~~~~~~~~~~~~~~~~~~~~ To address the concern that linearity is too strong of an assumption even in smaller subsamples we conduct a non-parametric analysis. For this we follow Calonico et al. (2014). The main difference in terms of implementation is that we drop individuals for whom either water or electricity consumption is too high to be eligible for health insurance. Comparing the results reported in Tables A.17 through A.19 to the ones in Tables 1 though 3 we again find that the strongest and most robust effects are on receiving medicines and hospital care and/or surgery, financed out-of-pocket, and positive effects on the level and the variability of out-of-pocket health care expenditures. Juntos and food aid program ~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our identification strategy is based on the assumption that discontinuities at the eligibility threshold can be attributed to SIS. There are some programs whose presence could in principle challenge this assumption. One of them is Juntos, a conditional cash transfer program. It combines a geographic targeting of the poorest districts with individual targeting, based on the IFH index and the presence of children up to the age of 14. However, Juntos is a rural program and our study focuses on the Lima Province, and our data confirm that no individual in the sample belongs to Juntos. Besides, there is a number of food aid programs oriented to the poor. To be precise, they are oriented to different groups of the population, such as mothers, children and school students. Our data show that 29% of the individuals of our sample receive at least support from one of them.49 Importantly, since these programs do not use SISFOH's targeting rules and in particular not the IFH index, it is unlikely that a discontinuity at the eligibility threshold can be attributed to them. Our finding in Section 7.3 that household expenditures do not exhibit a discontinuity at the insurance threshold provides additional support for this interpretation. Re-interpreting our results as intent-to-treat effects or lower bounds of the effect sizes ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In our analysis, we estimate the effect of becoming eligible for SIS due to crossing the eligibility threshold. As we describe in Section 5, we control for ineligibility that is due to other reasons and explain why we therefore face a sharp regression discontinuity design and estimate the average effect of becoming eligible. We argue in the Introduction and in Section 2.2 that this effect is essentially the effect of insurance coverage, because enrolling involves filling in a form and if eligible, individuals can usually come back the next day to receive treatment. It could nevertheless be that individuals do not know whether or not they are actually covered by health insurance. This by itself would not be a problem if they would always try to enroll and then learn that they are actually not eligible. If they wrongly believe that they are not eligible and therefore do not even try to enroll, then they may behave as if they are not covered by health insurance. Our analysis, however, assumes that they are covered. Consequently, the effects we estimate can be re-interpreted as intent-to-treat effects or lower bounds of effect sizes for those who know that they are covered. To see this, suppose that among those individuals with a value of the index that is very close to the threshold, 40% of the eligible individuals believe that they are not eligible and therefore the effect of becoming eligible is zero for them. The effect we then estimate is the intent-to-treat effect, which is a weighted average of the zero effect for those 40% who wrongly believe that they are not covered and the actual effect for the remaining 60% of the individuals. The intent-to-treat effect is always of the same sign but smaller in magnitude than the actual effect for those who know that they are eligible and in that sense we are estimating lower bounds of the effect sizes.50,51","Until recently, large parts of the population in developing countries did not have access to public health insurance. While it is commonly believed that the effects of health insurance coverage are positive, opportunities to control for selection by exploiting natural experiments or by conducting field experiments are rare, and therefore we still lack empirical evidence on its impact on health care utilization and out-of-pocket expenditures. Besides, it is not yet fully understood through which channels health insurance coverage ultimately leads to better health outcomes and to what extent it is possible to encourage individuals to invest into preventive care. In this paper, we use rich survey data from Peru to study the effects of the large-scale social health insurance program called “Seguro Integral de Salud” (SIS). The SIS program is targeted to poor individuals working in the informal labor market. We make use of the institutional details that give rise to a sharp regression discontinuity design. We estimate the effect of insurance coverage on a wealth of measures for health care utilization and health expenditures. We find strong effects of insurance coverage on arguably desirable, from a social welfare point of view, treatments such as visiting a hospital and receiving surgery and on forms of care that can be provided at relatively low cost, such as medical analysis in the first place and receiving medication. Effects on preventive care are much less pronounced. This is not surprising, as the system does not provide any extra incentives to actually use them. Furthermore, we find positive effects of health insurance coverage on the level and the variability of out-of-pocket spending that are mostly driven by increased spending for medicines and for hospital care and/or surgery, resulting from supply limitations. Based on this evidence, we develop two arguments that are less common in economics. First, access to health care centers leads to increased awareness about health problems. Once covered, individuals see a doctor and learn about the needs they were unaware of. Second, this even generates a willingness to pay for services that are in short supply, which in the context of Peru is a potentially desirable form of supplier- induced demand. By spending out-of-pocket, individuals reveal their preference for medical care. Overall, the evidence suggests that when compared to health care systems in other developing countries, the Peruvian one is a notable exception. It seems to reach its goal to provide access to medical care to a sizable fraction of the poor. A key determinant of this success seems to be that the monetary cost of enrolling is zero, instead of being small but positive, which it is elsewhere. As of now, there is no evidence on the effects this will have on objectively measured health, but it is imaginable that increased access will ultimately lead to better health outcomes."],["Using quarterly data for the U.K. from 1993 through 2012, we document that the extent of worker reallocation across occupations or industries (a career change, in the parlance of this paper) is high and procyclical. This holds true after controlling for workers' previous labour market status and for changes in the composition of who gets hired over the business cycle. Our evidence suggests that a large part of this reallocation reflect excess churning in the labour market. We also find that the majority of career changes come with wage increases. During the economic expansion wage increases were typically larger for those who change careers than for those who do not. During the recession this is not true for career changers who were hired from unemployment. Our evidence suggests that understanding career changes over the business cycle is important for explaining labour market flows and the cyclicality of wage growth. --------------------------------------------------------------------------------","One of the most important functions of the labour market is to pair the right set of workers with the right set of jobs. This assignment process, however, is slowed down by frictions that impede the reallocation of labour resources. For example, moving costs, re- training, learning about one׳s ability, information frictions about the location of workers or jobs, among others, can be important barriers for efficient resource reallocation. The result of these frictions is that we observe large concurrent flows of workers changing jobs directly from employer-to-employer as well as through spells of unemployment. As documented by Davis (1987) and Jolivet et al. (2006), among others, this excess churning is a common feature of all labour markets in OECD countries. The extent of reallocation is not necessarily constant over the business cycle. In one view, recessions are times in which the labour market is “cleansed” by speeding up the reallocation of workers, something that was prevented from occurring by frictions during the proceeding expansions (See, for example, Lilien, 1982; Mortensen and Pissarides, 1994; Caballero and Hammour, 1994; Groshen and Potter, 2003; Jaimovich and Siu, 2014). This view is appealing because it provides a possible explanation for why unemployment is persistently high in recessions. It simply takes workers time to switch, e.g., from jobs in industries and occupations for which demand is in secular decline to jobs in growing segments of the labour market. However, this is not the only view of the reallocative effects of recessions. Barlevy (2002) argues that, since employment-to-employment transitions are large and procyclical, economic expansions, rather than recessions, are times in which labour resources tend to reallocate to better uses. In his view recessions have a “sullying” rather than “cleansing” effect on reallocation. In this paper, we study two specific dimensions of reallocation: occupational and sectoral mobility of workers. If recessions have an important reallocative impact then occupational and sectoral mobility of workers are likely to be two important channels through which this reallocation occurs.1 In this context we interpret a career as a sequence of jobs a worker has in the same industry and occupation. A career change is a case in which a worker changes employer and starts a new job in either a different industry or occupation from the one he or she was previously employed in. We focus on career changes in the U.K labour market over the period from 1993 to 2012. The U.K. is an interesting country to look at for our purposes because it has one of the most flexible labour markets in Europe and exhibits one of the highest levels of worker turnover in the OECD (see Jolivet et al., 2006). This high level of turnover suggests that the U.K. labour market facilitates reallocation at a higher rate than those in other European countries. Fig. 1 shows the evolution of the U.K. unemployment rate during the period that we study, from 1993 through 2013. It shows that this period can be split up into four distinct episodes. The first episode is a period of economic expansion until 2001, during which the unemployment rate declined by about 4 percentage points.2 The second is a period of slow growth following 2001, when the U.K. economy skirted a recession and the unemployment rate blipped up marginally. The third episode is the economic expansion from 2002 until the start of the Great Recession in 2008, in which the unemployment rate remained centered around 5%. Lastly, the Great Recession and its aftermath make up the final episode. Fig. 1 shows that the unemployment rate increased by 3 percentage points during that period. It is the number and rate of industry and occupation changes, as well as the associated wage changes, in this final episode that we compare with the earlier parts of our sample. For this, we use individual- level data from the U.K. Quarterly Labour Force Survey. We present our evidence at two levels of detail. In the first part of our analysis we focus on aggregate patterns and uncover facts on (i) the extent of career changes in the labour market and (ii) how they fluctuate over the business cycle. In the second part we look closer at individual-level patterns that can shine a light on what drives these career changes. In this part we document (i) who change careers, (ii) which industries and occupations they come from and go to, and (iii) whether they do so at higher or lower wage gains than those who switch employers but stay in the same career. Five main findings emerge from our analysis of the U.K. Labour Force Survey. The extent of career changes is high: A worker who changes employers has around a 50% chance of switching to another occupation or industry. The rates of career changes are remarkably similar for those that change employers with or without an intervening spell of non-employment. Career changes in large part reflect excess churning in the labour market: the actual net mobility across industries and occupations due to career switches only amounts to 10% and 15% of the overall flows between occupations and industries respectively. This evidence on career mobility is in line with Longhi and Taylor (2011) who, using the same data source as us, find that the extent of occupational mobility in the U.K. is high.3 The U.K. is not an exception. Industry and occupational mobility rates are also high in the United States (see Moscarini and Thomsson, 2007; Moscarini and Vella, 2008; Kambourov and Manovskii, 2008; Hobijn, 2012, for example.) Career changes decrease in recessions: The total number of workers that change careers and the probability of a career change are procyclical. Moreover, for a worker, the probability of a career change is also procyclical, whether conditioning on changing employers directly, or on experiencing an intervening spell of non-participation, or a spell of unemployment. In this sense the cyclicality of career changes in the U.K. is similar to that in the U.S. For the U.S. Murphy and Topel (1987), Carrillo-Tudela et al. (2014), and Carrillo-Tudela and Visschers (2015) have all documented that the occupational and industry mobility is procyclical.4 Moreover, just like in the U.S., excess churning in the U.K. is the main driver of the cyclicality of overall mobility across occupations or industries. This is because employer-to-employer transitions, that account for the bulk of this churning, are procyclical. Moscarini and Thomsson (2007), Moscarini and Vella (2008) and Kambourov and Manovskii (2008), document these dynamics for the U.S. labour market. Characteristics of career changers: Career changes are more likely for (i) those workers actively searching for a job, (ii) those that made voluntary transitions (i.e. those who ‘resigned׳ from jobs, or gave up for ‘family or personal reasons’, as opposed to those that were made ‘redundant’ or ‘dismissed’) and (iii) those workers that work part-time or as temps. Though models of on-the-job search with multiple job types (as in Pissarides, 1994; Akerlof et al., 1988; Barlevy, 2002; Menzio and Shi, 2011; Hagedorn and Manovskii, 2013; Moscarini and Postel-Vinay, 2013, among others) do not specifically focus on career changes, and do not include a formal occupational or industry choice, they do imply that quits are procyclical. Our evidence suggests that many of these quits in the U.K. result in career switches. This is, however, not only the case for employment-to-employment transitions. Career changes are also very common for hires out of non-employment. In terms of underlying demographics, young workers and women are more prone to change careers than their older and male counterparts. Even after accounting for these characteristics, the propensity to change careers for workers that start a new job remains procyclical. Thus, our results are not due to changes in the composition of who gets hired over the business cycle. Career paths: Across occupations, career changes that involve an upgrade in the skill level are more likely through direct employer-to-employer transitions. On the contrary, career changes that involve a step down in skill level are more likely after spells of non-employment. Further, career changes tend to move workers from routine to non-routine employment. Our results also show that these movements did not accelerate during the Great Recession. Wage changes upon career changes: The majority of career changes come with wage increases and these wage increases tend to be bigger than for those workers that change jobs but remain in the same career. The wage gains for those who got hired out of unemployment and changed occupations fell during the recession and became smaller than the wage gains of those who did not change occupations. Several studies have linked wage gains to employer-to-employer transitions (Akerlof et al., 1988; Hagedorn and Manovskii, 2013). Our evidence here suggests that such wage gains disproportionately get realized by workers changing careers rather than continuing in the same one. These findings provide evidence as to which theories would be able to best explain labour market reallocation through occupational and industry mobility of workers. Our evidence shows that outcomes for career changers are different from those who remain in the same career when changing jobs. This suggests that understanding career changes over the business cycle is important for explaining the cyclicality labour turnover and wage growth. Most current models of labour turnover, like those that allow for on-the-job search mentioned above, provide theories of why turnover is highly procyclical. Though these theories have heterogenous jobs, none of them explicitly considers a career change decision. Recent models, like Carrillo-Tudela et al. (2014) and Groes et al. (2015), do contain a career change margin and help us better understand the incidence of career changes over the business cycle and across the income distribution, respectively. Taken together, the facts we document are consistent with the view that the Great Recession and its aftermath has affected workers across a large set of industries and occupations, with a broad-based shortfall in economic activity preventing workers from pursuing alternate careers at substantial wage gains. In this sense, our results are consistent with the “sullying” effect of recessions put forward by Barlevy (2002). Of course, career changes are only one form of reallocation of labour and other resources. Thus, our results do not imply that recessions have no cleansing effect at all but rather that such a cleansing is not happening through worker reallocation across occupations and industries. This is important, because it means we find little support in the U.K. data for recent theories of job polarization (Jaimovich and Siu, 2014) that point to occupational mobility between routine and non-routine jobs during recessions as the major driving force of the secular decline in routine jobs. The rest of the paper is structured as follows. In the next section we discuss the Quarterly U.K. Labour Force Survey, the definitions of the main variables, as well as the level of aggregation of the industry and occupational classifications that we use. In Section 3 we present the aggregate evidence and focus on broad patterns in the level and cyclicality of career changes in the U.K. In Section 4 we present individual-level evidence and discuss what it suggests about the reasons for career switches. Finally, we end with a brief discussion of the theoretical implications of the facts we document in Section 5.","The data we use are from the U.K. Quarterly Labour Force Survey (LFS) and cover the period 1993Q1–2012Q3. The LFS has a rotating panel structure, depicted in Fig. 2, in which individuals that live on the sampled address are followed for a maximum of 5 quarters, also referred to as waves. Each quarter, one-fifth of the sample of addresses is replaced by an incoming rotation group, or cohort. From this sample, we consider all male workers between 16 and 65 years of age and all female workers between 16 and 60 years of age with an ongoing career.5 In each wave, the respondents provide information about, among other things, their labour market status as well as their occupation and the industry they work in if they are employed. If non-employed, they provide the occupation and industry of their previous job.6 Because we are interested in those workers who switch employers and potentially change careers, and because non-employed workers provide information on previous employment, we need observations on workers only for two consecutive quarters. Thus, we use the two-quarter (2Q) longitudinal sample of the LFS. Fig. 2 depicts two quarters of this sample as long-dashed rectangles, labeled “2Q”. As can be seen from the figure, because of the rotating panel structure and sample attrition, the 2Q sample is smaller than the quarterly cross-section. It consists of about 60,000 individuals each quarter.7 Occupation and industrial classifications: To code occupations, the U.K. LFS uses the Standard Occupational Classification (SOC). The occupational coding system was redefined in 2001, from the SOC 1990 to the SOC 2000, which was used until the end of 2010. A drawback of this revision is that the SOC 1990 and SOC 2000 are not fully compatible. To reduce potential incompatibility errors we focus on mobility across 1-digit or major occupational groups. These groups are listed in Table 1 for both the SOC 1990 and SOC 2000. At this level of aggregation, the disagreement between the two SOC is of 26.5%. The disagreement between the two classifications introduces a level shift in some of the occupational series at the time of the switch from SOC 1990 to SOC 2000. To correct for this shift, we adjust all 5-quarter centered moving average series by running an OLS regression on the log of the corresponding series with respect to a linear time trend, the log of output per worker and a dummy which takes a value of zero before 2000Q4, and one after. We then use the coefficient estimate of the dummy variable (irrespectively if it was significant or not) to adjust the series up to 2000Q4.8 To code industries, the U.K. LFS uses the Standard Industrial Classification (SIC). In this case the U.K. LFS does provide homogenised industry information for workers for the entire sample period based on the SIC 1992.9 We focus on industrial mobility on broad industrial sectors, which roughly corresponds to a one-digit aggregation level, with 17 categories displayed in Table 2. Wage analysis: For the last part of our analysis, we also consider the change in wages when workers switch occupations or industries. The wage measure we use is the self- reported gross weekly earnings, deflated using the CPI. Individuals in the LFS only report their wages in the first and fifth waves. These are depicted by the circles labeled “W” in Fig. 2. Because they report their wages one year apart, we can calculate annual wage growth for these workers. However, to do so requires us to follow these workers for the full five quarters that they are in the LFS. This sample is known as the five-quarter longitudinal sample and is depicted by the short-dashed rectangle labeled “5Q” in the figure. This sample contains, on average, about 11,000 individuals. Using this sample we condition the wage analysis on employer changes through employment, unemployment or inactivity based only on uninterrupted spells.10 We aggregate all these transitions to analyse the wage changes among all workers. Level and probability of career changes ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ We record a career change when a worker changed employer and reported an occupation or industry in the new job that is different from the occupation or industry reported in the last job held. Then, what is flagged as a career change depends on the level of aggregation of the occupation and industry classifications used. Because we use the major occupation and industry classifications discussed above, the career changes we flag capture a substantial change in the nature of a worker׳s job.11 Net mobility ~~~~~~~~~~~~ Theories that emphasize the cleansing effect of recessions on the labour market emphasize how downturns accelerate the shift in labour market resources from segments that are in structural decline to those that are on a positive long-run trend. These are theories that focus on the net mobility of workers across professions and sectors. To put this net mobility in the context of the magnitude of overall flows in the labour market, we follow Davis and Haltiwanger (1992) and analyze excess reallocation. That is, we quantify by how much the total gross reallocation measured by the flows introduced in the previous subsection exceeds the minimum flows needed to achieve the net shift in the observed allocation of workers across occupations and industries.","In this section we investigate both the level as well as the cyclical fluctuations of the incidence of career changes in the U.K. labour market. In the first subsection we focus on the level and report long-run averages over our whole sample period. In the second subsection we shift our focus to how the prevalence of career changes moves over the business cycle. Long-run averages ~~~~~~~~~~~~~~~~~ The U.K. labour market displays a surprising degree of churning. Over our sample period, the sum of career movers and stayers is on average 1.3 million per quarter. This amounts to 4.5% of the U.K.׳s working age population. Of those who get hired and have a previous career, 43% come directly from a previous employer, 29% are hired out of unemployment, and 29% were out of the labour force. These numbers are in line with Gomes (2012). What is even more striking is the high share of these hires that involve a career change. Table 3 shows the average fraction of these hires that we classify as a career change. As can be seen from the top row of the table, 49% of those workers with a previous career who start a new job do so in a different (major) occupation from which they worked in before. This fraction is even higher for industries, for which the majority, 53%, of such hires involve a switch in major industry. The similarities in the extent of career changes across occupations or industries arises mostly because the majority of career movers change occupations and industries at the same time. For example, on average 75% of workers who changed occupations also changed industries and 70% of workers who changed industries also changed occupations. Though high, these numbers are in line with evidence for the United States. For example, Carrillo-Tudela et al. (2014), using data from the Current Population Survey, and Carrillo-Tudela and Visschers (2015), who rely on the Survey of Income and Program Participation, both find that about half of the hires in the United States involve a career change as well. One caveat is important to note. Reporting errors, more so for occupations than for industries, are common in surveys like the U.K. LFS. If estimates from other datasets are applied to our results for the U.K. LFS, then, maybe even as much as a quarter, of the career moves that we measure could be due to workers misreporting their occupation and/or industry in the survey.13 However, even if this is true, this would still mean that about a third of all hires of persons with previous work experience involves them changing either the industry or profession that they work in. Even after such a drastic downward adjustment, this would imply that more than one percent of the U.K. working age population switches careers every quarter. Rows 2 and up of Table 3 list the probability of a career change conditional on the labour market status of the worker in the quarter before she or he starts a new job. As can be seen from the table, the average probability of a career change is around 50% for each of these types of hires. Two groups of workers stand out as having a higher probability of switching careers than others. The first consists of workers who make an EE transition and who actively searched for the new position in the old job. These are more likely workers who actively pursue a voluntary change in their career path. To be specific, career or job changes are categorised as voluntary when workers report in the LFS that they left their previous employer because they “resigned”, went to “education or training” or “gave up for family or personal reasons”. Involuntary career or job changes are made by those workers who left their last job because they were “dismissed”, “made redundant/took voluntary redundancy”, “temporary job finished” and “gave up work for health reasons”. Finally, workers in the other group are those who left their last job because they “took early retirement”, “retired” and due to “other reasons”.14 Active search encompasses all activities that involve the worker to contact or actively pursue job opportunities rather than browse job opportunities that are available. This is the definition of job search that defines a person without a job as being unemployed. The specific LFS answers that result in a person being classified as an active searcher are listed in the Appendix. The second group of workers with a higher probability of moving to a different career are those who were unemployed for two quarters or more in the quarter before they started their new jobs. These transitions most likely reflect involuntary career decisions that occur in long spells of unemployment. Such career changes are often emphasized as driving up the natural rate of unemployment in the short-run in the wake of a recession due to mismatch in the labour market. Recent studies show that mismatch can only account for a small part of overall fluctuations in the unemployment rate.15 Most studies of mismatch in the labour market compare the composition of job openings by industry and occupation with the composition of the pool of unemployed workers. This assumes that it is the pool of unemployed workers that are required to make all the adjustments to make the skill composition of the labour supply adjust to the composition of skills demanded. It turns out that more than half of the workers that get hired out of unemployment end up making such an adjustment. Moreover, our results suggest that the large number of EE career switchers helps to accelerate this adjustment process. By providing a measure of the gap between the skill requirements needed to fill the stock of job openings and the skill composition of the pool of unemployed, measures of mismatch are a proxy for the net amount of reallocation needed in the labour market to equilibrate the supply of and demand for skills. However, gross mobility between careers far exceeds net mobility. The average net mobility rates, nmt, over our sample period are 10% for occupations and 13% across industries.16 This echoes the findings for the U.S. of Jovanovic and Moffitt (1990), Kambourov and Manovskii (2008) and Auray et al. (2014), who show that net mobility accounts for only a small proportion of gross mobility across industries and occupations. Cyclical fluctuations ~~~~~~~~~~~~~~~~~~~~~ Whether recession are times of accelerated or of relatively slow reallocation in the labour market can, of course, not be gleaned from the long-run averages we reported so far. To answer this question we now present evidence on the fluctuations, in deviation from these averages, in the extent and probabilities of career changes over our sample period. Above, we have focused on comparing the Great Recession with the previous episodes in the data. The procyclicality of the level and probability of career changes that we documented, however, is also robust to other ways of business cycle accounting. For example, it also shows up if one uses the Hodrick and Prescott (1997) filter to distinguish between trend and cycle in the unemployment rate and the time series plotted in Figs. 3 and 4.19 One possible explanation for the procyclicality of the propensity to change careers out of unemployment is the increased incidence of workers being recalled to their previous job during downturns. For example, Fujita and Moscarini (2012) find that, in the U.S., those workers that become unemployed after being permanently separated from their previous jobs are much more likely to make an occupational change than those that were on layoff and recalled within 3 months. However, in the UK such recall practice is minimal and, hence, is thus not likely to affect the results presented here. What could be more pertinent is that, on the supply side, those workers who get laid off in recessions would first look for a job that is similar to the one they lost and only slowly broaden their search.20 However, as Carrillo-Tudela et al. (2014) argue, workers take into account that they may be less likely to start a particularly successful career path during a recession, which reduces their incentives to change careers at any duration. On the labour demand side, because of the increased size of the pool of unemployed workers in recessions, employers would be more likely to find candidates that more closely match the career profile they are looking for. Some studies, like Ravenna and Walsh (2012) and Sedláček (2014), suggest that employers also get more selective in their hiring practices during downturns. Such an increase in the pickiness of employers about who they hire in downturns also affects the opportunities of those who are employed and are looking to change jobs and pursue a different career. These effects could result in a decline in the fraction of EE transitions that result in a switch in industry or occupation during recessions, as can be seen from the long-dashed line in Fig. 4. Another way to gauge the relative importance of these effects is to look at the fluctuations in net mobility, NM, over the business cycle. Net mobility for both occupations and industries is plotted in Fig. 5. If recessions had a major “cleansing” effect that resulted in a substantial shift in workers from occupations and industries in secular decline to those for which demand is booming, then net mobility would increase during the recession as well during the subsequent recovery. This is because during the recovery workers would, gradually perhaps, find jobs in careers different from those that they were in before. It is exactly this slow adjustment during the recovery that is often pointed to as a source of the jobless recoveries from the last three recessions in the U.S. (Groshen and Potter, 2003; Jaimovich and Siu, 2014). However, as Fig. 5 shows, there is no such persistent spike in net mobility. Net mobility briefly went up at the onset of the Great Recession, but then declined to levels rather lower than typical values in the period 2001–2008Q1. While the early rise coincided with the wave of layoffs described by Elsby and Smith (2010), by the end of the recession net mobility rate had fallen deeply, however. From this low level, net sectoral mobility started to increase again during the 2010–2011 recession, only reaching pre-recession levels at the end of the second recession. The increase in net mobility in 2010 and 2011 is mainly due to workers flowing towards services sectors. The main contributors to this increase are all in the service sector (in order of importance): (i) Real estate, renting and business activities; (ii) Health and social work; (iii) Education; (iv) Wholesale and Retail Trade including Repairs; and (v) Transport, storage and distribution. This evidence on net mobility, together with that on the level and probability of career changes presented above, is in line with Barlevy׳s (2002) interpretation of the role of business cycle for labour market dynamics, here for career changes, rather than job changes. He argues that, because labour turnover is higher during expansions than during downturns, the reallocation of labour market resources is procyclical rather than countercyclical. Our interpretation of the above results is that, in terms of worker reallocation across occupations and industries, recessions do not appear to be times of accelerated labour market reallocation which is prevented from happening during expansions due to frictions. Instead, in a recession, workers seem to stay put in their respective occupations and industries when labour market opportunities for them dry up during downturns.","In this section we dive into the details underlying these aggregates and use additional information from the U.K. LFS to analyse the reasons for the career changes, who changes careers, what they do before and after the career change, and how the change affects their wages. This turns out to yield further evidence supportive of the “sullying effect” of recessions through the lenses of career changes. Reasons for career change ~~~~~~~~~~~~~~~~~~~~~~~~~ Unfortunately, the U.K. LFS survey does not directly ask respondents who take jobs in a different occupation or industry about the specific reason for their career change. However, some of the questions asked allow us to indirectly infer some of the potential reasons. In particular, we revisit the questions we first focused on in Table 3. That is, for those who move directly from one employer to another we consider whether this move was voluntary and whether or not they had been actively searching for a job before they switched. For those who were unemployed in the quarter before they started their new job, we consider the duration of their unemployment spell in that quarter. Because EE flows account for the bulk of the turnover in Fig. 3, we focus on the evidence for this switchers first. Fig. 6 divides up the EE flows into movers and stayers and classifies them by whether or not they made a voluntary EE switch, panels (a) and (c), and by whether they were actively searching on the job before they made the switch, panels (b) and (d). The first thing that stands out from the figure is that the bulk of EE transitions are voluntary. Moreover, the vast majority of EE transitions is not the result of the worker actively searching for another job but rather of the worker getting a job offer without searching. We interpret these two facts as suggesting that a lot of job changes are voluntary quits that could occur as result of employers contacting workers. Recent evidence for the U.S. also shows that many workers get hired without ever reporting to be actively looking for a job (see Topa et al., 2014; Carrillo-Tudela and Visschers, 2015, for example). This contrasts with the common perception, as expressed in Jaimovich and Siu (2014), that recessions are times of accelerated involuntary structural transformation. During such times a large number of workers supposedly gets displaced from jobs that will never come back and thus are forced to look for and take jobs in sectors and occupations different from those they worked in before. One possible explanation for why the incidence of career changes among hires out of unemployment does not spike in the recession is that workers that get displaced from jobs that are in secular decline might decide to drop out of the labour force rather than to switch careers. This is especially a concern in the United States, where the labour force participation rate dropped by more than 3 percentage points in the five years after the start of the Great Recession.22 Such flows to inactivity, however, are not likely to be important in the U.K. where the labour force participation rate actually increased between 2007 and 2012. Who changes careers? ~~~~~~~~~~~~~~~~~~~~ Of course, the discussion in the previous subsection focuses on the Great Recession versus the rest of the sample. In addition, the evidence presented does not condition on other factors that might be correlated with the variables used to proxy for different reasons for a career change. Here we show that the procyclicality of the probability of career changes, shown in Figs. 4 and 7, is statistically significant even if one considers the whole sample and also corrects for factors that affect the probability of a career switch. We do so by presenting Probit estimates derived from a model where the dependent variable is whether or not the hire of a worker with previous work experience results in a career change. The explanatory variables include a set of worker characteristics, properties of the job the worker is hired in, and variables that proxy for the potential reasons for why the worker changed careers or not. Because the availability of some of the variables related to the reasons for the career change depends on the labour market status of the worker before he or she accepted the new job, we present the Probit estimates not only for all hires but also condition them on what labour market status the worker had in the quarter before starting the new job. The estimation results are presented in Table 4.23 In terms of the effects of human capital on the probability of a career change, we find that age decreases the probability of a career change, suggesting the importance of on-the-job human capital accumulation. Educational attainment, however, affects occupations and industries differently. Across occupations, high and medium skilled workers have a higher probability of a career change than low skilled workers (our reference category). Across industries, we find that low skilled workers have a higher probability of a career change than medium and high skilled workers. These results seem to arise from differences in the impact of skill levels by employment status. Across occupations, it is only the unemployed for which high and medium skilled workers have a higher probability of a career change. Across industries, low skilled workers have a higher probability of changing career when mobility is through employment or inactivity, but not through unemployment. Table 4 also shows the effects of different types of job characteristics on the probability of a career change. This probability increases if the worker obtains a part-time versus a full-time job or if the worker obtains a temporary versus a permanent job.24 Women have a higher probability of a career change than men. Furthermore, the larger the household someone is part of, the less likely a person is to change careers. That is, Hm is lower for persons who are married or cohabitate. It also decreases, although not significantly, in the number of children. The Probit estimates also reaffirm the results found in Table 3 and Figs. 3, 6, and 7. We find that for employed workers, career changes are more likely among those employed workers that made voluntary EE transitions and among those that were actively searching for a job (our baseline category with respect to all the search channels). Unemployed workers are more likely to make career changes than employed (our baseline category) or inactive workers, while a career change through unemployment is more likely to occur at longer unemployment spells. Using individual-level data in the Probit regression allows us to shine a more detailed light on search method workers employed to find their new jobs and how it affects their chance of changing careers. In particular, the explanatory variables listed in Rows 18 through 21 get at this.25 We find that those workers who find jobs responding to ads are more likely to change careers than those who find jobs through other means. Conditioning on the worker-, job-, and search- characteristics does not erase the significance of the procyclicality of career changes. This suggests that the business cycle movements in occupational and industry mobility of workers are not the result of the composition of the group of workers with a previous career that gets hired changing with the cycle. The higher sensitivity of occupational switches compared to industry switches to the aggregate unemployment rate is offset by the higher sensitivity of occupational mobility with respect to the regional component of the unemployment rate, reported in Row 2 of Table 4. Taking the results of Rows 1 and 2 of Table 4 together both occupational as well as industry mobility comove very significantly with labour market conditions. Origins and destinations ~~~~~~~~~~~~~~~~~~~~~~~~ Another way to gauge the reasons for career switches is to consider what type of job in which industry and occupation workers come from and what type of job they end up in. This is what we explore in this subsection. We focus on three aspects of the origins and destinations of career changers in our data. The first is whether the jobs are full- or part-time. The second is what industry and occupation career changers come from and which ones they go to. Finally, we refine the occupation analysis by considering whether the occupations are routine or non-routine. Full- versus part-time jobs: So far, we have documented that most career changes result from voluntary labour turnover and that the share of career changes that is voluntary is procyclical. That is, during downturns a higher fraction of career changes is involuntary (see Fig. 6). This cyclical behaviour of voluntary career changes is mirrored by the extent to which occupational mobility results in full- or part-time jobs. Career changes turn out to be an important mechanism through which workers move between part-time and full-time jobs and, on net, contribute positively to part-time and to full-time job flows.27 On average 65% of hires resulted in a full-time job and 35% of hires resulted in a part-time job during the 1993–2007 period. These hires are disproportionately people who change occupations.28 Career movers on average get a full-time job in 60% and a part-time job in 40% of the time. For those that switch directly between employers we know both their full-time status before and after they get hired and can thus infer whether their full-time status changed when switching jobs. Using these data, we find that on average 13% workers making an EE transition move from part- time into full-time employment, while 7% move from full-time to part-time employment during the 1993–2007 period. The bulk of changes in the full-time nature of work, in either direction, involves a career change. Of those who moved from part-time into full- time employment, 66% changed careers; while from those that moved from full-time to part- time employment 59% changed careers. During the Great Recession, however, the incidence of part-time work increased. On average 37% of hires now resulted in a part-time job, while 63% of hires resulted in a full-time job. Consistent with this, the net contribution of career changes to part-time-to-full-time flows declined during the same period.29 Thus, if we would consider part-time jobs to be typically less desirable than full-time jobs, then the shift in the full-time/part-time composition of career movers׳ new jobs during the recession reflects a relative worsening of outcomes associated with changing careers in downturns and thus a deceleration of the pace with which workers move to higher quality jobs during those periods. Note, however, that the shift in the full-time/part-time composition is much less pronounced than the shift in terms of voluntary versus involuntary turnover, depicted in Fig. 6. Industries and occupations: Above, we suggested that transitions from part-time to full-time jobs are generally considered a step up the job ladder while the reverse are considered a step down. To paint a more detailed picture of the job ladders that career changers are on, we consider the origins and destinations of their career moves here in terms of industry and occupation. We do so by constructing industry and occupation transition matrices for workers׳ career changes. These matrices provide useful information on the mobility patterns of workers as they shed light on the potential importance of individual occupations or industries in driving overall mobility. Table 5 shows the transition matrix for workers changing careers across occupations.30 This matrix shows that all occupations exhibit a high degree of mobility. The dark-shaded cells list the fraction of hires that get hired in the same major occupation as they were working in before. Looking at the numbers for all hires, labeled as “Total”, the probability of a career change ranges from 61% for sales occupations to 32% for professional occupations. Across occupations, however, we observe some clustering by skill level. To show this, we group together those occupations that require similar skill levels. This results in three groups of high-, medium-, and low-skilled occupations. The first two groups consist of three major occupation codes and the last group consists of two major occupation codes. Career changes within each of these groups are highlighted in light grey as the block diagonal in the transition matrix. As can be seen, the transition probabilities in the grey cells tend to be higher than those in the other cells. There are two destination occupations that are notable exceptions to this pattern. First, a substantial number of career changes out of high-skill occupations result in jobs in “Clerical and administrative” jobs. Second, the miscellaneous ninth category absorbs a large number of career switchers from middle-skilled jobs.31 Although we observe similar non-diagonal probabilities between rows in the transition matrix, we also observe that workers are more likely to stay within their skill category or move to the highest skill category after an EE transition and more likely to move to a lower skill category through a UE or IE transition.32 These patterns suggest that workers tend to move more often to occupations that demand skills closer to the ones they can supply. However, conditional on moving to a different skill category, workers are more likely to make career changes that involve an upgrade in the skill level through direct EE transitions, while career changes that involve a lower skill level are more likely through spells of non-employment. This evidence reinforces the view that occupational mobility through EE transitions are more likely to be voluntary career changes in which workers mostly pursue upward career moves, while occupational mobility through non-employment are more likely to be involuntary career changes. Routine and non-routine occupations: One particular type of occupational mobility that has been emphasized in the recent literature is that between occupations that involve routine and those that involve non-routine tasks. The distinction between these two types of occupations is relevant for the “Polarization” hypothesis (see Autor, 2003, Acemoglu and Autor, 2011, Autor and Dorn, 2013, among others). This hypothesis is that, over the last decades, job tasks that can be captured easily by a set of explicit of simple instructions or rules, i.e. ‘routine tasks’, have been increasingly taken over by computers and machines. As a result, employment in those occupations in which workers are mainly executing routine tasks, summarily called ‘routine occupations’, has declined. In its place, employment has risen at the bottom of the wage distribution, in occupations that require physical labour, yet with tasks that cannot easily be captured in routines to be automated. This includes simple service jobs that require physical eye–hand coordination and physical navigation, typically under the heading ‘non-routine manual’ jobs. Employment has also risen higher in the wage distribution, where tasks require knowledge acquisition and creative thinking, with jobs put under the ‘non-routine cognitive’ header.33 Jaimovich and Siu (2014) argue that this secular process of job polarization accelerates during recessions when many routine jobs are permanently destroyed and workers in those jobs are forced to pursue other careers. In this way, they claim, the cycle is actually the trend, since this type of job polarization during recessions is not reversed during expansions. To consider whether job polarization is happening in the U.K. labour market and to what extent it is reflected in workers switching from careers in routine to non-routine occupations, we split up the post-2000 data by occupation into routine and non-routine occupations, following Acemoglu and Autor (2011). The second column of Table 1 contains a marker that signifies which SOC 2000 occupations are classified in which category. Fig. 8 shows employment in routine occupations as both a share of the working age population as well as of total employment. The figure shows that the share of employment in ‘routine occupations’ has steadily declined in the U.K., similar to that in the U.S. (Jaimovich and Siu, 2014). However, there was no acceleration in this trend during the Great Recession, as the “trend-is-the- cycle” hypothesis would suggest. In fact, using more formal regression-based techniques we find no significant cyclical component in the routine share series plotted in Fig. 8. This is in line with the evidence for the U.S. in Foote and Ryan (2014). Fig. 9 shows the time series of career changes that result in a switch between routine and non-routine occupations. The first thing that stands out from this figure is the excess churning we already saw in terms of the net mobility measure in Fig. 5. The net change in routine employment induced by these career switches is negative and contributes to the trend decline shown in Fig. 8. Just like in the U.S. (Cortes et al., 2014) IE and UE flows contribute the bulk of this net decline. Most importantly, however, is the observation that the share of routine to non-routine career switches does not increase significantly during the recession, indicating that, in terms of career switches, there is no evidence that the long-run downward trend in the share of routine employment accelerates during recessions. In fact, the overall turnover between these two categories of occupations seems to have declined in the recession. Of course, the data in Fig. 9 only includes workers who have been employed before at some point, and are hired again. This means that adjustment in the overall level of routine employment could also come about by a diminishing inflow into routine occupations by labour market entrants, and by an increased outflow of retirees from these occupations is not visible in our statistics. Cortes et al. (2014), for example, emphasize that such a cohort effect is an important driver behind the trend decline in routine employment in the United States. However, the lack of a cyclical pattern in Fig. 9 suggests that this cohort effect most likely also does not fluctuate a lot over the business cycle. Thus, our analysis for the U.K. is supportive of the same conclusion that Albanesi et al. (2013) draw for the U.S.; weakness in the labor market in the Great Recession was shared by non-routine and routine occupations alike, did not disproportionately affect routine occupations, nor did it accelerate the secular decline in routine jobs. Wage gains ~~~~~~~~~~ Thus far, we have shown that career switches make up a substantial fraction labour market turnover, and of voluntary turnover in particular. Recent theoretical (Hagedorn and Manovskii, 2013) and empirical (Daly et al., 2012) studies have emphasized the importance of voluntary turnover and employer-to-employer transitions for understanding the cyclical behaviour of wage growth. Our data suggest that distinguishing between career switchers and stayers would refine our understanding of wage growth over the business cycle even more. To see why, consider Tables 6 and 7, which summarize the distribution of percent real wage changes for job switchers, conditional on moving careers or staying in the same career, for the whole sample as well as for the three main periods in our sample.34 Because we are interested in wage changes, our analysis only includes hires for which we have data in waves 1 and 5 of the survey, depicted in Fig. 2. In particular, that means that for workers who flow through unemployment, we only have wage changes for those with an unemployment spell shorter than 4 quarters. Long-run perspective: Table 6 shows the probability that the hire of a worker with previous work experience results in a wage gain. The table lists this probability conditional on whether the hire involves a change in career and on the level of the wage earned in the previous job, measured in terms of the percentile of the wage distribution. The probability of a positive wage gain is much higher for workers who earned a low wage in their previous jobs. More importantly, for those workers this probability is also higher when they change careers than when they did not. For workers making an above-median wage, however, the probability of obtaining a positive wage growth when changing employer is closer to 30% but now is higher for those who do not change careers. This suggests that a large part of the voluntary career mobility through employer-to-employer moves that we document is workers moving up the job ladder to progress their careers.35 Where Table 6 provides information about the sign of the wage change, the columns for the “Whole sample” in Table 7 show the distribution of the magnitude of wage changes.36 The first takeaway from this table is the large degree of dispersion in wage growth that results from a change in employers. Below the 50th percentile of each distribution, workers can experience large negative wage losses when moving employers, while above the 50th percentile workers experience large wage gains.37 The most striking feature of the distributions shown in Table 7 is that the dispersion of wage gains is larger for career movers than for career stayers. This also holds true when we condition on whether the worker changed employers through an intervening spell of unemployment or not. Relative to stayers those who changed careers have higher wage growth at and above the 50th percentile of the wage growth distribution; while the opposite happens below the 50th percentile. This evidence again supports our interpretation that workers typically change careers for wage gains bigger than for those that stayed in the same occupation. It might seem counterintuitive that career changes through unemployment do tend to lead to positive wage gains that are larger than those obtained by unemployed workers who will stay in the same career. However, this evidence is not inconsistent with a theory in which these potentially larger wage gains can only be obtained after a costly reallocation process which only becomes worthwhile after job prospects in the original career have deteriorated sufficiently (see, for example, Carrillo-Tudela et al., 2014). Cyclical patterns: The last six columns of Table 7 show how the distribution of wage changes varies over different business cycle episodes in the U.K. labour market. Across occupations and industries the wage growth distribution of those workers that change careers through unemployment shifts down during the recession. The decrease is stronger across occupations than industries. Further, the shift in the wage growth distribution of those who changed occupations through unemployment is sufficiently big that their wage gains are now below the wage gains of career stayers even at the 75th percentile of the wage growth distribution. In contrast, the wage growth distribution of workers that changed employers directly through an EE transition or those that changed employers through unemployment but did not undertake a career change, do not seem to respond as much to business cycle conditions.38 The evidence presented suggests that career changers have a higher probability of a substantially large wage increase than career stayers. However, during the recession the wage gains of occupational changers decrease to the point that, for unemployed workers, these have become smaller than the wage gains from changing employer in the same occupation. As argued, for example, in Carrillo-Tudela et al. (2014), the decrease in the gains of reallocation can help explain the drop in the probability of a career change during the recession, documented in Section 3.2. Thus, the procyclicality of the incidence of career changes and the associated wage gains that we document suggest that adding a career change margin to our models of labour market fluctuations will help improve our understanding of the, not well-understood, link between unemployment, labour turnover, and aggregate wage growth.39","Overall, the patterns in the UK LFS suggest that in good times career changes imply a chance to improve a worker׳s position in the labour market. In downturns the gains associated with career changes appear to diminish. From a theory perspective one can build on Carrillo-Tudela et al. (2014) and Wiczer (2013) to reconcile these patterns using a framework that incorporates heterogeneity in labour market conditions, costly mobility choices between labour markets (career changes) and business cycle shocks. In such a framework, fluctuations in the expected net returns to a career change induce workers to adjust their mobility choices. In downturns, when net returns are low, workers decide to stay in their careers and wait for conditions to improve instead of changing to a new occupation or industry. In such a framework two motives for job mobility can arise: (i) workers may move to other jobs because their current employment conditions worsen while outside opportunities stay the same; (ii) workers may move because outside opportunities improve while current employment conditions are unaffected. Although both reasons may be at work, they are not necessarily two sides of the same coin. Aggregate conditions may interact differently with the idiosyncratic shocks to workers׳ current employment than with the stochastic arrival of new employment opportunities in different occupations or industries. In these models, adverse shocks to current employment could then generate ‘involuntary’ transitions, through which workers try to recover the loss of prospects in their current job. Increased opportunities elsewhere could draw workers to ‘voluntarily’ change their jobs and careers. The ‘pull’ of the latter kind of opportunities can be especially strong in booms, in line with the evidence presented in this paper; while the mobility ‘push’ associated with the shocks behind ‘involuntary’ transitions could be especially relevant in recessions. Taken together, career changes are different from other hires in terms of their cyclicality, their associated (wage) gains and the cyclical variations in these gains. Incorporating a career-mobility dimension in equilibrium business cycle models of the labor market can be a promising direction to contribute to our understanding of the overall behaviour of labour turnover and wage growth over the business cycle, and could help guide better policy responses to business cycle fluctuations."],["The relative return to strategies that augment inputs versus those that reduce inefficiencies remains a key open question for education policy in low-income countries. Using a new nationally-representative panel dataset of schools across 1297 villages in India, we show that the large public investments in education over the past decade have led to substantial improvements in input-based measures of school quality, but only a modest reduction in inefficiency as measured by teacher absence. In our data, 23.6% of teachers were absent during unannounced school visits, and we estimate that the salary cost of unauthorized teacher absence is $1.5 billion/year. We find two robust correlations in the nationally-representative panel data that corroborate findings from smaller-scale experiments. First, reductions in student-teacher ratios are correlated with increased teacher absence. Second, increases in the frequency of school monitoring are strongly correlated with lower teacher absence. Using these results, we show that reducing inefficiencies by increasing the frequency of monitoring could be over ten times more cost effective at increasing the effective student-teacher ratio than hiring more teachers. Thus, policies that decrease the inefficiency of public education spending are likely to yield substantially higher marginal returns than those that augment inputs. --------------------------------------------------------------------------------","Determining the optimal level and composition of public education spending is a key policy question in most low-income countries. Many education advocates believe that low-income countries need substantial increases in public education spending to meet enrollment and learning goals (UNESCO, 2014); others argue that public sector inefficiencies leave considerable room for improvement within existing education budgets, and that fiscal constraints make it imperative to improve the efficiency of public expenditure (World Bank, 2010). However, the data to assess the relative importance of these contentions remains sparse, in part, due to the difficulty in detecting and measuring inefficiencies in public spending. In this paper, we study one striking measure of public sector inefficiency - teacher absences - with panel data collected 7 years apart in India at a time of sharp increases in education spending. A large portion of this increase was accounted for by the salary cost of hiring teachers to reduce the student-teacher ratio in public schools. As a policy alternative to hiring more teachers, we show that reducing teacher absences by increasing school monitoring could be over ten times more cost effective at reducing the effective student-teacher ratio (net of teacher absence). Thus, while the default approach to improving education in low-income countries is input- augmentation, our results suggest that investing in reducing inefficiencies may yield much greater returns. India presents a particularly salient setting for our analysis. It has the largest primary education system in the world, catering to over 200 million children. Further, over the past decade, the Government of India has invested heavily in primary education under the Sarva Shiksha Abhiyan (SSA) or “Education for All Campaign.” Partly financed by a special education tax, this national program sought to correct historical inattention to primary education and led to a substantial increase in annual spending on primary education across several major categories of inputs including school infrastructure, teacher quality, student-teacher ratios, and school feeding programs.1 However, the public education system in India also faces substantial governance challenges that may limit the extent to which this additional spending translates into improved education outcomes. Our indicator of systemic inefficiency - teacher absence - presents a particularly striking indicator of weak governance. A nationally-representative study of over 3000 public primary schools across 19 major Indian states found that over 25% of teachers were absent from work on a typical working day in 2003 (Kremer et al., 2005). Although administrative data from the government's official records suggest that SSA has led to an improvement in various input-based measures of school quality, there is little evidence on whether these investments have translated into improvements in education system performance, both with respect to intermediate metrics such as teacher absence and final outcomes such as test scores.2 Our study of this nationwide campaign to improve school quality in India uses a new nationally-representative panel dataset of education inputs and outcomes that we collected in 2010. We constructed this dataset by revisiting a randomly-sampled subset of the villages originally surveyed in 2003 (see Kremer et al. (2005)) and collecting detailed data on school facilities, teachers, community participation, monitoring visits by officials, and teacher absence rates. Thus, in addition to reporting updated estimates of teacher absence, and independently-measured summary statistics on input-based measures of school quality, we are able to correlate changes in input-based measures of school quality with changes in teacher absence. The panel data help mitigate concerns arising from fixed unobserved heterogeneity at the village-level, and let us study how the sharp increases in public education spending over the last decade have affected school quality. We find significant improvements in almost all input-based measures of school quality between 2003 and 2010. The fraction of schools with toilets and electricity more than doubled, and the fraction serving mid-day meals nearly quadrupled. There were significant increases in the fraction of schools with drinking water, libraries, and a paved road nearby. The fraction of teachers with college degrees increased by 41%, and student-teacher ratios (STR) fell by 16%. The fraction of teachers not paid on time fell from 51 to 22%, and the fraction of teachers reporting the existence of teacher recognition programs increased from 50 to 81%. Finally, the frequency of school inspections and parent-teacher association (PTA) meetings increased significantly. However, reductions in teacher absence rates were more modest. The all- India weighted average teacher absence in rural areas fell from 26.3 to 23.6%. 3 While increased teacher hiring brought the STR down from 47 to below 40, the effective STR (ESTR), after accounting for teacher absences was still over 50 (having reduced from 64 in 2003 to 52 in 2010). The variation in teacher absence across states remains high. At one end, several top performing states have teacher absence rates below 15%, while at the other end, the poorest performing state, Jharkhand, has a teacher absence rate of 46%. Our panel-data analysis, where we correlate changes in village-level teacher absence with changes in teacher and school characteristics, and administrative and community-level monitoring, yields two robust correlations. First, reductions in the school-level student- teacher ratio (STR) are correlated with an increase in teacher absence, suggesting that the potential benefits from investing in more teachers and lower STR may be partly offset by an increase in teacher absence. Second, better top-down administrative monitoring is strongly correlated with lower teacher absence. Absence rates were 6.5 percentage points lower in villages with regular public school inspections relative to those without, which is a 25% reduction in overall absence and a 40% decline in unauthorized absence.4 One way to estimate the cost of teacher absence is to calculate the salary cost paid by the government to teachers for days of work that they did not attend. We estimate this fiscal cost to be over $1.5 billion per year, which is around 60% of the entire revenue collected from the special education tax used to fund SSA in 2010.5 Teacher salaries typically account for over 80% of non-capital education spending (Dongre et al., 2014), and the most expensive component of the recently passed Right to Education (RtE) Act in India is a commitment to reduce STR from 40:1 to 30:1, by hiring more teachers at an additional cost of $5 billion/year. Using the most conservative panel-data estimates of the correlations between increased monitoring and reduced teacher absence, we estimate that improving school governance (by hiring more supervisory staff) could be over ten times more cost effective at increasing effective student-teacher ratio (net of teacher absence) than hiring more teachers. These calculations suggest that the marginal returns to investing in an inefficiency-reduction strategy (through better monitoring and governance of the education system) are likely to be much higher than a typical input-augmentation strategy. This paper makes several contributions to the literature on public economics in low-income countries. First, teacher absence is now widely used as a governance indicator in education in low- and middle-income countries.6 We update estimates of teacher absence in rural India from 2003 and show that despite substantial increases in education spending over the last decade, improvements on this key measure of governance have been more modest. While corruption in education spending has been shown to hurt learning outcomes (Ferraz et al., 2012), our results highlight the importance of also focusing on governance issues that lead to significant amounts of ‘passive’ waste and inefficiency on an ongoing annual basis, but may not obtain as much media attention as one-off corruption scandals (Bandiera et al., 2009; World Bank, 2010). Second, the fact that decreases in STRs are correlated with increased teacher absence underscores the importance of distinguishing between average and marginal rates of corruption and waste in public spending. Niehaus and Sukhtankar (2013) propose this terminology in the context of wages paid in a public-works program in India and find that marginal rates of leakage are much higher than average rates. We find the same result in the context of teachers and show that the effective absence rate of the marginal teacher hired is considerably higher than the average absence (because of the increased absence among existing teachers). This result, from a large all- India sample, mirrors smaller-sample experimental findings in multiple settings. Duflo et al. (2015), and Muralidharan and Sundararaman (2013) present experimental evidence (from Kenya and India) showing that provision of an extra teacher to schools led to an increase in the absence rate of existing teachers in both settings. In other words, additional spending on school inputs (of which teacher salaries are the largest component) was correlated with increased inefficiency of spending. Third, improvements in top-down administrative monitoring (inspections) are more strongly correlated with reduced teacher absence than improvements in bottom-up community monitoring (PTA meetings), consistent with experimental evidence on the relative effectiveness of administrative and community audits on reducing corruption in road construction in Indonesia (Olken, 2007). More broadly, a growing body of experimental evidence points to the effectiveness of audits and monitoring (accompanied by rewards or sanctions) in improving the performance of public- sector workers and service providers (including Olken (2007) in Indonesia; Duflo et al. (2012) in India; and Zamboni and Litschig (2016) in Brazil). Our panel-data estimates using data from an “as is” nationwide increase in monitoring of schools provide complementary evidence to smaller-scale experiments and suggest that investing in better governance and monitoring of service providers may be an important component of improving state capacity for service delivery in low-income countries (Besley and Persson, 2009; Muralidharan et al., 2016). Finally, recent research has pointed to ‘misallocation’ of capital and labor in low-income countries as an important contributor to lower total factor productivity (TFP) in these settings (Hsieh and Klenow, 2009), and has also documented that a plausible reason for this misallocation is that ‘management quality’ is poorer in low-income countries, and that public-sector firms are managed especially poorly (Bloom and Van Reenen, 2010). Our results provide a striking example of weak management and misallocation in publicly-produced primary education in India (a sector that accounts for over 3% of GDP in spending). In particular, our estimates suggest that reallocating a portion of the $5 billion/year increase in education spending budgeted for hiring more teachers towards measures focused on reducing teacher absence (for instance, by hiring more supervisory staff) may be a much more cost effective way of increasing effective teacher-student contact time. Thus, misallocation is likely to be a first-order issue in this setting, and reallocating education spending towards better governance may substantially increase TFP in publicly-produced education.7 The rest of this paper is organized as follows: Section 2 discusses our empirical methods and analytical framework. Section 3 reports summary statistics on school inputs and teacher absence. Section 4 presents the cross-sectional and panel regression results. Section 5 discusses the fiscal costs of weak governance and compares the returns to investing in better monitoring with that from hiring more teachers. Section 6 discusses policy implications, and Section 7 concludes.","The nationally-representative sample used for the 2003 surveys, which our current study uses as a base, covered both urban and rural areas across the 19 most populous states of India, except Delhi. This represented over 95% of the country's population. The 2010 sample covered only rural India. The sampling strategy in 2010 aimed to maintain representativeness of the current landscape of schools in rural India, and to maximize the size of the panel. We met these twin objectives by retaining the villages in the original sample to the extent possible, while re-sampling schools from the full universe of schools in these villages in 2010, and conducting the panel analysis at the village level.8 Enumerators first conducted school censuses in each village, from which we sampled up to three schools per village for the absence surveys. During fieldwork, enumerators made three separate visits to each sampled school over a period of 10 months from January–October 2010.9 Data on school infrastructure and accessibility, finances, and teacher demographics were collected once for each school (typically during the first visit, but completed in later visits if necessary), while data on time-varying metrics such as teacher and student attendance and dates of the most recent inspections and PTA meetings were collected in each of the three visits. We also assessed student learning with a test administered to a representative sample of fourth grade students in sampled schools. See Appendix A and Tables A1 – A3 for further details on sampling and construction of the village-level panel data set. Teacher absence was measured by direct physical verification of teacher presence within the first fifteen minutes of a survey visit. Data collected during the school census were used to pre-populate teacher rosters for the sampled schools, so that enumerators could look for teachers and record their attendance and activity immediately after their arrival at the school.10 Once teacher attendance was recorded, all other data were collected using interviews of head teachers and individual teachers.11 We record teachers as absent on a given visit if they were not found anywhere in the school in the first fifteen minutes after enumerators reached a school. We consider all the teachers in the school to be absent if the school was closed during regular working hours on a school day, and respondents near the school did not know why the school was closed or mentioned that the school was closed because no teacher had arrived or they had all left early.12 To be conservative in our measure of absence, we exclude all school closures due to bad weather, school construction/repairs, school functions and alternative uses of school premises (for instance, elections). We also exclude all part-time teachers, teachers who were transferred or reassigned elsewhere, or teachers reportedly on a different shift. We construct a school infrastructure index by adding binary indicators for the presence of drinking water, toilets, electricity, and a library. We construct a remoteness index by taking the average of nine normalized indicators of distance to various amenities including a paved road, bus station, train station, public health facility, private health clinic, university, bank, post-office and Ministry of Education office. A lower score on the remoteness index represents a better connected school. During each survey visit, enumerators referred to written school records to note the date of the most recent school inspection, and the date of the most recent parent-teacher association (PTA) meeting. Average parental education of children in a school is computed from the basic demographic data collected for the sample of fourth- grade students chosen for assessments of learning outcomes. For most of the analysis in this paper, we use the village as our unit of analysis and examine mean village-level indicators of both inputs and outcomes because a large number of new schools had been constructed between 2003 and 2010, including in villages that already had schools. This school construction resulted from a policy designed to improve school access by ensuring that every habitation with over 30 school-age children had a school within a distance of one kilometer. Thus, to ensure that our sample was representative in 2010, and at the same time amenable to panel data analysis relative to 2003, we constructed the panel at the village level, with a new representative sample of schools drawn in the sampled villages.13 All the results reported in this paper are population weighted and are thus representative of the relevant geographic unit (i.e., state or all-India). Changes in inputs ~~~~~~~~~~~~~~~~~ The data show considerable improvements in school inputs between 2003 and 2010 along three broad categories - teacher qualifications and working conditions, school facilities, and monitoring (Table 1 - Panels A–C). The fraction of teachers with a college degree increased from 41 to 58%, the fraction reporting that they were getting paid regularly rose from 49 to 78%, and the fraction reporting the existence of teacher recognition schemes rose from 50 to 81%. The fraction of teachers who report a formal teaching credential fell from 77 to 68%, largely due to a significant increase in the hiring of contract teachers (who are not required to have teaching credentials) in several large states. In our data, the fraction of teachers on a temporary contract or ‘contract teachers' increased from 6 to 30%. School facilities and infrastructure improved on almost every measure. The fraction of schools with toilets and electricity more than doubled (from 40% to 84% for toilets and 22% to 45% for electricity); the fraction of schools with functioning mid-day meal programs nearly quadrupled (from 22% to 79%); the fraction of schools with a library increased by over 35 %(from 51% to 69%), and almost all schools now have access to drinking water (96%). Initiatives outside the education ministry to increase road construction have also led to increased proximity of schools to paved roads increasing the accessibility of schools for teachers who choose to live farther away. Relative to the distribution observed in 2003, a summary index of school infrastructure improved by 0.9 standard deviations.14 We also find improvements in both ‘top-down’ administrative and ‘bottom-up’ community monitoring of schools over this period. The fraction of schools inspected in the three months prior to a survey visit increased from 38% to 56%. The extent of community oversight of schools, measured by the frequency of PTA meetings also increased: The probability that a PTA meeting took place during the three months prior to a survey visit increased from 30% to 45%. Overall, Table 1 (Panel A–C) confirms that the Government of India's increased focus on primary education in the past decade did lead to significant improvements in input-based measures of school quality, as well as administrative and community monitoring. Changes in teacher absence ~~~~~~~~~~~~~~~~~~~~~~~~~~ We now turn to changes in teacher absence. Table 1 (Panel D) shows that the population- weighted national average teacher absence rate for rural India fell from 26.3 percent% to 23.6%, a reduction of 10%. Since students receive reduced teacher attention when teachers are absent, we divide the STR by “1 - teacher absence rate” to obtain the effective student teacher ratio (ESTR). Although the all-India STR had been reduced to below 40 in this period, the effective STR after accounting for teacher absence was still over 52. We present state-level data on teacher absence rates and ESTR for 2003 and 2010 in Table A4.15 Chaudhury et al. (2006) find a strong negative correlation between GDP/capita and teacher absence rates (both across countries and within Indian states). Hence, one way to interpret the magnitude of these changes is to compare them with the expected reduction in teacher absence that may be attributed simply to the economic growth that has taken place in this period. Using a growth accounting (as opposed to causal) framework, we can decompose the change in teacher absence into a component explained by changes in GDP/capita (as a proxy for ‘inputs') and one explained by a change in governance (a proxy for TFP). Cross-sectional estimates from the 2003 data suggest that a 10 percent increase in GDP/capita is associated with a 0.6 percentage point reduction in teacher absence.16 In the period between 2002 and 2010, real GDP/capita in India had grown by 38%. Thus, growth in GDP/capita over this period should have by itself contributed to a reduction in teacher absence of 2.4%. Our estimate of the change in teacher absence rate is exactly in this range, and suggests that the reduction of teacher absence we document is consistent with a proportional increase in ‘inputs' into education, but a limited improvement in TFP in this period. We discuss the policy implications of this result in the conclusion. Stated reasons for absence, teaching activity, and official records ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In cases where a teacher was not found in the school, enumerators asked the head teacher (or senior-most teacher present) for the reason for absence. These stated reasons are summarized in Table 2 (Panel A). Two categories of clearly unauthorized absence (school closure during working hours and no valid reason for absence) account for just under half the cases of teacher absence (48%), which provides a lower bound on the extent of unauthorized absences of 11.3 percentage points. The two other categories of stated absence (authorized leave and official duties) that account for 52% of the observed absence are potentially legitimate but cannot be verified. While head teachers may overstate the extent of official duties to shield absent colleagues, they should have no reason to understate it. We can, therefore, reasonably treat the stated reasons for absence as an upper bound for duty-induced absence. This yields the important finding that one commonly cited reason for teacher absence - namely, that teachers are often asked to perform non-teaching duties such as conducting censuses and monitoring elections - is a very small contributor to the high rate of observed teacher absence. Table 2 - Panel A shows that official non-teaching duties account for less than 1% of observations and under 4% of the cases of teacher absence (these results are unchanged from 2003). In cases where the teacher was present, enumerators recorded the activity that the teacher was engaged in at the point of observation: 53% of teachers on the payroll were found to be actively teaching, and another 4% were coded as passively teaching (defined as minding the class while students do their own work). Just over 19% of teachers were in school but were either not in the classroom or not engaged in any teaching activity while in the classroom (Table 2 - Panel B). Thus a total of 42% of teachers on the payroll were either absent or not teaching at the time of direct observations.17 Finally, enumerators also recorded whether a teacher had been marked as present in the log-books on the day of the visit and also on the previous day, and we see in Table 2 - Panel C that going by these records would suggest a much lower teacher absence rate of 16% using the same day's records, or as low as 10.2% using the previous day's records (this was not collected in 2003).18 These data highlight the importance of measuring teacher absence by direct physical verification as opposed to official records on log books. Correlates of teacher absence in 2010 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 3 presents village-level cross-sectional regressions between indicators of school quality and teacher absence in 2010. Column 1 shows the mean level of each covariate in the sample, columns 2–4 present the coefficients on each indicator in individual regressions with the dependent variable being teacher absence, while columns 5–7 do so in multiple regressions that include all the variables shown in Table 1 as regressors. We first show the regressions with no fixed effects, then with state fixed effects, and finally with district fixed effects. The comparison of results with and without state fixed effects is important for interpretation. Many indicators of school quality vary considerably across states in a manner that is likely to be correlated with other measures of governance and development as well as the history of education investments in these states. On a similar note, while primary education policy is typically made at the state level, there is often important variation across districts within a state based on historical as well as geographical factors (Banerjee and Iyer, 2005; Iyer, 2010). Thus, specifications with district fixed effects that are identified using only within-district variation are least likely to be confounded by omitted variables correlated with historical or geographical factors. However, there may still be important fixed omitted variables across villages (such as the level of interest in education in the community) that are correlated with both measured quality of schools and teachers as well as teacher absence. We therefore present the cross-sectional regressions in Table 3 for completeness and focus our discussion on the village-level panel regressions presented in Table 4. Overall, there are few robust correlations across all specifications except that schools that have been inspected recently have lower rates of absence. One important result in the correlations is that there appears to be no significant relationship between teacher salary and the probability of teacher absence. Since salary data were not collected in the 2003 survey, this variable is not included in the panel analysis below. Correlates of changes in teacher absence between 2003 and 2010 ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Since Eq.(1) differences away fixed unobserved heterogeneity at the village level (and therefore at the state and district level as well), the inclusion of state and district fixed effects in the specification controls for average state and district specific changes over time in both the left-hand and right-hand side variables. Thus our panel results with state and district fixed effects are least likely to be confounded with time- invariant and time-variant omitted variables.19 However, it is also worth noting that such a specification biases us against detecting small effects. First, first-differencing leaves us with less variation in the explanatory variables, which will increase standard errors. Second, to the extent there is measurement error in the explanatory variables, first differencing would also increase the attenuation bias. This is why we focus our discussion and interpretation of the results on the ones that are robustly significant and do not treat lack of evidence of significant effects as strong evidence in favor of null effects. Nevertheless, the results in Table 4 suggest that several plausible narratives for the reasons for teacher absence seen in the cross-sectional data reported in Kremer et al. (2005) are not supported in the panel data regressions. In particular, unlike in Kremer et al. (2005), we find no correlation between changes in school infrastructure or proximity to a paved road and teacher absence. We also find no correlation between changes in teacher professional qualifications or professional conditions (such as regularity of pay) and changes in teacher absence.20 We find two robust relationships in the panel regressions, where we define ‘robust’ as correlations that are significant in both individual and multiple regressions; significant in all three main specifications (no fixed effects, state fixed effects, and district fixed effects) and consistent across all specifications (we cannot reject that the estimates are the same across specifications). We discuss these two results below. Reductions in STR are correlated with increased teacher absence First, villages that saw a reduction in student-teacher ratio (STR) have significantly higher rates of teacher absence. A 10% reduction in STR is correlated with a 0.5% increase in average teacher absence, and these estimates remain stable when we include state and district fixed effects and are unchanged when we include a full set of controls (also measured in changes). Changes in STR reflect changes in enrollment as well as in the number of teachers, and a higher STR may affect teacher absence through both enrollment and number of teachers. First, having more students enrolled may increase the cost to teachers of being absent since there are more students (and parents) who may complain. Second, the most common outcome for students when their teacher is absent is that they are combined with other classes/grades whose teachers are present.21 Thus, having more teachers in the school may make it easier for teachers to be absent (since other teachers can handle their class).22 These correlations should not be interpreted as causal (for instance, student enrolment may decline in response to increased teacher absence), but they are consistent with a causal relationship between increased teacher hiring and increased absence of existing teachers that has been established experimentally in India (Muralidharan and Sundararaman, 2013) and other low-income countries such as Kenya (Duflo et al., 2015). Our results provide complementary evidence and greater external validity to these experimental results, and suggest that the benefits of additional teacher hiring to reduce STR may be attenuated by increased teacher absence (in contexts with weak governance of education systems). Increasing monitoring is correlated with reduced teacher absence The second robust result in the panel data estimates is the strong negative correlation between improved school monitoring and teacher absence. In each of the three visits to a school, enumerators recorded the date of the most recent inspection, and we average across the three visits across all the sampled schools in the village to construct the variable “Probability of being inspected in last 3 months”, which ranges from zero (none of the schools in the village were inspected in the prior three months in any of the three visits) to one (all the schools in the village were inspected in the prior three months in all of the three visits). We find that villages where the probability of inspection in the past three months increased from zero to one had a reduction in average teacher absence of between 6.4 and 8.2 percentage points (a 27–35 percent reduction in teacher absence).23 While these results are based on correlations, we present several pieces of evidence consistent with a causal effect of increased school inspections on reduced teacher absence. First, we look at the categories of stated reasons for absence (official duty, authorized leave, and unauthorized absence), and find that increases in inspection probability are correlated only with reductions in unauthorized teacher absence, but not with reductions in teacher absence due to either official duty or authorized leave (Table 5). Second, we examine the extent to which changes in inspection frequency can be explained by other observable factors, and find that there are no correlations between changes in inspection frequency and changes in other observable measures of school quality that are significant across our three standard specifications (Table A5). Third, we use the technique developed by Altonji et al. (2005) to show that the ratio of unobservable to observable correlates of changes in teacher absence would have to be over a factor of 10 for these results to be completely explained by omitted variables (Table A6). Given the very rich data we have on observable changes in school quality, and the fact that our estimates are unchanged even after including state and district fixed effects, this is unlikely to be the case.24 Finally, these results are also consistent with experimental evidence from India that finds significant reduction in teacher absence in response to improved monitoring and rewards linked to better teacher attendance (Duflo et al., 2012). This experimental study, however, was carried out in a small sample of informal schools in one district in India. Thus, our estimates using nationally-representative panel data of rural public schools across 190 districts provide complementary evidence that improved ‘top down’ administrative monitoring may have a substantial impact on reducing unauthorized teacher absence. In contrast, there is less evidence that increases in ‘bottom up’ monitoring by the community (measured by whether the PTA had met in the past 3 months) are correlated with reductions in teacher absence (Table 4). This is consistent with the experimental results reported in Olken (2007) on the impacts of monitoring corruption in Indonesia. These results should not be interpreted as suggesting that bottom-up monitoring cannot be effective, since it is also likely that they reflect differences in the effective authority over teachers possessed by administrative superiors (high) versus parents (low). PTAs in India typically do not have authority to appoint or retain regular civil-service teachers, and they cannot sanction teachers for absence or non-performance (Banerjee et al., 2010). Inspectors and administrative superiors, on the other hand, possess considerable authority over teachers. Their powers include the ability to demand explanations for absence, to issue verbal or written warnings, to make adverse entries in teachers' performance records, to recommend against a pay increment, to suspend a teacher, and in extreme cases to initiate proceedings to fire a teacher (see Ministry of Education (1964-1966) for a detailed discussion of the design of the Indian school inspection system and the powers it provides inspectors). While it is rare for teachers in India to actually get fired for absence (Kremer et al., 2005), and also true that politically- connected teachers can evade sanctions for absence (De and Dreze, 1999; Kingdon et al., 2014), the teacher service rules include several provisions that make it possible for inspectors to significantly raise the costs of teacher absence and thereby reduce it. A striking recent example of how a motivated school inspector in India was able to reduce teacher absence is provided by Anand (Feb. 19, 2016).25 In interpreting the result on school inspections, it is useful to consider why there might be variation in the frequency of inspections across villages and what this would imply for a causal interpretation. One obvious explanation is that inspectors are more likely to visit more accessible villages, but the data do not support this hypothesis since there is no correlation between changes in the remoteness index and changes in inspection rates (Table A5). District-level interviews on school governance in India suggest two important reasons for the variation in inspection frequency. The first is staffing. Districts are broken down further into administrative blocks, and schools within blocks are organized into clusters. School supervision is typically conducted by “block education officers” and “cluster resource coordinators”. We find that a significant fraction of these posts are often unfilled. For instance, in 19% of the cases (where we have data) even the position of the “District Education Officer (DEO)”, the senior-most education official in a district, was vacant (Centre for Policy Research, 2012).26 Further, there is high turnover in the education administration (the average DEO had a tenure in office of just one year) creating periods when the positions are vacant during transitions. The lack of supervisory staff at the block-level is even more acute, as 32% of these positions were estimated to be vacant in 2010 (the year of our survey) even by an official government report (13th JRM Monitoring Report, 2011). Our interviews suggest that these staffing gaps at the block and cluster level are the most important source of variation in inspection frequency within districts, since blocks and clusters without supervisory staff are much less likely to get inspected. The second source of variation in inspections is the diligence of the concerned supervisory officer. Even if all the positions of supervisory staff were filled, there would be variation in the zealousness with which these officers visited villages/schools, which might lead to some areas being inspected more often than others based on whether they were in the coverage area of a more diligent officer or not. However, since supervisors are typically assigned a coverage area of clusters or blocks that comprise many villages, variation in monitoring frequency that is driven by supervisor-level unobservable characteristics is unlikely to be correlated with other village-level characteristics that are also correlated with absence. Of course, this source of variation has implications for thinking about the likely effectiveness of hiring new supervisory staff (some of whom may be less diligent). We discuss these in Section 5.3. Teacher absence and student learning outcomes ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ While not very precise, the estimates in Table 6 (column 5) suggest that both reductions in log(STR), and reductions in log(1-Absence) matter equally for improved test scores. Column 6 of Table 5 shows that once log(ESTR) is controlled for there is no independent effect of teacher absence on learning outcomes, suggesting that the main mechanism by which teacher absence affects learning outcomes is through increasing the ESTR. The stronger relationship between teacher absence and student learning outcomes seen in columns 2 and 3 (that do not include state or district fixed effects) suggests that teacher absence is likely correlated with other measures of education governance at the state and district levels, and highlights why our preferred specifications are the ones with district fixed effects. Our data, which are collected seven years apart and have only mean village-level test scores, are not ideal for studying the impact of teacher absence or other school characteristics on test scores (the ideal specifications would use annual panel data on student test scores matched to these characteristics and estimate value- added models of student learning). But it allows us to present suggestive evidence on the negative correlations between teacher absence and student learning outcomes that are consistent with other studies using better data that find similar results.28 The results in Table 6 also help illustrate that teacher absence can attenuate the benefits of reducing STR, and that reducing effective STR can be done both by reducing STR and by reducing teacher absence. We consider the relative cost effectiveness of these approaches in the next section. The fiscal cost of teacher absence ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ High levels of teacher absence translate into considerable waste of public funds since teacher salaries are the largest component of education spending in most countries, including India. One way of estimating these costs is to calculate the total salary cost paid to teachers for days of work that they were expected to attend, but do not. Note that this is not a cost that would be saved if teacher absence were to be reduced (since the full teacher salaries would be paid in either case). However, it is standard in the corruption literature to measure the cost of corruption by the amount of public expenditure that does not reach its intended goal (often referred to as ‘leakage’), and to measure the impact of interventions to reduce corruption by quantifying the reduction in leakage, even if there is no reduction in fiscal outlay (Reinikka and Svensson, 2004;Reinikka and Svensson, 2005;Niehaus and Sukhtankar, 2013;Muralidharan et al., 2016).29 We follow a similar approach here by first quantifying the salary cost of absence as an estimate of ‘leakage’ in education spending, and then using these costs as the metric to evaluate alternate policy approaches to reducing ESTR. Calculating the cost of teacher absence requires us to estimate and exclude the extent of legitimate absence from our calculations. As part of the institutional background work for this project, we obtained teacher policy documents from several states across India. Analysis of these documents indicates that the annual allowance for personal and sick leave is 5% on average across states. This is close to the survey estimate of 5.9%(Table 2), but we use the official data since the stated reasons may be over-reported. Estimating the extent of legitimate absence due to ‘official duty’ (outside the school) is more difficult because there are no standard figures for the ‘expected’ level of teacher absence for official duties. Policy norms prescribe minimal disruption to teachers during the school day and stipulate that meetings and trainings be carried out on non-school days or outside school hours. Since we are not able to verify the claim that teachers were on official duty, and there is evidence that head teachers try to cover up for teacher absences by claiming that these are due to ‘official duties', our default estimate treats half of these cases as legitimate. This gives us a base case of legitimate absence of 8%(5% authorized leave, and 3% official duty). We also consider a more conservative case where the legitimate rate of absence is 10%. This 8–10% range of legitimate absence also makes sense because the fraction of teacher observations that are classified as either ‘authorized leave’ or ‘official duty’ is in this range for the five states with the lowest overall absence rates - even treating the stated reasons for absence as being fully true (tables available on request). To estimate the cost of teacher absence, we use teacher salary data from our surveys and use administrative (DISE) data on the number of primary school teachers by state.30 We provide three estimates of the fiscal cost of teacher absence based on assuming the rate of legitimate teacher absence to be 8, 9, and 10 percent% respectively, and these calculations suggest that the annual fiscal cost of teacher absence is around Rs.81 –93 billion, which is around US$1.4–1.6 billion/year at 2010 exchange rates (Table 7 - Panel A). Calculating the returns to better governance in education ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Using the results in Table 4, we calculate the returns to a marginal increase in the probability of a school being inspected. We make the following assumptions: (a) enough supervisory staff are hired to increase the probability of a school being inspected in the past 3 months by 10 percentage points (relative to a current probability of 56%); (b) increasing inspection probability by 10 percentage points would reduce mean teacher absence across the schools in a village by 0.64 percentage points (the most conservative estimate of the correlation between increased inspection probability and reduced teacher absence from Table 4); (c) the full cost (salary and travel) of a supervisor is 2.8 times that of a teacher; (d) a supervisor works 200 days per year and can cover 2 schools per day.31 The results of this estimation are presented in Table 7 (Panel B) and we see that the cost of hiring enough supervisors to increase the probability of a school being inspected by 10 percentage points is Rs.448 million/year (see Table A8 for state-level calculations). However, the reduction in wasted salary from this investment in terms of reduced teacher absence amounts to Rs.4.5 billion/year, suggesting that investing in better monitoring would lead to a reduction in ‘leakage’ of teacher salaries (defined as salary payments for days when teachers do not attend work) that is around ten times greater than the cost of increasing monitoring by hiring more supervisory staff. Input augmentation versus inefficiency reduction ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ To compare the relative cost effectiveness of hiring more teachers (input augmentation) versus hiring more supervisors to reduce teacher absence (inefficiency reduction) as a way of reducing the ESTR, we calculate the salary cost of hiring more teachers to achieve the same reduction in ESTR that we estimate would be obtained by increasing the inspection probability by 10 percentage points. We estimate this to be Rs. 5.7 billion/year (Table 7 - Panel B; Table A8 provides detailed state-level calculations), and see that increasing the probability of inspection would be 12.8 times more cost effective at reducing ESTR than doing so by hiring more teachers (on the current margin).32 The difference in the relative cost effectiveness of the two policy options is large enough that hiring more supervisors rather than teachers is likely to be a more cost effective way of reducing ESTR (on the current margin) even if the supervisors were to work less efficiently than assumed in these calculations. For instance, if supervisors were absent at the same rate as teachers (say 25 %), allocating marginal funds to hire an additional supervisor would still be nearly ten times more cost effective at reducing ESTR than using those funds to hire an additional teacher.33","The main caveat to using our results to recommend a universal policy of hiring more supervisors to scale up the frequency of school inspections is that our estimates are based on correlations and may not be convincing enough to warrant a universal scale up. Nevertheless, it is worth noting that both our key results - the correlation between increased monitoring and reduced teacher absence, and the correlation between lower STR and increased teacher absence - are consistent with experimental evidence from smaller- scale, which increases our confidence in their validity. Further, our estimates are based on an expansion of existing system of inspections, and use nationwide panel data (which mitigates omitted variables concerns) representing close to a billion people, and complement results from smaller-scale randomized experiments warranting them greater external validity for several reasons. First, while our results support results from smaller randomized experiments, there is evidence that experimentally-estimated positive results of interventions that are implemented by NGOs may not be replicated when the programs are implemented by governments (Banerjee et al., 2008). Second, there is also evidence of site-selection bias where implementing partners are more likely to be willing to rigorously evaluate programs in locations where they are more likely to be successful (Allcott, 2015). Finally, even in the absence of such a bias, most experiments are conducted in very few sites, and may yield imprecise treatment effects (for inference over a larger population) in a setting where unobserved site-specific covariates may interact with the treatment (Pritchett and Sandefur, 2013).34 Thus, even if small-scale experiments are unbiased within sample, they may be biased and also imprecise for population-level inference. In other words, there is likely to be a trade-off between the potential omitted variable bias in our panel-data estimates on one hand, and the advantages of greater precision, “as is” implementation, and unbiased site selection on the other. We do not attempt to quantify this trade-off in this paper since we have no objective basis of doing so. However, one way of reconciling this trade-off is to conduct a substantial nationwide expansion of school inspections by hiring more staff in the context of a large experimental evaluation. From a decision-theoretic perspective, our results are strong enough to support such a policy even if there is only a 1% chance that our estimates are causal. In Appendix B, we formally show that, barring extreme priors, a policy-maker interested in lowering effective student-teacher ratio will find it cost-effective to invest in or scale-up monitoring of teachers.","The central and state governments in India have considerably increased spending on primary education over the past decade. We contribute towards understanding the impact of these substantial nationwide investments in primary education in India by constructing a unique nationally-representative panel data set on education quality in rural India. We find that there has been a substantial improvement in several measures of school quality including infrastructure, student-teacher ratios, and monitoring. However, teacher absence rates continue to be high, with 23.6% of teachers in public schools across rural India being absent during unannounced visits to schools. Using village-level panel data, we find two robust correlations in the panel data that provide external validity in nationally- representative data to results established in smaller-scale experiments. First, reductions in student-teacher ratios are strongly correlated with increased teacher absence, suggesting that increased spending on hiring additional teachers was accompanied by increased inefficiency, which may limit the extent to which additional spending may improve outcomes. Second, increases in the frequency of inspections are strongly correlated with lower teacher absence, suggesting that of all the investments in improving school quality, the one that was most effective in reducing teacher absence was improved administrative monitoring of schools and teachers. We calculate that the fiscal cost of teacher absence is over $1.5 billion per year, and estimate that investing in improved governance by increasing the frequency of monitoring would be over ten times more cost effective at increasing student-teacher contact time than doing so by hiring additional teachers. In interpreting our results, it may be useful to think of the performance of the education system (measured by the level of teacher absence) as comprising two components - ‘inputs' into the production of education that expand with income growth (such as school infrastructure, class size, and teacher salaries), and the efficiency of the use of these inputs (which would correspond to the TFP of education production). Our results show that the Indian education system has made significant progress on the former, but made less progress on the latter. They also suggest that pivoting public expenditure away from simply augmenting inputs towards policies that increase the efficiency of inputs may considerably increase the productivity of education spending, and thereby enable achievement of improved human capital outcomes at any given level of per-capita income. One promising way of reducing inefficiency is improving school governance and achieving such a reallocation of resources would be to expand the existing system of administrative monitoring of teachers and schools by hiring more supervisory staff. Our calculations indicate that such a marginal expansion could (on the current margin) have a significant impact on reducing teacher absence, and that this would be highly cost effective in terms of reducing the fiscal cost of weak governance. More broadly, our results suggest that the returns to investing in state capacity to better monitor the implementation of social programs in low-income countries may be quite high, and that at the very least there is a strong case for expanding such programs in the context of large experimental evaluations of “as is” implementation to obtain more precise estimates of their benefits.35"],["This paper proposes a geometric delineation of distributional preference types and a non-parametric approach for their identification in a two-person context. It starts with a small set of assumptions on preferences and shows that this set (i) naturally results in a taxonomy of distributional archetypes that nests all empirically relevant types considered in previous work; and (ii) gives rise to a clean experimental identification procedure - the Equality Equivalence Test - that discriminates between archetypes according to core features of preferences rather than properties of specific modeling variants. As a by-product the test yields a two-dimensional index of preference intensity. --------------------------------------------------------------------------------","Many economists׳ default assumption is that all agents are exclusively motivated by their own material self-interest. This assumption is in sharp contrast to both day-to-day experience and empirical evidence gathered by psychologists and experimental economists in the last decades. This has aroused renewed interest in theories of other-regarding preferences, where arguments beyond material self-interest enter the decision maker׳s utility function.1 Typical examples of such arguments are other people׳s (material) well- being (as in distributional preferences models),2 others׳ opportunities and expected or observed behavior (as in reciprocity models),3 others׳ payoff expectations (as in guilt aversion models),4 or others׳ other-regarding concerns (as in type based models).5 The present paper focuses on the first of the above mentioned subclasses, i.e. on distributional (or “social”) preferences, where besides one׳s own material payoff the (material) well-being of others enters an agent׳s utility function. Distributional preferences have been shown to be behaviorally relevant in important market and non-market environments – see Sobel (2005) and Fehr and Schmidt (2006) for excellent surveys. The current paper adds to this literature by proposing (i) a simple classification of distributional preference types that nests almost all major classifications of archetypes discussed in the economic and the social psychology literature; and (ii) a simple identification procedure based on the classification. Identification of distributional preferences has been the topic of numerous papers, of course – see Kerschbamer (2013) for a thorough review of the literature. These pioneering studies – which have greatly advanced our understanding of non-selfish behavior – suffer from at least one of two methodological shortcomings. First, the tests employed typically discriminate between the members of a somewhat arbitrary list of distributional types; and secondly, the identification procedures typically rely on strong structural assumptions.6 Regarding the former dimension – the set of distributional types tested for – previous studies either start with a given list of types, or they employ a test design that allows discriminating only between the members of a limited set of types.7 For instance, the path-breaking dictator-game study by Andreoni and Miller (2002) distinguishes between selfish, Leontief and perfect substitutes preferences, plus weak incarnations of those types; the follow-up study by Fisman et al. (2007) employs a richer design and discriminates between self- interested, lexself (lexicographic for self over other), social welfare and competitive types plus some mixes thereof; the pioneering discrete choice study by Engelmann and Strobel (2004) tries to disentangle efficiency concerns (defined as surplus maximization), maximin preferences and (two modeling variants of) inequality aversion; Blanco et al. (2011) discriminate between selfish and various intensities of piecewise linear inequality aversion; Charness and Rabin (2002), Cabrales et al. (2010) and Iriberri and Rey-Biel (2013) allow for self-interested, social welfare, difference-averse and competitive preferences; and the ring-test – originally developed by social psychologists to assess “social value orientations”8 and recently used by economists to identify type and intensity of distributional concerns9 – discriminates between altruists, cooperators, individualists, competititors, aggressors, martyrs, masochists and sadomasochists. Turning to the second dimension – the structural assumptions imposed – the identification procedures employed in previous studies typically rely on strong assumptions regarding the form of the utility or motivational function meant to represent preferences. For instance, the ring-test is based on the assumption of linear preferences; the studies by Cabrales et al. (2010), Blanco et al. (2011) and Iriberri and Rey-Biel (2013) employ identification procedures based on the piecewise linear model originally introduced by Fehr and Schmidt (1999) as a description of self-centered inequality aversion and later extended by Charness and Rabin (2002) to allow for other forms of distributional concerns and thereby assume piecewise linearity; and Andreoni and Miller (2002), Fisman et al. (2007) and Cox and Sadiraj (2012) check consistency with – and estimate parameters of – standard or modified constant elasticity of substitution (CES) utility functions. Summing up the above discussion we conclude (i) that there is neither an agreement in the literature on what the relevant set of distributional basic motivations – defined as the manner in which people care about the (material) well-being of others – is, nor on how to delimitate distributional types; and (ii) that existing studies employ identification procedures that rely on strong structural assumptions as, for instance, linearity, piecewise linearity or standard or modified CES forms. By using a systematic approach based on a small set of primitive assumptions on preferences, the present paper offers an improvement in both dimensions. It shows (i) that this set of assumptions naturally results in a well delineated, mutually exclusive and comprehensive distinction between nine archetypes of distributional concerns; and (ii) that this set gives rise to a simple non-parametric experimental test – the Equality Equivalence Test (EET) – that discriminates between the archetypes according to core features of preferences rather than properties of specific modeling variants or functional forms. As a byproduct the test yields a two-dimensional index of preference intensity. While the primary purpose of this paper is methodological, the experimental results obtained in an implementation of the EET also produce some substantive insights. For instance, the result that – consistent with the theoretically appealing assumption that distributional preferences are convex – about 95% of the subjects reveal (weakly) more benevolent (less malevolent) preferences in the domain of advantageous than in the domain of disadvantageous inequality. A second interesting detail is that beyond selfish subjects, the empirically most frequent distributional archetypes are those who exhibit (at least weakly) positive attitudes towards others in both domains (i.e., altruism and maximin), while archetypes that imply a negative attitude in at least one of the domains are by far less important empirically (the behavior of less than a fourth of the subjects is consistent with any form of inequality aversion, for instance, and the choices of less than 7% of the subject population are consistent with spite).10 The rest of the paper is organized as follows: Section 2 presents the assumptions on which the analysis is based and argues that those assumptions are fulfilled by all major modeling variants of distributional preferences discussed in the economic and the social psychology literature. Section 3 introduces the proposed classification of preference types based on the rate an agent is willing to trade between own monetary payoff and the monetary payoff of another. Section 4 presents the proposed identification procedure – the “Equality Equivalence Test” (EET). It starts (in Section 4.1) by conveying the intuition behind the proposed identification approach and explaining its similarity to the Certainty Equivalence Test. Section 4.2 presents the symmetric basic version of the test, and Section 4.3 discusses several extensions. In Section 4.4 a two-dimensional index for identifying the archetype and characterizing the intensity of distributional concerns – the (x, y)-score – is introduced, and a graphical representation of the type–intensity distribution is proposed. Section 4.5 relates the (x, y)-score to other measures of type and intensity of distributional concerns. Section 5 illustrates the working of the EET by reporting experimental results generated with the symmetric basic version of the test, and Section 6 concludes. Implementation issues for the case where the test is used as a tool in experimental economics (to address research questions in which distributional preferences are expected to shape behavior, to control for subject pool effects, or to help to interpret data from other unrelated experiments) are discussed in Appendix A. Appendix B contains the instructions of the experiment reported in Section 5.","Let a=(m, o) denote an income allocation that gives material payoff m (for “my”) to the decision maker (DM or “agent”) and material payoff o (for “other”) to the other person. The space of feasible income allocations is assumed to be the non-negative orthant of R2 and is denoted by A. Throughout we assume that the DM is equipped with a preference relation over income allocations, which we denote by ≽. Technically, ≽ is a binary relation on A, allowing the DM to compare pairs of allocations a, a*∈A. We read a≽a* as “the DM weakly prefers allocation a to allocation a*” and denote the asymmetric and the symmetric part of ≽ by ≻ and ~, respectively.11 For the DM’s preferences we require: (completeness, transitivity and continuity): The DM’s preference relation on income allocations is complete, transitive and continuous. That is, for ≽ it holds that: for every pair a, a′∈A, either a≽a′, or a′≽a (or both); for every triple a, a′, a*∈A, if a≽a′ and a′≽a*, then a≽a*; for every two sequences a1, a2, a3,… and a′1, a′2, a′3,… in A, if the sequence a1, a2, a3,… converges to a and the sequence a′1, a′2, a′3,… converges to a′, and if ai≽a′i for each i, then a≽a′. Completeness (i.e., the first part of Assumption 1) requires that the DM can compare any two income allocations; transitivity (the second part) adds the requirement that the preferences of the DM are internally consistent; and continuity (the last part) says that the DM′s preferences do not exhibit “jumps”, with, for example, the DM preferring each element in the sequence a1, a2, a3,… to a′, but suddenly reversing her preferences at the limiting point of the sequence. While ordering (completeness and transitivity) is important for the arguments below (as it is for substantial parts of economic theory), continuity is not.12 As shown by Eilenberg (1941) the three parts of Assumption 1 together imply that the DM′s preferences can be summarized by means of a continuous utility or motivational function u(m, o) that assigns a real- valued index to every (m, o)∈A. (strict m-monotonicity): The DM′s preference relation on income allocations is strictly monotonic in the own material payoff. That is, comparing any two income allocations (m, o) and (m′, o) in A with the same level of o, (m, o)≻(m′, o)⇔m>m′ and (m, o)~(m′, o)⇔m=m′. Strict m-monotonicity requires that – holding the material payoff of the other person constant – the DM strictly prefers more own material payoff to less own material payoff. This is quite a natural assumption. It is violated, for instance, if the DM is willing to burn her own monetary payoff because she feels bad whenever she has (much) more than the other person. Such behavior is essentially never observed in experiments. In terms of utility representation, Assumption 2 translates to the requirement that for every (m, o)∈A and ∆∈R++ we have u(m+∆, o)>u(m, o). (piecewise o-monotonicity): The DM′s preference relation between two income allocations that have the same own material payoff for the DM but different payoffs for the other person depends only on whether the DM is ahead or behind. That is, comparing any two income allocations (m, o) and (m, o′) in A with the same level of m and o<o′, the DM′s preference relationship between (m, o) and (m, o′) (i.e., whether ≻, ≺, or ~ holds) is constant for all o, o′, m such that o>m and is also constant for all o, o′ m such that m>o′ (but potentially different between the two domains). Piecewise o-monotonicity requires that the DM′s general attitude towards the other person (i.e., whether she is benevolent, neutral, or malevolent to the other) depends only on whether the other person has more or less monetary payoff than the DM herself. In terms of utility representation, it translates to the requirement that for every ∆∈R++ the sign of the difference u(m, o+∆)−u(m, o) is constant for all (m, o)∈A with o>m and is also constant for all (m, o)∈A with o+∆<m (but potentially different between the two domains). Piecewise o-monotonicity is both permissive and restrictive, depending on the perspective. It is permissive because it allows for all major variants of distributional preferences that have been discussed in the economic literature – see the discussion at the end of this and in the next section. Piecewise o-monotonicity is also restrictive because it implies (i) that preferences only depend on monetary outcomes, not on the way they are achieved (this is the defining feature of distributional preferences); and (ii) that the reference point for the evaluation of allocations (if one is used) is an equal-material-payoffs allocation. Ad (i): The implication that preferences only depend on monetary outcomes is likely to be violated in many important applications. For instance, in strategic interactions (where the other person has an opportunity to move and thereby a possibility to influence the payoff of the DM) beliefs about intentions behind observed or expected action choices of the other person potentially play a role (see the literature on reciprocity and related concepts cited in Footnote 2). Also, in some games beliefs about the payoff expectations of the other person seem to influence behavior (see the literature on guilt aversion and related concepts cited in Footnote 3). Furthermore, in a richer environment, where agents have more information on each other, beliefs about the other-regarding concerns of the other person may play a role (as in the literature on type-based models cited in Footnote 4). Finally, features of the situation (such as context, entitlements, properties of the outcome generating process, etc.) or the DM (such as a code of conduct, or a preference for honesty) might shape behavior. Knowing that all those factors might be behaviorally relevant in a richer environment, it seems important that distributional preferences are identified in a non-strategic setting and a neutral frame to avoid confounds. This is not to say that distributional preferences are unimportant in richer environments, of course, but rather that they cannot be unambiguously identified there. Ad (ii): Some distributional archetypes discussed in real life and in the literature (most importantly, inequality aversion and egalitarian motives; maximin, Rawlsian and Leontief preferences; and envy) are inevitably defined in terms of a “reference location”, where the DM׳s general attitude towards the other changes from preferring higher payoffs for the other to preferring lower payoffs. In theory, this reference location can be anything (an interval, a point, or whatsoever), and it can differ among individuals. In existing models of reference-dependent distributional concerns, the reference location is a point, and the point is the egalitarian one for all individuals (see, for instance, Bolton, 1991; Mui, 1995; Fehr and Schmidt, 1999; Bolton and Ockenfels, 2000; or Charness and Rabin, 2002). While Assumption 3 is more agnostic than existing models of reference-dependent distributional concerns, it is still restrictive.13 For instance, there might exist individuals who consider it fair to get 20% more than others but unfair to get 30% more. Assumption 3 does not allow for this. While it would be feasible, in principle, to generalize Assumption 3 (and the test relying on it) so as to allow for heterogeneous reference points, this would seriously impair simplicity and transparency: Ultimately the aim of the paper is to propose a classification of subjects in distributional preference types that is helpful in organizing experimental data. For that purpose we need some kind of clustering and not a different distributional type for each single individual. Stated differently, as any model the approach proposed here is by design an abstraction of reality, and hence is deliberately constructed so as to not explain some behavior, in return for parsimony. While parsimony calls for a unique reference point, it does not suggest equality as the reference point. Equality is suggested by normative considerations and by empirical evidence. The normative basis of equality as a reference point is discussed in some detail in Konow (1993) and in the working paper version of this article (Kerschbamer, 2013). Regarding empirical evidence Andreoni and Bernheim (2009, p. 1607f) cite several studies showing that equal sharing is common in the context of joint ventures among business firms, partnerships among professionals, share tenancy in agriculture, and bequests to children. They also provide evidence indicating that equality is a frequent outcome of negotiation and conventional arbitration in the field. In lab-experiments the assumption that the egalitarian outcome is somehow focal among subjects who change their general attitude towards others at some point seems even more natural than in the field: Subjects enter the laboratory as equals, their roles are assigned randomly and they have absolutely no information about each other. It seems therefore quite plausible that those subjects who attribute special meaning to an allocation (again, nothing in Assumption 3 requires them to do so) do this to the egalitarian one. And there is indeed considerable support for this assumption in existing experimental data. For instance, one of the stylized facts in standard dictator games is precisely that a sizeable fraction of the subject population voluntarily cedes exactly half of the pie to the recipient, and that very few subjects cede more (Camerer, 1997). This result survives even in experiments where the action space is continuous and where the price for giving is quite high (see Andreoni and Miller, 2002, for instance). The frequency of equal divisions is even higher in ultimatum games, where expectations about the “reference point” of the recipient enter the picture (see Camerer, 2003). While all this evidence indicates that the egalitarian outcome has something special for a substantial fraction of subjects, it does not tell us anything about the exact fraction of subjects for whom this is the case.14 But this is exactly (one of) the question(s) the proposed test aims to address. (strict equal- material-payoff-monotonicity): The DM′s preference relation on income allocations is strictly monotonic in both payoffs along the ray m=o. That is, comparing any two income allocations (m, o) and (m′, o′) in A with m=o and m′=o′, (m, o)≻(m′, o′)⇔m>m′ and (m, o)~(m′, o′)⇔m=m′. Strict equal-material-payoff-monotonicity requires that more preferred allocations are reached when the payoffs of both agents are increased along the 45° line. In terms of utility representation it translates to the requirement that for every z∈R+ and ∆∈R++ we have u(z+∆, z+∆)>u(z, z). In combination with strict m-monotonicity, strict equal-material-payoff-monotonicity essentially rules out extreme forms of spite by putting an upper bound on the malevolence of the DM along the ray m=o. As is easily checked, almost all (modeling) variants of distributional preferences discussed in the economics literature satisfy Assumptions 1–4, notable exceptions being lexself preferences (discussed by Fisman et al., 2007) which – in a strict interpretation – violate the continuity part of Assumption 1, and maximin (or Rawlsian, or Leontief) preferences (discussed by Andreoni and Miller, 2002; Charness and Rabin, 2002; or Engelmann and Strobel, 2004, for instance) which – in their purest form (but not in the form typically discussed in the literature) – violate strict m-monotonicity.15","This section introduces a simple graphical classification of distributional preferences based on the four assumptions introduced in the previous section. Referring to Fig. 1, the preference of a DM is classified by characterizing the indifference curve that runs through the reference point r=(e, e). The choice space is divided into six relevant subsets, {x1, x2, x3} and {y1, y2, y3}. Here, y1 is the area below the 45° line and to the left of the vertical line through the reference point, y2 is the section of the vertical line that lies below the reference point, and y3 is the area to the right of the vertical line and below the horizontal line through the reference point. The subsets {x1, x2, x3} are defined similarly. Note that Assumptions 2 and 4 together imply that the indifference curve that runs through the reference point cannot pass through any of the two shaded areas in Fig. 1. The preference type of the DM is now classified by the subsets that contain the DM’s indifference curve that runs through r=(e, e). Given Assumptions 1–4, one section of the indifference curve necessarily runs through one (and only one) of the x subsets, while the other section necessarily runs through one (and only one) of the y subsets. Therefore, it is simple to see that there are nine possible archetypes of distributional preferences given the proposed division in subsets. The nine archetypes are defined in Table 1 and a typical indifference curve of each archetype is displayed in Fig. 2. Let me shortly discuss the core features of different distributional preference types rattling around in the literature and how they fit into the proposed template. First consider selfish or own-money-maximizing preferences. They can be considered as a degenerated version of distributional preferences where an agent’s well-being neither increases nor decreases in the monetary payoffs of other agents. Thus, the core property of selfish preferences in a two-person context is that indifference curves in (m, o) space are vertical. Referring to Fig. 1 this means that a selfish DM’s indifference curve through r=(e, e) must run through the subsets x2 and y2 (as indicated in Table 1). The well-being of an altruistic agent increases in the monetary or utility payoffs of other agents (Becker, 1974; Andreoni and Miller, 2002); the well-being of an efficiency loving or surplus maximizing agent (Engelmann and Strobel, 2004), the well-being of an agent with perfect substitutes preferences (Andreoni and Miller, 2002) and the well-being of an agent with social welfare preferences (Charness and Rabin, 2002; Fisman et al., 2007) increases in the (weighted or unweighted) sum of payoffs. In all cases, well-being increases in o everywhere. Thus, indifference curves in (m, o) space are negatively sloped everywhere (if o increases m has to decrease to hold the agent indifferent) meaning that (in terms of Fig. 1) the indifference curve of an altruistic DM must pass through x1 and y3. An agent is spiteful (Levine, 1998), or competitive (Charness and Rabin, 2002), or status seeking or interested in relative income (Duesenberry, 1949), if her well-being decreases in the payoffs of others everywhere; so the core property of such preferences is positively sloped indifference curves in (m, o) space. In terms of Fig. 1 this means that a spiteful DM’s indifference curve through r=(e, e) must run through the subsets x3 and y1. The well- being of an envious or grudging agent decreases in the payoffs of agents who have more, but is unaffected by the payoffs of agents who have less (the role of envy has been emphasized by Bolton, 1991 and Mui, 1995, for instance); thus, the core property of envious preferences is positively sloped indifference curves in the domain of disadvantageous inequality and vertical indifference curves in the domain of advantageous inequality (yielding the combination x3, y2).16 The well-being of an agent with maximin preferences (Engelmann and Strobel, 2004), Rawlsian preferences (Charness and Rabin, 2002), or Leontief preferences (Andreoni and Miller, 2002; Fisman et al., 2007) increases in the lowest of all agents’ payoffs. Thus, its defining feature in a two-person context is that indifference curves in (m, o) space are negatively sloped if inequality is advantageous and vertical otherwise (yielding the combination x2, y3). An agent is inequity or inequality averse (Fehr and Schmidt, 1999; Bolton and Ockenfels, 2000), or difference averse (Charness and Rabin, 2002; Fisman et al., 2007), or egalitarian (Dawes et al., 2007; Fehr et al., 2008) if she incurs a disutility when other agents have either higher or lower payoffs (as in the model by Fehr and Schmidt, 1999), or when the agent’s payoff differs from the average payoff of all agents (as in Bolton and Ockenfels, 2000). Consequently, the defining feature of inequality averse or egalitarian preferences in a two-person context is negatively sloped indifference curves in the domain of advantageous and positively sloped indifference curves in the domain of disadvantageous inequality (yielding the combination x3, y3). The opposite constellation, benevolence in the domain of disadvantageous inequality combined with malevolence in the domain of advantageous inequality, is referred to as equality aversion (by Hennig-Schmidt, 2002, for instance), or as equity aversion (e.g. by Charness and Rabin, 2002 and by Fershtman et al., 2012). Its defining feature in a two-person context is that indifference curves in (m, o) space are positively sloped below and negatively sloped above the 45° line (translating to x1, y1). Table 1 lists and Fig. 2 displays two further archetypes of distributional preferences, “kick down” and “kiss up”. Those types have not been discussed in the literature and are included for completeness only: Kick-down or bully-the-underlings preferences imply malevolence towards agents who have lower and neutrality towards agents who have higher payoffs. Thus, the defining feature of such preferences in a two-person context is that indifference curves in (m, o) space are positively sloped in the domain of advantageous inequality and vertical in the domain of disadvantageous inequality (implying the combination x2, y1).17 The opposite constellation, benevolence towards agents who are better off combined with neutrality towards those who are worse off, is called kiss-up or crawl-to-the-bigwigs preferences and such preferences imply negatively sloped indifference curves in the domain of disadvantageous inequality and vertical indifference curves in the domain of advantageous inequality (implying the combination x1, y2). Note that the nine types listed in Table 1 and displayed in Fig. 2 are well delimitated, mutually exclusive and comprehensive. Also note how the four basic assumptions introduced earlier enter the picture: ordering and continuity translate into existence and uniqueness of indifference curves through any point in (m, o) space; strict m-monotonicity means that upper contour sets are to the right of an indifference curve (the arrows in Fig. 2); piecewise o-monotonicity requires that the general attitude of the DM (i.e., whether she is benevolent, neutral or malevolent) changes at most once – when crossing the equal- material-payoff line; and strict equal-material-payoff-monotonicity excludes indifference curves that fall on only one side of equal-material-payoff line. Thus, Assumptions 1–4 together naturally result in the distinction between the nine mutually exclusive and comprehensive archetypes listed in Table 1 and displayed in Fig. 2, meaning that qualitatively there is no room left for additional types.18 Before proceeding it seems important to address the potential critique that the nine archetypes defined here are not really new. This is correct, of course. The main contribution of the present paper is not to introduce new preference types; one of the goals is rather to derive the number and core properties of preference types from a small set of primitive assumptions on preferences. This stands in contrast to previous studies which either start with a given list of types or a specific model of preferences. A second – related – critique is that a list of archetypes similar to the one presented in Table 1 could also be obtained by working off the possible sign combinations of the two parameters in the piecewise linear model originally introduced by Fehr and Schmidt (1999) as a description of self-centered inequality aversion and later extended by Charness and Rabin (2002) to allow for other forms of distributional concerns. If one is willing to assume that subjects have preferences of this very specific form then this critique is justified. However, a major point in the current paper is exactly that there is no need to impose such a tight structure. This is true both for the type delineation introduced in this and the elicitation procedure proposed in the next section. Stated differently, all modeling variants of distributional preferences satisfying the four assumptions introduced in Section 2 and all distributional archetypes tested for in previous experiments fall into one of the nine categories defined here. This is also true for the Charness and Rabin model. On the other hand, there are many models of distributional preferences in the economic literature that do not fit into the piecewise linear framework of Charness and Rabin – the altruism models by Andreoni and Miller (2002), Cox et al. (2007) and Cox and Sadiraj (2012), the envy model by Bolton (1991), and the inequality aversion model by Bolton and Ockenfels (2000) are prominent examples. Idea of the Equality Equivalence Test ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As mentioned earlier, the four basic assumptions introduced in Section 2 not only naturally result in a classification of distributional preference types that nests all major behavioral types discussed in the literature, but also give rise to a clean identification procedure (a “test”) that does not rely on unnecessary structural assumptions. This subsection explains how the test works and motivates its name (Equality Equivalence Test). Given Assumptions 1–4, the DM’s type can be determined by identifying the location of the two sections of her indifference curve through the reference allocation r=(e, e), the section that passes the domain of disadvantageous inequality (the area above and to the left of the 45° line through the reference point) and the section that passes the domain of advantageous inequality (the area below and to the right of the 45° line). Theoretically, this can be done by exposing the DM to only four binary choices. Take points r, p1 and p2 in Fig. 3. Suppose we ask the DM to decide subsequently between p1 and r and between p2 and r. If the DM decides for the p allocation in both choices then she reveals p1≽r and p2≽r; thus, for the domain of disadvantageous inequality her indifference curve through r=(e, e) must run through x1. Similarly, if the DM reveals r≽p1 and p2≽r (by deciding for r in the first binary choice and for p2 in the second) then her indifference curve is in x2.19 And if the DM reveals r≽p1 and r≽p2 (by deciding for r in both choices) then her indifference curve is in x3.20 By exposing the DM in addition to binary choices between r=(e, e) and two points on the horizontal line below r (one to the left and one to the right of the vertical line through r) the location of the second part of her indifference curve through r – that is, the part that lies below the 45° line – can be determined. This is the idea behind the EET. Note that the test proposed here is in many respects similar to the Certainty Equivalence Test (CET) used in experimental economics (and beyond) as a means to elicit risk attitudes (see Dohmen et al., 2010 for a recent application). With both procedures the DM is exposed to a short sequence of binary decision-making problems, where one of the two options is held constant across the binary choices. In the CET the recurring option is a coin-flip lottery (that is, a lottery with two possible outcomes occurring with the same probability) and the option that changes across choices is a safe amount of money. If the researcher is only interested in qualitative information about the risk attitude of a subject, then exposing her to just two binary choices – one in which the safe amount of money is just below the expected value of the lottery and another in which it is just above – is sufficient: If the subject decides for the lottery in both cases she is classified as risk-loving, if she decides for the lottery in the former choice and for the safe amount in the latter then she is classified as risk-neutral, and if she decides for the safe amount in both choices then she is classified as risk averse. This is very similar to the minimalist version of the EET described above, the main difference being that in the latter the attitude of the DM has to be elicited for two domains, for the domain of advantageous inequality and for the domain of disadvantageous inequality. An implication of this latter difference is that the minimum test size of the EET is four binary choices, while the minimum test size of CET is just two binary choices. The minimal version of the CET (as described in the previous paragraph) gives only qualitative information about the risk attitude of the DM (it discriminates only between three types of DM – risk-averse, risk-neutral and risk-loving, where risk-neutrality cannot be identified exactly but only “with arbitrary precision”), just as the minimal version of the EET described previously gives only qualitative information about the distributional attitude of the DM (it discriminates only between the nine archetypes of distributional concerns listed in Table 1, where vertical parts of an indifference curve cannot be identified exactly but only “with arbitrary precision”). The standard implementation of the CET differs from the minimal version described above in two respects: First it exposes subjects to more than one binary choice where the safe amount of money is higher (lower, respectively) than the expected value of the lottery; and secondly it includes one binary choice where the safe amount exactly equals the expected value of the lottery. The symmetric basic version of the EET (to be introduced in the next subsection) shares these two features: In terms of Fig. 3 (and focusing on the domain of disadvantageous inequality) it exposes subjects (i) to more than one choice between an option with the qualitative feature of p1 (p2, respectively); and (ii) to one choice where the alternative to the recurring reference point is located exactly on the intersection of the horizontal line above the reference point and the vertical line through the reference point. With both tests the aim of the former modifications (in comparison to the minimal version) is to get information about preference intensity while the latter modification is intended “to give a sign to neutrality” (see below). The overall goal of the CET is to identify the safe amount that generates indifference to a given gamble. With a list of binary choices the point of indifference cannot be identified exactly. However, by keeping the lottery constant and increasing the safe amount systematically from one choice to the next the researcher can identify the “switching point” of the subject, i.e., the binary choice where the subject switches from the lottery to the safe alternative. This switching point gives a range for the point of indifference and thereby for the certainty equivalent of the subject to the given lottery. Suppose a subject decides for the lottery in all choices where the expected value of the lottery is higher than the safe amount and for the safe amount in all choices where it is lower. Then the behavior of the subject is consistent with risk neutrality. However, it is also consistent with a low degree of risk aversion and with a low degree of risk loving. By exposing the subject in addition to a choice where the safe amount corresponds to the expected value of the lottery the researcher “attaches a sign to risk neutrality”. The overall goal of the EET is the identification of the locations of two points of indifference to the reference allocation, one for the domain of advantageous inequality, the other for the domain of disadvantageous inequality. With a list of binary choices the points of indifference cannot be identified exactly. However, by keeping the symmetric reference point and the material payoff of the other person in the asymmetric allocation constant across binary choices (in a given domain) and increasing the material payoff of the DM systematically from one choice to the next the researcher can identify the “switching point” of the subject in the respective domain, which gives a range for the point of indifference of the DM in the domain under scrutiny.21 As will be shown in Section 4.3, this information can be used to construct a two-dimensional index representing both the archetype of distributional concern and the preference intensity (conditional on the chosen vertical distance between r and the horizontal line). Suppose a subject decides for the symmetric reference point in all choices where her material payoff in the reference allocation is higher than her payoff in the asymmetric allocation and for the asymmetric allocation in all choices where it is lower. Then the behavior of the subject (in the domain under investigation) is consistent with selfishness. However, it is also consistent with a low degree of benevolence and with a low degree of malevolence. By exposing the subject in addition to a binary choice where her payoff is the same in the reference point and in the asymmetric allocation, we elicit her impartial distributional preference thereby “attaching a sign to selfishness”.22 Given the many similarities between CET and EET it probably does not come as a surprise that the two also share many pros and cons (in comparison to econometric elicitation techniques). The main advantages of the two tests are (i) that they are simple and short as they merely require subjects to complete a comparatively short sequence of binary decision making problems, properties that facilitate comprehension by experimental subjects and serve the experimenter’s need to limit the duration of experimental sessions; (ii) that they are parsimonious as they rely on a small set of comparatively mild primitive assumptions on preferences; (iii) that they are general as they directly tests the core features of preferences rather than concrete models or functional forms; (iv) that they are flexible as test size and test design can easily be fine-tuned to the research question of interest; (v) that they are precise because they identify the preference type with arbitrary precision and also give an index of preference intensity; and (vi) that they minimize experimenter demand effects as subjects are asked to make binary decisions in a neutral frame and do not have the option to do nothing. The main disadvantages of the two tests in comparison to econometric elicitation techniques are (i) that the switching point(s) of a subject give(s) only a range for the point(s) of indifference, which implies that “neutrality” cannot be identified exactly but only “with arbitrary precision”; (ii) that the assumptions on which the approaches rely are not directly tested; (iii) that the index of preference intensity for a given subject and the distribution of types that is inferred from a sample of subjects depend on the chosen parameterization of the test; and (iv) that they provide no measure of uncertainty of a subject’s elicited preference type. We discuss this latter issue further in Section 4.5. The symmetric basic version of the Equality Equivalence Test ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ As explained above the EET exposes subjects to a series of diagnostic binary choice problems. In the (symmetric) basic version of the test the family of binary choices is characterized by four positive integers, e, g, s and t, where e determines the locus of the equal-material-payoff allocation (m, o)=(e, e); g is a “gap” variable characterizing the vertical distance between (e, e) and the two horizontal lines in Fig. 3 (see Fig. 4); in order to avoid zero or negative monetary payoffs we restrict g to values strictly smaller than e; s is a “step size” variable characterizing the horizontal distance between two adjacent points on a line; t≥1 is a “test size” variable determining the number of steps (of size s) which are made to the left and to the right starting from the point just above or below (m, o)=(e, e); in order to preserve advantageous and disadvantageous inequality we impose the restriction t≤g/s. In total the symmetric basic version of the EET consists of 4t+2 binary decision problems. In each decision problem the subject is asked to decide between two alternatives (named Left and Right), each involving a payoff pair – one payoff for the subject (the DM) and one for the (randomly matched, anonymous) other subject (the passive person). For expositional purposes the decision problems are separated into two blocks, the disadvantageous inequality block (X-List) and the advantageous inequality block (Y-List). Within each block the decision problems are presented as rows in a table. In each decision problem one of the two alternatives (the alternative “Right”, say) is the (recurring) equal-material-payoff allocation (m, o)=(e, e). For the disadvantageous inequality block the second alternative in each decision problem (the alternative “Left”) is constructed as shown in Table 2. The construction of the second alternative for the advantageous inequality block is similar, the only difference being the material payoff of the passive person for the alternative LEFT, which is now e−g (instead of e+g). An important feature of the EET is that within each of the two blocks the material payoff of the passive person in the asymmetric allocation is held constant, while the material payoff of the DM increases monotonically from one choice to the next. Together with the fact that the symmetric allocation remains the same in all choices, this design feature guarantees that strict m-monotonicity is enough to make sure that when facing the choice between Left and Right within a given block, each individual switches at most once from Right to Left (and never in the other direction). In Section 4.4 I will use the two switching points of a subject to construct a two-dimensional index representing both archetype and intensity of distributional concerns (conditional on the chosen test parameters – see Section 4.5 for a discussion). As previously mentioned the EET allows for discrimination between the nine archetypes at any arbitrary precision. More specifically, the researcher needs to define when an agent should be considered as egoistic in a particular domain (this is the meaning of arbitrary precision). Suppose we define an agent to be egoistic in a particular domain if she is not willing to give up c Cents in order to change the material payoff of the passive person by 1$. Then the appropriate EET has to be such that c=100s/g⇔s=cg/100 meaning that we can choose the remaining parameters of the test freely. Extending and refining the Equality Equivalence Test ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The working paper version of this article (Kerschbamer, 2013) proposes three modifications of the symmetric basic version of the test that might help to shed light on more specific research questions. The first modification replaces the symmetric step-size in the basic version by an asymmetric one (where the step size is small at the center but grows larger when moving away from the center) in order to increase the power of the test to discriminate between selfish and different variants of non-selfish behavior without increasing the size of the test or decreasing the discriminatory power of the test at the borders. The second modification extends the X-List to the left and the Y-List to the right in order to address the question whether there are subjects who (in the relevant range) put more weight on the material payoff of the passive person than on their own material payoff. The third modification is a multi-list version of the EET where subjects are asked to complete two or more X- and Y-lists distinguished by the size of the gap variable g. This modification is intended to gain more insights on the exact shape of indifference curves in (m, o)-space. Identifying archetype and characterizing intensity of distributional concerns: the (x, y)-score ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ This subsection describes a method to identify the archetype and to characterize the intensity of the distributional preference of a subject based on her choices in the symmetric basic version of the test. It then proposes a procedure to represent the type- intensity distribution of a given subject pool graphically. Step 1 (consistency check): As argued above an individual whose preferences satisfy strict m-monotonicity has at most one switch from Right to Left (and no switch in the other direction) in each of the two tables. Step 1 is to eliminate all subjects that fail this basic consistency check (in an implementation of the symmetric basic version of the test – see Section 5 for details – less than 5% of the subjects failed the consistency check). Step 2 (defining scores): Represent each subject with consistent behavior by an (x, y) tuple defined as follows: The variable x (x-score) summarizes the behavior of the individual in the disadvantageous- inequality related block (X-List) and is defined as (t+1.5) points minus the row number in which the individual decides for the first time for the asymmetric allocation (that is, for the payoff vector on the left hand side). If an individual always decides for the symmetric (or egalitarian) allocation, we take the convention that she decides for the first time for the asymmetric allocation in the (2t+2)th row, so that she gets an x-score of −(t+0.5). For instance, if in the test version displayed in Fig. 4 (where t=2) an individual decides for the symmetric allocation in the first row of the X-List and for the asymmetric allocation in the second (and in all other) row(s) then she gets an x-score of 3.5–2=1.5. The variable y (y-score) summarizes the behavior of the subject in the advantageous-inequality related block (the Y-List) and is defined as the row number in which the individual decides for the first time for the asymmetric allocation minus (t+1.5) points. If an individual always decides for the symmetric allocation, we take again the convention that she decides for the first time for the asymmetric allocation in the (2t+2)th row; she then gets a y-score of t+0.5. Note that the definition of the two scores implies that each of them can take on 2(t+1) different values (see Table 3); thus, the proposed test allows for 4(t+1)² different (x, y)-scores. Also note that a positive (negative) x-score corresponds to benevolence (malevolence) in the domain of disadvantageous inequality, while a positive (negative) y-score corresponds to benevolence (malevolence) in the domain of advantageous inequality. Furthermore, the magnitude of the x-score (y-score, respectively) is an ordinal index of the intensity of distributional preferences in the domain of disadvantageous inequality (advantageous inequality, respectively).23 Step 3 (representing relative frequencies of types): Represent the absolute or relative frequencies of the different (x, y)-scores in an axis of abscissas as shown in Fig. 5. Relation of (x, y)-score to parameters in piecewise linear model and to WTP ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Note that the piecewise linear model implies that the DM’s willingness to pay (WTP) for income increases (or decreases) of the passive person is piecewise constant (WTP=uo/um, where the subscripts denote partial derivatives). In the domain of disadvantageous (advantageous) inequality we have WTPd=σ/(1−σ) (WTPa=γ/(1−γ), respectively); if σ≥0 (γ≥0, respectively) then this term gives the own-money amount the DM is willing to give up in the domain of disadvantageous inequality (advantageous inequality, respectively) in order to increase the other person׳s material payoff by a single unit; symmetrically, if σ<0 (γ<0, respectively) then −σ/(1−σ) (−γ/(1−γ), respectively) gives the own-money amount the DM is willing to give up in the domain of disadvantageous inequality (advantageous inequality, respectively) in order to decrease the other person׳s material payoff by a single unit. Thus, within the piecewise linear model x-scores translate into WTPd as shown in the right-most column of Table 4 (again the translation for the y-score is similar except that WTPd is replaced by WTPa and strict inequalities are replaced by weak ones). It is important to note that using estimates of the parameters of the piecewise linear model (or estimates of the piecewise constant WTP for changes in the income of the other) as a cardinal metric for distributional preferences does not necessarily mean assuming piecewise linear preferences: In experimental set ups, where stakes tend to be small, the estimates are probably best interpreted as linear approximations of the true values. This interpretation is especially valid when the parameters of the piecewise linear model are estimated from the raw data using the McFadden (1974) random utility specification. In addition to yielding a cardinal metric that is comparable across studies, estimating the parameters of such a structural model has several other practical advantages as well:24 As McFadden׳s random utility specification allows for noisy decisions, subjects with inconsistent choices do not have to be dropped. This may be crucial when a test design with high resolution (i.e. with large gap variable g and small step size s) is used, or when the binary choices are presented to the subjects one-at-a-time in random order (see Appendix A for a discussion of implementation issues). The parameters׳ standard errors enable statistical tests; for example, to check whether a subject׳s deviations from purely selfish behavior are statistically significant. The structural model could also be applied in the context of a finite mixture specification. This would allow the experimenter to identify the prevalent social preference types and to endogenously classify each subject into the type that fits her behavior best. The parameters of the piecewise linear model can also be estimated when the test is applied in its multi-list variant (introduced in Kerschbamer, 2013) where the (x, y)-score is no longer available.","Here the data from a paper-and-pen experiment based on the symmetric basic version of the test is reported. The experiment was conducted in paper-and-pen (and several other design features reported below were applied) to convince subjects that neither other experimental subjects nor the experimenters could identify the person who has made any particular decision. This was done in an attempt to minimize the impact of experimenter demand and audience effects. See List (2007) for a discussion on experimenter demand effects and Hoffman et al. (1994), Andreoni and Petrie (2004), and Andreoni and Bernheim (2009) for experimental evidence indicating that – depending on the experimental design – audience effects might have a large impact on subjects׳ behavior in dictator-game like situations.25 Experimental procedures ~~~~~~~~~~~~~~~~~~~~~~~ Five experimental sessions were conducted manually (i.e., in pen-and-paper) at the University of Innsbruck in autumn 2009. Forty subjects who had not participated in similar experiments in the past were invited to each session using the ORSEE recruiting system (Greiner, 2004). Since not all subjects showed up in time, 192 (instead of the invited 200) subjects from various academic backgrounds participated in total, and each subject participated in one session only. After arrival, subjects assembled in one of the two laboratories and individually drew cards with ID numbers (which remained unknown to other participants and the experimenters). Then instructions were distributed and read aloud.26 Instructions informed subjects (i) that there are two roles in the experiment, the role of an “active person” and the role of a “passive person”; (ii) that there is exactly the same number of active and passive subjects in the experiment and that roles are assigned randomly; (iii) that each active person is matched with exactly one passive person and vice versa, and that at no point in time a participant will get to know anything regarding the identity of the person she/he is matched with; (iv) that active persons are called to make a series of ten binary decisions that determine not only their own earnings from the experiment but also the earnings of the passive person they are matched with; (v) that passive persons do not have a decision to make in the experiment and that their earnings will depend exclusively on the decisions of the active person they are matched with; (vi) that only one of the ten choice problems of each active person will be relevant for cash payments; and (vii) that cash payments could be collected the day after the experiment at one of the secretaries who also handles the cash payments for other experiments (to ensure that the amount a subject earns cannot be linked to her/his decisions). Then subjects were randomly assigned to one of the two roles; active persons stayed in the same room while passive persons were escorted to the adjacent laboratory. In both rooms subjects were seated at widely separated computer terminals (computers were switched off) with sliding walls. Active persons were handed out a form consisting of two pages – an empty cover sheet and a decision sheet as described in the next paragraph – and they were asked to fill out the decision sheet in private. Passive persons received a form consisting of three pages – an empty cover sheet and a two-page questionnaire unrelated to the experiment – and they were asked to complete the questionnaire in private. After the tasks in both rooms had been completed, for each active person one of the choice problems was randomly selected via a manual device – a bingo ball cage handled by the active person – for the purpose of cash-payment generation. The payoff-relevant decision problem was written on the cover page of the active person and the person was given the opportunity to take (in private) a look at her/his choice in the payoff-relevant decision problem. Now subjects in both rooms were asked to label (in private) the cover sheet of their document with their ID number. Then participants in both rooms were called to put their documents (again in private) in boxes before leaving the room. Anonymous cash payments started the next day – giving experimenters the opportunity to manually match active with passive persons in the meantime. Participants presented the card with their ID number to an admin staff person, who did not know who did what for which purpose nor how cash payments were generated, and they got their earnings in exchange (the fact that cash payments would be made that way was clearly indicated in the instructions). On average subjects earned approximately 11 Euros plus a show up fee of 4 Euros. Experimental design ~~~~~~~~~~~~~~~~~~~ The symmetric basic version of the test was implemented with e=10, g=3, s=1, t=2 and with experimental currency units corresponding to Euros. Thus, each active person (96 in total) was exposed to 10 binary decision problems with (10, 10) as the recurring equal-material- payoff allocation. The decision problems were presented in two tables, 5 in the X-Table (disadvantageous inequality) and 5 in the Y-Table (advantageous inequality). The design of the two tables was similar to that of Table 2. Discussion ~~~~~~~~~~ The experiment reported here uses the fixed-role-assignment protocol, where roles (active DM and passive person) are assigned ex ante, and only active DMs decide while passive persons do nothing. While this protocol seems to be the cleanest one from a theoretical point of view, it is not practicable when the test is intended as a tool to be added to arbitrary other experiments, since the preferences of half of the subjects remain unclassified. An easy way to have a measure of all subjects׳ social preferences is to either use the role-uncertainty protocol (where each subject decides in the role of the active DM, and only later subjects get to know whether their decision is relevant – as in Engelmann and Strobel, 2004 and in Blanco et al., 2011, for instance), or the double-role- assignment protocol (where each subject decides, and each subject gets two payoffs, one as an active DM and one as a passive person – as in Andreoni and Miller, 2002 and in Fisman et al., 2007, for instance). Appendix A discusses some pros and cons of the different protocols. A second issue worth discussing regards the implemented test version. The experiment reported here uses the symmetric basic version of the test (which has equidistant step sizes in the binary choice lists) with a relative low resolution (i.e., a relatively high value of the quotient s/g). As is evident from the results, however, this form of the test yields a classification that is coarser than some researchers might find ideal. To address this issue, either an asymmetric test version with small step sizes in the center and larger step sizes in the periphery (as suggested in Section 4.3) could be used, or the power of the symmetric version of the test to discriminate between selfish and different variants of non-selfish behavior could be increased by increasing g, keeping the rest of the test as it is (remember the discussion on “identification with arbitrary precision” in Section 4.2). A third implementation issue regards the presentation of tasks. In the paper-and-pen experiments reported here the binary decision tasks were presented to the subjects in ordered lists (similar to the lists often used in risk- attitude elicitation tasks). In computer-aided experiments presenting the binary decisions one-at-a-time in random order (i.e., each binary decision on an own screen) might be an attractive alternative. Appendix A discusses this issue further. A final point worth addressing regards the comparison of subjects according to the intensity of preferences. The implemented standard version of the test uses only one size of g. If subjects differ in the shape of indifference curves, their relative ranking regarding preference intensity may depend on g. For instance, one person might be more altruistic than another if g is small, but less altruistic than the other if g is large. More generally, if distributional preferences are non-linear, any results regarding the relationship between intensity of distributional preferences and behavior in another experiment will depend to some degree on the level of g chosen for the test. So, if a researcher is interested in correlating the intensity of benevolence or malevolence in the two domains (i.e., the x- and the y-score) with behavior in another experiment it seems advisable to adapt the parameters of the test to the parameters in that other experiment.","This paper has proposed a geometric delineation of distributional preference types and a non-parametric approach for their identification in a two-person context. Major advantages of the proposed Equality Equivalence Test (EET) over previous ones are (i) that it is simple and short as subjects’ task is to make a small set of diagnostic choices without feedback; (ii) that it is parsimonious as it relies on a small set of primitive assumptions; (iii) that it is general as it directly tests the core features of different types of distributional preferences rather than concrete models or functional forms; (iv) that it is flexible as test size and test design can easily be fine-tuned to the research question of interest; (v) that it is precise as it identifies the archetypes of distributional concerns with arbitrary precision and also gives an index of preference intensity; and (vi) that it minimizes experimenter demand effects as subjects are asked to make binary decisions in a neutral frame and do not have the option to do nothing.30 Those features together suggest that the EET might be suitable as a tool in experimental economics to disentangle the impact of distributional preferences from that of other factors thereby helping to interpret data from other (unrelated) experiments (similar to the choice list tests used to elicit risk attitudes; see Holt and Laury, 2002, or Dohmen et al., 2010, 2011).31 That the EET is indeed suitable for that purpose has been shown in two recent studies: Balafoutas et al. (2012) investigate in a standard lab experiment the relationship between distributional preferences and competitive behavior and find (a) that distributional archetypes (as assigned by the proposed test) differ systematically – and in an intuitively plausible way – in their response to competitive pressure, in their performance in a competitive environment and in their willingness to compete; and (b) that controlling for the effects of distributional preferences, as well as for risk attitudes and some other factors, closes the large gender gap in competitive behavior found in earlier studies (by Niederle and Vesterlund, 2007, 2010, for instance). This is an important finding because it indicates that the gender gap in competitiveness is largely driven by mediating factors (potentially accessible to policy intervention) and not by gender per se. Hedegaard et al. (2011) examine in a large-scale internet experiment the impact of distributional concerns on the contribution behavior in a standard (linear) public goods game and find (a) that distributional archetypes differ systematically – and in an intuitively plausible way – in their contribution behavior; and (b) that accounting for the differences explains roughly half of the gap between actual behavior of subjects in the lab and the theoretical benchmark derived under the assumption that players are rational and selfish (and that this fact is common knowledge). Again, this is an important finding because it helps to disentangle the impact of distributional concerns on the behavior of subjects in social dilemma games from that of other factors – as beliefs on others׳ behavior or intentions, for instance. Together the findings in those studies clearly indicate that associating subjects with one of the proposed archetypes of distributional concerns has explanatory value and that the proposed test is indeed a valid control instrument in experimental economics. Given that the EET does not provide a measure of uncertainty of a subject’s classification (in the sense of a counterpart of the standard error of an estimated parameter of a structural model), a systematic investigation of the test–retest reliability of the EET would be an interesting area of future research. The results of a recent study suggest that this reliability is high: Balafoutas et al. (2014) compare experimentally the revealed distributional preferences of individuals and teams by exposing subjects to the EET under two different decision-making regimes: an individual regime and a team regime. The authors employ a mixed within- and between-subjects design in two sets of sessions run in two consecutive weeks: In the first week all subjects are exposed to the EET in the individual regime; in the second week some subjects are again exposed to the individual regime, while the rest make their choices in the EET in the team regime. This design feature allows addressing the test–retest reliability issue by comparing the choices of subjects who face the individual regime twice across the two weeks. The authors show (in Table 5) that elicited preference types remain remarkably stable over the two weeks. Beyond its potential to act as a control instrument in experimental economics, other potentially fruitful applications of the EET include (a) investigating the stability of distributional preferences over different domains (for instance, a potential shortcoming of the approach proposed here is its focus on the two-agents case; investigating whether the preferences revealed in that context carry over to a richer environment is surely an important issue);32 (b) investigating possible links between distributional preferences and other forms of other-regarding preferences (for instance, “Are altruists more or less likely to be motivated by positive or negative reciprocity?”, “Do altruism and altruistic rewarding (or altruistic punishment) go together or are they mutually exclusive ways to reach the same goal – promoting private provision of public goods?” 33, or “Is the test-based classification of subjects in distributional-preference types somehow correlated with the propensity to be motivated by trust?”) ;34 and (c) applying the EET (together with tests for risk and time preferences and for personality traits) in experiments with large demographic variation (age, gender, income, education) or with a representative sample of the population to detect patterns and correlations (for instance, “Are distributional preferences and risk attitudes or time preferences somehow related?”, “Are there gender differences in the distribution of archetypes?”35 or “What is the impact of age and income on distributional preferences?”). Beyond economics the proposed test might help to address important research questions in biology and psychology as, for instance, “What determines human altruism (or spite)?” or, “What drives altruistic punishments and rewards?”. For those and many other interesting research questions, identification of distributional preference types in a “clean” environment appears to be a natural first step. The proposed EET seems to be well suited for this purpose. Turning back to the quote at the start of the paper the hope is that it turns out to be “as simple as possible, but not one bit simpler”."],["We estimate the impact of a carbon tax on manufacturing plants using panel data from the UK production census. Our identification strategy builds on the comparison of outcomes between plants subject to the full tax and plants that paid only 20% of the tax. Exploiting exogenous variation in eligibility for the tax discount, we find that the carbon tax had a strong negative impact on energy intensity and electricity use. No statistically significant impacts are found for employment, revenue or plant exit. © 2014. --------------------------------------------------------------------------------","The rise of climate policy on government agendas around the world has stirred a renewed interest in the optimal design of large-scale regulation of environmental externalities. Climate change – the “ultimate commons problem” (Stavins, 2011) – is caused by anthropogenic emissions of greenhouse gases (GHG) such as carbon dioxide (CO2) and is expected to have severe ecological and economic consequences (IPCC, 2007). Mitigating climate change will require substantial abatement of GHG emissions from all core economic sectors (Pacala and Socolow, 2004). The choice of appropriate policy instruments for each of these sectors is essential for minimizing the overall economic costs of mitigation with given technologies (static efficiency), and for stimulating technological innovations that will further reduce mitigation costs in the future (dynamic efficiency). This paper evaluates the performance of one such instrument, a tax designed to curb industrial CO2 emissions, in a panel of manufacturing plants. Manufacturing is a major contributor to GHG emissions around the world.1 Since most manufactured goods are tradable, there is a risk that regulated firms will lose international competitiveness, shed part of their labor force or even exit. These concerns have been fueling vehement opposition towards regulation and left their mark on the design of the policies implemented so far. Command- and-control policies have long been the predominant form of environmental regulation in the manufacturing sector, and their impacts have been studied extensively in the context of air pollution.2 On theoretical grounds, economists have favored market-based instruments such as taxes and tradable permit schemes because they are more efficient in both the static and dynamic senses (e.g. Montgomery, 1972; Milliman and Prince, 1989; Tietenberg, 1990). However, empirical evidence on the impacts of market-based environmental regulation on manufacturing is scarce, especially when it comes to carbon emissions.3 For example, the European Union Emissions Trading Scheme (EU ETS), the largest cap-and-trade system for carbon emissions worldwide, is overdue for a microeconometric evaluation (Martin et al., 2013b). While carbon taxes have been implemented in various EU countries, their rigorous evaluation has proven difficult, be it because of the lack of suitable microdata or because of the lack of a compelling identification strategy.4 This paper fills the void by analyzing the Climate Change Levy (CCL) package — the single most important climate change policy that the UK government has unilaterally imposed on the business sector so far (HM Government, 2006). The package consists of a carbon tax – the CCL – and a scheme of voluntary agreements available to plants in selected energy intensive industries. Upon joining a Climate Change Agreement (CCA), a plant adopts a specific target for energy consumption or carbon emissions in exchange for a highly discounted tax liability under the CCL. While the CCL package is still in place today, our analysis focuses on the first three years following its introduction in 2001, thereby avoiding overlap with the EU ETS. During the period of analysis, the CCL added 15% to the energy bill of a typical UK business (NAO, 2007) and the discount granted under a CCA amounted to 80% of the tax rate. Given its scope and institutional context, the CCL package provides a unique opportunity to study the effects of a carbon tax in an industrialized economy. We use longitudinal data on manufacturing plants to estimate the impact of the CCL on energy use, emissions and economic performance. Our identification strategy is to compare changes in outcomes between fully-taxed CCL plants and CCA plants. A naïve difference-in-differences (DiD) estimator would likely be biased because the plants that were eligible for CCA participation could self-select their tax regime. However, plants were only eligible if they emitted pollutants subject to environmental regulation pre-dating the CCL. The variation in eligibility across plants can hence be exploited to instrument for the tax rate. We implement this idea in an IV framework where the reduced form is a DiD regression of plant outcomes on eligibility. For this approach to be consistent, it must be true that differences between eligible and non-eligible plants are not systematically related to changes in outcome variables over the treatment period. While this assumption is not testable, we show that there are no significant trend differences between eligible and non-eligible firms in the pre-treatment period. In addition, we exploit the panel structure of the dataset to control for pre-trends directly in the regression. Firms in the control group were not only entitled to a tax discount, but they also faced a reduction target for energy consumption or carbon emissions. Although these targets could have placed binding constraints on the plant's production choices, the fact that massive over compliance occurred right from the start suggests otherwise. In fact, a large degree of flexibility was built into both the target negotiation process and the compliance review. If targets were nonetheless stringent, then our estimate represents a lower bound on the full price effect of the tax differential between the two groups of plants. With this approach, we find robust evidence that the CCL had a strong negative impact on energy intensity, particularly at larger and more energy intensive plants. An analysis of fuel choices at the plant level reveals that this effect is mainly driven by a reduction in electricity use and translates into a negative impact on CO2 emissions. In contrast, we do not find any statistically significant impacts of the tax on employment, revenue (gross output) or total factor productivity (TFP). In addition, we examine extensive-margin adjustments and find no evidence that the CCL accelerated plant exit. While the regression-based test we use does not have much power to detect small negative impacts on these outcomes, our results do not substantiate worries about devastating effects of the CCL on the competitiveness of UK manufacturing, which gave way to a costly exemption scheme.5 Over the past two decades, carbon taxes and their effects on industrial competitiveness have been a matter of political debate in many industrialized countries. By conducting the first ex-post analysis of the causal impact of such a tax on manufacturing, our study provides much-needed empirical evidence on the impacts of large-scale regulation aimed at pricing pollution. It does so in the context of climate change – an area where regulatory stringency is bound to increase in the near future – and with a focus on manufacturing, the principal engine of growth in the emerging economies and still a cornerstone of employment in post-industrial economies. The remainder of the paper is structured as follows. Section 2 describes the CCL package in detail and reviews previous research on the tax. Section 3 describes the research design and econometric framework. Section 4 describes the data sources and summarizes the dataset used for the analysis. Section 5 reports the main results and presents several robustness checks. Section 6 examines heterogeneous impacts, aggregate effects and estimates the impact of the CCL on exit. Section 7 concludes. The Climate Change Levy and Climate Change Agreements ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Since the 1990s the UK has adopted a series of increasingly ambitious targets for climate policy. In addition to a 12.5% reduction of GHG emissions from 1990 levels to be achieved under the Kyoto Protocol, the Blair administration promised to reduce CO2 emissions by 19% by 2010 and by 60% by 2050.6 When the CCL package was implemented in 2001, it constituted the single-most important policy aimed at achieving these goals.7 The CCL is a per unit tax payable at the time of supply to industrial and commercial users of energy. It was first announced in March 1999 and came into effect in April 2001. Taxed fuels include coal, electricity, natural gas, and non-transport liquefied petroleum gas (LPG). For each fuel type subject to the CCL, Table 1 displays the tax rates per kilowatt hour (kWh) equivalent, the average energy price in Pound Sterling paid by manufacturing plants in 2001 and the implicit carbon tax. Energy tax rates vary substantially across fuel types, ranging from 6.1% on coal to 16.5% on natural gas.8 While the tax establishes a meaningful price incentive for energy conservation overall, it is immediately seen that carbon contained in gas and electricity is taxed at almost twice the rate as carbon contained in coal.9 Other fuel types were tax-exempt precisely because of their low carbon content, such as electricity generated from renewable sources and from combined heat and power. Hence, rather than a pure carbon tax the CCL is a tax on energy with non-uniform rates, shaped by a mixed bag of fiscal and regulatory goals. Similar to other European governments that had introduced energy taxes during the 1990s, the UK government set up a scheme of negotiated agreements, the CCAs, in order to mitigate possible adverse effects of the CCL on the competitiveness of energy intensive industries. By participating in a CCA, facilities in certain energy intensive sectors can reduce their tax liability by 80% provided that they adopt a binding target on their energy use or carbon emissions. Defined either in absolute terms or relative to output, these targets were negotiated at two levels. In an ‘umbrella agreement’, the sector association and the government – represented at the time by the Department for Environment, Food, and Rural Affairs (DEFRA) – agreed upon a sector-wide target for energy use or carbon emissions in 2010 and on interim targets for each two-year compliance period. At a lower level, ‘underlying agreements’ stipulate a specific reduction to be achieved by a ‘target unit’, i.e. a facility or group of facilities in a sector with an umbrella agreement. DEFRA originally negotiated 44 umbrella agreements with different industrial sectors, including the ten most energy intensive ones.10 While the primary objective of both the CCL and the CCAs is to enhance the efficiency of energy use in the business sector, the two instruments represent fundamentally different approaches. The levy provides a price signal at roughly 15% of energy prices faced by the typical business in 2001 (NAO, 2007). If energy demand is price sensitive, the increased relative price of energy should lead to a reduction in energy consumption. In terms of CO2 emissions, this effect could be offset in part by a shift towards more carbon-intensive fuels. In contrast, the CCA combines a very diluted price signal of 0.2 × 15% = 3% of energy prices faced by the typical business with quantity regulation, mostly in the form of efficiency targets. This target affects the plant only if it places a binding constraint on the trajectory of energy use during the remaining economic lifetime of the plant. If this is not the case, the plant faces weaker incentives for energy conservation than it would under the full tax rate. Moreover, since most targets are specified in terms of energy units rather than carbon emissions, there is no guarantee that even a stringent energy target leads to emission reductions. How stringent are the targets negotiated in the CCAs? ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ In theory, an omniscient government can choose a combination of tax discount and reduction targets so as to induce at least as much abatement as under the full tax rate (Smith and Swierzbinski, 2007). In reality, however, the government is unlikely to have perfect information about firm-specific abatement cost, especially if firms worry that sharing this information with the government weakens their bargaining position in the target negotiations. What is more, the government might not have been willing to drive a hard bargain for fear of jeopardizing international competitiveness and exacerbating distortions in marginal abatement cost (de Muizon and Glachant, 2003; Smith and Swierzbinski, 2007). A closer inspection of the negotiation, monitoring and enforcement of CCA targets yields a number of reasons to believe that the targets did not place binding constraints on firm behavior. First, the government may have “double counted” carbon savings from the CCA scheme (ACE, 2005). On average, CCA targets were supposed to improve energy efficiency by 11% between 2000 and 2010. This figure is well above the 4.8% improvement the government expected to occur under a “business as usual” (BAU) scenario (AEAT, 2001). However, alternative BAU scenarios were much closer to the CCA target, projecting energy efficiency of all UK industry to improve by 9.5% (DG Transport and Energy, 1999) or even 11.5% when taking into account the effect of the CCL (DTI, 2000). Second, there was massive overcompliance with CCA targets. Combined annual carbon savings in all CCA sectors were substantially larger than the 2010 target throughout the first three compliance periods. At the end of the first compliance period in 2002, CCA sectors reported savings of 4.5 MtC — almost twice the target amount of 2.5 MtC to be achieved by 2010.11 Consistent with this, the proportion of compliant target units was high, rising from 88% in the first compliance period to 98% and 99% in the second and third compliance periods, respectively (AEAT, 2004, 2005, 2007). CCA participants that did not meet their target could attain compliance by buying emission allowances on the UK Emissions Trading Scheme (UK ETS), a carbon market that was operational between 2002 and 2006. Allowance prices in this market remained below the implicit carbon tax rates given in Table 1.12 Third, the lower bound on compliance cost is zero. This is because facilities were re- certified for the reduced tax rate even if they had missed their target, provided that the sector as a whole met its target. In 2004, this was true of approximately 250 non- compliant target units (NAO, 2007). Finally, a large degree of flexibility in both the target negotiations and the compliance review further limited the stringency of CCA targets. For instance, CCA sectors could choose their own baseline year for the target indicator. More than two thirds of all sectors chose a baseline year prior to 2000 (in some cases going as far back as 1990), allowing them to count carbon savings unrelated to the CCA towards target achievement (NAO, 2007). Furthermore, targets could be adjusted ex post to reflect a more energy intensive product mix, declining output, or other ‘relevant constraints’.13 Because of this, and for the reasons given above, it appears unlikely that the negotiated CCA targets placed binding constraints on energy use by the average CCA company.14 Previous evaluations of the CCL package ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Several evaluations of the CCL package were conducted at different stages of its implementation. In the 2000 Regulatory Impact Assessment, the government projected that the CCL instrument alone would achieve carbon savings of at least 2 MtC in 2010 against BAU projections (HMCE, 2000). This estimate was based on a model of business energy use maintained by the Department of Trade and Industry (DTI). An interim evaluation study, commissioned by DEFRA at the end of the second commitment period in 2004, finds evidence that the announcement of the CCL package in March 1999 reduced energy demand in the service and public sectors, but not in manufacturing (Cambridge Econometrics, 2005). The authors of the study identify this “announcement effect” as a structural break in an error correction model of quarterly energy demand (see Agnolucci et al., 2004, for more details). A series of simulation studies uses a macroeconometric model of the UK economy to assess the CCL package. An important result is “that the energy (and therefore carbon) saving and energy-efficiency targets would have been met without the CCAs” (Cambridge Econometrics, 2005, p. 7), which confirms the conclusion drawn above on the lack of target stringency. Since model simulations of the CCL package give rise to much smaller carbon savings than official estimates computed for the first compliance period (AEAT, 2004), Ekins and Etheridge (2006, p. 2079) conclude that “the CCL package as implemented […] achieved a greater carbon reduction than a no-rebate CCL would have done by itself”. They attribute this to managers becoming aware of more cost-effective efficiency enhancement projects as they started to benchmark their energy use. To be sure, the existence of such an “awareness effect” depends on whether the official carbon savings were real and not just a consequence of AEAT's (2001) pessimistic BAU scenario. In another simulation study on the impact of the CCAs on output and employment, a large effect of the CCAs on sectoral energy demand – averaging a 9.1% reduction in sectoral energy use by 2010 – is built into the model rather than estimated (Barker et al., 2007). These assessments of the CCL package highlight two fundamental challenges in policy evaluation, namely (i) to determine a valid baseline against which to measure the impact of a policy and (ii) to attribute any measured impact to this policy in a causal fashion. In studies that use simulated trajectories of energy use as a baseline against which to measure the impact of the CCL package, the validity of the results critically depends on the counterfactual baseline being true. In econometric studies based on time series data at the sector level, it is difficult to discern the effects of the policy from that of unobserved aggregate shocks.15 The present study is the first evaluation of the Climate Change Levy package to use longitudinal business microdata. We address the baseline problem by comparing changes in actual firm behavior under two types of policy regimes, thus purging the effect of aggregate shocks. Moreover, we identify the causal effect of the tax by exploiting exogenous variation in the eligibility rules for the tax rebate. The next section explains our research design in detail. Econometric model ~~~~~~~~~~~~~~~~~ Estimation of Eq. (1) recovers the full effect of the CCL if – as previous research has suggested – CCA targets did not impose binding constraints on firm behavior. If the converse is true, the estimated α falls short of the true price effect as control plants choose lower-than-optimal levels of energy so as to comply with their CCA target. Hence, the estimated parameter α can be regarded as a conservative estimate of the impact of the CCL. Fig. G.1 in the online appendix illustrates this point.16 In order to estimate α consistently, one needs to address the issue of non-random selection of plants into the control group. As we document in Section 4.2 below, CCA plants are, on average, older, larger and more energy intensive than CCL plants. Clearly, plants using large amounts of energy receive a larger absolute discount on their CCL liability which gives them a stronger incentive to join a CCA. In turn, as there are fixed costs of participating in a CCA, plants with low levels of energy use may find it more profitable not to join.17 This is illustrated in Fig. G.2a of the online appendix. In principle, selection effects can be addressed by adding further control variables, but selection might in part be driven by factors not directly observable to us. For instance, given two plants that initially use the same amount of energy, the plant with the steeper marginal abatement cost schedule has a stronger incentive to join the CCA (cf. Fig. G.2b in the online appendix for an illustration). Instrumental variable ~~~~~~~~~~~~~~~~~~~~~ Eligibility for CCA participation was granted to plants engaged in polluting activities regulated under the PPC act (listed in Appendix B.1). An eligible plant is comprised of at least one installation dedicated to the PPC activity, such as a blast furnace or cement kiln. The discounted rate of the CCL applies to all energy use at this installation.19 We define the instrumental variable Zi as an indicator variable that equals 0 for all plants containing at least one eligible installation, and 1 otherwise. The instrument is relevant because the eligibility of a plant for CCA participation ought to be correlated with its tax regime. Furthermore, the validity of using ∆Z as an instrument for ∆T in Eq. (2) rests on the identifying assumption that eligibility is orthogonal to shocks Δϵit that occurred after 2000. This assumption deserves a careful assessment. For instance, one might worry that plants could self-select into PCC activities in order to become eligible for the CCA. Since the entire CCL package was conceived and implemented in a mere two years, and eligibility rules were established only a year before implementation (in the 2000 Financial Act) it appears unlikely that firms switched technologies in the short run just because of the CCA discount. Moreover, if PPC regulated and non-regulated plants are subject to different trends in the outcome variables, the resulting IV estimates will be biased. In the empirical analysis to follow, we investigate this possibility by looking at pre-treatment trends but find no evidence of such differences. A visual examination of time series plots of various outcome variables (shown in Fig. 1 below) suggests no systematic differences in trends between eligible and non-eligible firms before 2001. The corresponding statistical test results are reported in panel B of Table 2 and fail to reject the hypothesis of common trends. Furthermore, our panel dataset allows us to directly control for differential trends in the outcome regressions. As we discuss in Section 5.3 below, this does not lead us to reject the hypothesis that outcomes in PPC firms and non-PPC firms followed a common trend before the introduction of the CCL. Also, the point estimates of the tax effect hardly change when controlling for pre-trends. Finally, the exclusion restriction also rules out the possibility that mandatory public disclosure of PPC pollution in the European Pollution Emissions Register (EPER) had a direct effect on the outcome variables. While this assumption is untestable, we are not aware of any evidence that EPER reporting requirements affected firm behavior in the UK.20 Moreover, the fact that pollution emissions in 2001 were published only in 2004 rules out any direct effects operating through the demand side. It is worth noting that the exclusive focus on pollution intensity when eligibility was first determined left many energy intensive industries ineligible for the tax discount. For instance, textile wet processing was an eligible activity thanks to its high pollution emissions, but not so dry processing which, although energy intensive, emits no pollution regulated under PPC. Similarly, both the production and the recycling of glass containers are very energy intensive processes. However, since only the former is pollution intensive, glass container recycling was not eligible for CCA participation.21 This institutional ‘glitch’ induces exogenous variation in the probability of treatment even within narrowly defined, energy intensive industrial sectors. Heterogeneity of the treatment effect ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ So far the treatment effect α was implicitly assumed to be homogeneous across plants. For the case of heterogeneous responses to treatment, Imbens and Angrist (1994) have shown that, under certain conditions, the IV estimator identifies the average treatment effect on “compliers”, i.e. on the subset of the treated for which a change in the instrument induces a change in treatment status. Although “compliers” need not be representative of all treated plants, an instrument based on a strict eligibility rule identifies the average treatment effect on the treated (ATT), simply because non-eligible plants cannot receive treatment (there are no “always-takers”). This result was first derived by Bloom (1984) and can be applied to our setting with only minor modifications to the interpretation. Recall that the treatment we consider is to pay the full tax rate, and that the instrumental variable indicates whether or not a plant is eligible for an exemption from treatment. All ineligible plants must pay the full tax rate, so that only eligible plants are able to escape the treatment (i.e. there are no “never takers”). In Section D of the online appendix, we show that the IV estimator identifies the average treatment effect on the non-treated plants (the ATNT), i.e. those that apply for a tax discount when given the opportunity. We shall refer to this sub-population in a more intuitive way as the group of “tax concerned” plants. As we explain in more detail below, we measure eligibility using data from the EPER database. These data cover all facilities with PPC emissions above certain reporting thresholds, whereas eligibility for a tax discount was granted regardless of the amount of emissions. In Appendix D we show that this has no effect on the interpretation of our estimates as long as firms below and above the reporting threshold do not differ systematically with respect to their treatment response and their probability of being tax concerned.","The compilation of a dataset suitable for the micro-econometric evaluation of the CCL required a major effort in terms of data collection, cleaning and matching. The result is a unique dataset that matches publicly available information on CCA participation and EPER coverage to production data from two confidential business datasets. Data sources ~~~~~~~~~~~~ The core dataset is the Annual Respondents Database (ARD) which is maintained by the Office for National Statistics (ONS) and can be accessed by approved researchers through the UK Data Service's secure access program.22 The ARD is an annual production survey that covers about 10,000 plants in the manufacturing sector.23 During the sample period, all plants with 250 employees or more (in some industries: 100 or more) had to report annually whereas smaller plants were included on a random basis (Barnes and Martin, 2002). The ARD provides information on the plant's age, number of employees, gross output (revenue), variable cost, capital stock, materials, and energy expenditures (inclusive of CCL payments). Detailed information on energy use is taken from the Quarterly Fuels Inquiry (QFI), a quarterly survey among a panel of about 1000 manufacturing plants managed by the ONS on behalf of DTI. The survey collects data on expenditures and quantities for all relevant fuel types, including medium fuel oil, heavy fuel oil, gas oil, liquefied petroleum gas (LPG), coal (graded, smalls), hard coke, natural gas, and electricity. We have data for the period from 1993 to 2004. The majority (83%) of the observations in the QFI can be matched to the ARD without difficulty because both surveys use the same underlying government business register IDBR as their sampling frame. However, due to random sampling in the ARD we do not have ARD data for all QFI plants.24 We gathered information on CCA participation from both the DEFRA and HM Revenue and Customs (HMRC) websites. Lists of facilities in the original sector agreements were downloaded from DEFRA's website. The agreements stipulate the certification periods and the sector targets along with the details on the calculation of the units of energy used and carbon emissions. They also contain a list of all facilities initially covered by the CCAs. Seven agreements lack sufficient information on the facilities covered by the CCA and thus had to be excluded from the analysis.25 The HMRC website provides, sector by sector, the list of facilities that have joined the CCA along with the date of publication.26 The lists are regularly updated and facilities that have resigned from the CCA are removed. We merged the DEFRA and HMRC lists to obtain a complete list of facilities that pay the reduced rate of the CCL. We match this information to the ARD and QFI by combining information on a plant's postcode and the UK Company Register Number (CRN). To construct the instrumental variable, we downloaded publicly available data from the European Pollution Emissions Register (EPER) which covers all European facilities regulated under the IPPC directive whose emissions exceed the reporting thresholds. The 2001 EPER file contains reporting thresholds and pollution discharges into air and water for 50 pollutants and covers 2397 facilities in 56 sectors of activity in the UK. We construct the instrumental variable NEPER as a dummy variable that equals one if a facility is not on the EPER list, i.e. it does not report emissions of any of the pollutants regulated under PPC legislation. A value of zero is assigned otherwise. Just like the treatment variable T, this variable is zero for all plants before 2001 and does not vary between 2001 and 2004. To match EPER facilities to plants in our dataset we use the same algorithm that we used for matching CCA participation data. Descriptive statistics ~~~~~~~~~~~~~~~~~~~~~~ Our regression sample comprises 6886 and 1079 plants in the ARD and QFI datasets, respectively.27 Table G.1 in the online appendix reports descriptive statistics. We calculate energy intensity as the share of energy expenditures in either gross output or variable costs (the sum of expenditures on materials, energy and wages), finding a substantial amount of dispersion between plants. For example, the energy expenditure share in gross output of a plant at the 90th percentile is seven times larger than that of a plant at the 10th percentile. We report both quantities consumed and expenditures paid for the fuel variables, after aggregating up some of the variables available in the QFI to obtain the categories liquid fuels (oil, petrol, and LPG), solid fuels (coal and coke) and natural gas (firm contract, interruptible contract, tariff). Moreover, we compute the share of natural gas in the consumption of both gas and electricity, and total CO2 emissions (in thousands of tonnes) on the basis of the fuel mix. The regression sample starts in 1999, because this is the first year for which energy expenditure data are available in the ARD, and covers the first two target periods that lasted from 2001 until 2004. This window of analysis avoids possible complications due to (i) an overlap with the EU ETS which affected approximately 500 CCA plants from 2005 onwards, (ii) adjustments of CCA targets for the third compliance period, and (iii) new entry of sectors in 2006 following changes in the eligibility rules. Table 2 displays the means of the main variables in the pre-treatment year 2000 (panel A) and the differences between year 2000 and 1999 (panel B), broken down by treatment and eligibility status.28 The treatment variable CCL takes a value of one if a plant pays the full tax rate and a value of zero if the plant participates in a CCA. Panel A shows that participation in CCAs is not random: CCA plants are, on average, older, larger and more energy intensive. For most of these plant characteristics, a t-test of equal group means for CCL and CCA plants rejects at the 1% significance level. Given this strong correlation between treatment status and observable plant characteristics, we cannot rule out that unobservable plant characteristics also influence selection. We address selection in levels by differencing out fixed unobserved plant characteristics in Eq. (2). To mitigate bias from selection on changes in the outcome variables, we instrument the difference regression using eligibility which is presumably exogenous to innovations in the outcome variables. This assumption is more credible if we find that eligible and non-eligible plants do not follow systematically different trends in terms of the outcome variables ahead of the treatment. We examine this in Fig. 1 by plotting average changes in the main outcome variables with respect to the year 2000, both for eligible and non-eligible plants, as well as by treatment status.29 This shows that trends were closely aligned when treatment was imminent. More formally, panel B of Table 2 reports the pre-treatment growth rates by treatment and eligibility status, along with the results of a t-test for group equality. The test never rejects at the 5% level, suggesting that differential pre-trends in outcome variables were not important. This mitigates concerns about changes in the outcome variables being confounded with unobserved attributes of eligible firms. Finally, selection bias might also arise if attrition rates are systematically related to treatment status. We investigate this in Section 6.3 below, finding no significant impact of the CCL on plant exit relative to CCA plants. Determinants of CCL status ~~~~~~~~~~~~~~~~~~~~~~~~~~ Table 3 reports the results from various regressions of CCL status on NEPER and other plant characteristics. Each regression is run in both the ARD and the QFI samples. Columns 1 and 4 report the marginal effects from a probit regression of CCL on NEPER in the cross section for the year 2001. The coefficients imply that a value of NEPER = 1 increases a plant's chances of paying the tax in full by 28.4% in the ARD sample and by 44% in the QFI sample. The results from the first-stage regression underlying the IV estimation of Eq. (2) in first differences are reported in columns 2 and 5. They corroborate that there is a robust positive and statistically significant relationship between the treatment variable and the instrument. Columns 3 and 6 display the results from a probit regression of CCL status in 2001 on various plant level controls evaluated at their 2000 levels. The coefficient estimates show that the simple correlations between CCL status and plant characteristics we found in Table 2 persist after controlling for sectoral differences. In particular, plants that were larger in terms of their capital and energy inputs prior to treatment were more likely to participate in a CCA. The coefficients on energy and gross output suggest that the same is true of more energy intensive plants. In Section B.2 of the online appendix we present further evidence pointing to size and energy intensity as the main determinants for take up among eligible plants. This is consistent with the notion that the 80% discount on the energy tax rate would allow only large and energy- intensive plants to accumulate enough tax savings to cover the fixed costs of CCA participation. Treatment effect of the CCL ~~~~~~~~~~~~~~~~~~~~~~~~~~~ The first two rows in panel A of Table 4 report the results for energy intensity measured as the share of energy expenditures in either gross output and variable costs, respectively. We find that the CCL caused plants to decrease their energy intensity relative to CCA plants. The point estimates from the IV regressions are − 0.181 for the former measure and − 0.211 for the latter. The effects are both economically and statistically significant. The importance of controlling for selection is evident from the sizable differences between the OLS and IV estimates. In particular, OLS estimation leads to an upward bias when estimating the effect of the CCL on the growth in energy intensity. This is because the OLS estimator does not correct for the self-selection of energy intensive plants into the low-tax regime, which we found in Table 3 above. As we show in Section 6.1 below, this type of plant responded more strongly to the CCL, causing bias towards zero in the OLS estimates. In rows 3 and 4, we break down the effect on energy intensity by looking at its components. The IV point estimates of − 0.095 for energy expenditure and 0.086 for gross output suggest that CCL plants both reduced energy and increased gross output so as to achieve the reductions in energy intensity reported in row 1. However, the point estimates are imprecise and lack statistical significance at conventional levels. This reflects the fact that both variables lump together prices and quantities, which are likely to move in opposing directions and thus attenuate the effect of higher energy prices.30 Furthermore, we obtain a positive but not statistically significant point estimate for employment of 0.082. We derive an estimate of the CCL impact on TFP from an augmented Eq. (1) which includes the production factors capital, labor, materials, and energy. This amounts to estimating a production function where the treatment variable captures the impact of the CCL on otherwise unexplained differences in TFP.31 The coefficients reported in row 6 are positive but small in magnitude and lack statistical significance. We thus cannot reject the hypothesis that the CCL had no effect on plant-level TFP. The evidence in panel A clearly shows that the CCL led to substantial reductions in plant-level energy intensity compared to the CCA. While the other coefficients are estimated less precisely, the point estimates are consistent with firms substituting labor for energy and increasing output prices in response to the energy price increase. In Appendix E we show that all the qualitative results in Panel A – including the stronger response of energy intensity than energy expenditures – can be generated by a simple equilibrium model with neo-classical production functions that exhibit a sufficiently large degree of substitutability between labor and energy. The CCL package was not part of any harmonized carbon tax scheme for Europe but a unilateral policy measure. As such, it may have had a detrimental effect on the competitiveness of UK industry. On the basis of the positive but insignificant point estimates we obtain for employment and gross output, however, we cannot reject the hypothesis that the CCL did not cause firms to shed jobs or lose revenue relative to CCA firms. While it seems plausible that the CCL lowered profits we cannot estimate this effect directly for lack of pertinent data. However, if profit losses were substantial they might have induced firms to shut down some plants. We examine this possibility in Section 6.3 below. From a climate-policy perspective, it is important to know whether reductions in energy expenditures in CCL plants actually occurred, whether they corresponded to reductions in energy consumption and whether they lowered carbon emissions. For example, instead of consuming less of all fuel types CCL plants might substitute towards fuels that are cheaper but also more polluting, such as coal. More detailed information on energy use is needed to address this issue, as the energy expenditures variable lumps together changes in the tax-inclusive price and quantity of energy, as well as the effects of substitution between different fuel types. Panel B of Table 4 reports results from regressions using quantity changes in energy consumption by fuel type which are available in the QFI sample. Although this sample is smaller than the ARD sample, we find economically and statistically significant evidence that the CCL caused plants to decrease their electricity use by 22.6%. For natural gas, solid fuels, and solid fuels as a share of total kWh consumed we obtain positive point estimates of the treatment effect.32 However, the coefficients are not estimated with enough precision to support conclusions about interfuel substitution. The significant decrease in electricity consumption among CCL plants translates into a decrease in carbon dioxide emissions ceteris paribus, but this could be offset by an increase in the consumption of other fuel types. The last row of Table 4 shows the impact of the CCL on total CO2 emissions, calculated as the sum of emissions across fuel types. The CCL is associated with a significant decrease in total CO2 emissions of 7.3% in the OLS regression. The point estimate increases slightly when going from OLS to IV, yet statistical significance is lost. We conjecture that this is due to the noisy estimates of the tax response for fuels other than electricity. In the absence of a larger sample that would enable us to estimate this effect with more precision, there are two possible ways of quantifying the effect of the CCL on carbon emissions. On the one hand, one can choose to disregard statistically insignificant coefficients altogether and conclude that the unchecked decrease in electricity consumption translates into a decrease in CO2 emissions of equal magnitude. On the other hand, a more cautious interpretation of the results is to use the point estimate of − 0.084 from the IV estimation which accounts for the possibility that some CCL plants switched into dirtier fuels such as coal. We thus conclude that the CCL – though not designed as a pure carbon tax – caused plants paying the full rate to reduce CO2 emissions by between 8.4% and 22.6% compared to plants that paid the reduced rate. In further regressions, reported in Section F of the online appendix, we interact the treatment indicator with year dummies so as to recover the time profile of the treatment response following the introduction of the CCL. This can reveal possible time delay in plants' responses to the treatment, or whether the treatment effect dies off after a while. We find that the tax has the largest effect on the ARD outcome variables in the first two years of treatment. While the negative impact of the CCL on electricity use is statistically significant from 2002 onwards, the point estimates for natural gas and coal are usually not statistically significant at the 5% level.33 Balanced sample Our sample is an unbalanced panel for a number of reasons: random sampling of smaller plants in the ARD, plant births and deaths, and missing responses from some plants in some years. As the set of plants in the sample changes slightly from year to year, the time profile of the treatment effect might reflect – at least in part – the changes in sample composition rather than the dynamic response to the CCL. Another potential problem with the unbalanced panel is that the results could be dominated by potentially more extreme responses of exitors. To address these concerns, we estimate the model with time interactions in a subset of “stayer” plants with observations in all years after 1999. The results are summarized in Tables G.4–G.6 in the online appendix. Since the sample size drops by about half in both samples, some of the estimated treatment effects lose statistical significance. However, the qualitative findings remain similar to the ones estimated on the full sample. Controlling for pre-treatment trends Our identification strategy relies on the (untestable) assumption that differences between eligible and non-eligible plants are not systematically related to changes in outcome variables over the treatment period. In Section 4.2, we have shown that pre-treatment trends did not differ across these groups in a statistically significant way, meaning that our estimates are unlikely to confound the impact of the treatment with pre-existing differences. To corroborate this, we include a time-invariant eligibility dummy in Eq. (2) so as to directly control for unobserved trends in the outcome variables, separately by eligibility status. 34 Tables G.7 and G.8 in the online appendix show that this yields qualitatively similar results, albeit less statistically significant ones in later years. Since the coefficients on the eligibility dummy are statistically insignificant for all outcome variables except solid fuels, we do not include them in our preferred specification. Common support regression Despite our IV strategy there might be concern that results are driven by a fundamental heterogeneity between treated (eligible) and non-treated (non- eligible) plants. Therefore, as a robustness test we restrict the control group to a common support which is identified by the predicted probability of a plant in the control group to receive treatment.35 We construct this common support sample by dropping plants that do not belong to the central 80% of the propensity score distribution, while also balancing the covariates between the treatment and the control group.36 The results obtained for the common support sample are reported in Table G.9 of the online appendix. For the ARD variables in panel A this leads to slightly larger point estimates, suggesting that heterogeneity within the treated group is not a major problem. In the smaller QFI dataset, about half of the sample needs to be dropped, but this entails no qualitative change to the results. The impact of the CCL in different subsamples ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Our discussion so far has focused on the average effect of the CCL on non-treated plants. It is useful to know how this effect varies across plants with certain characteristics. For example, the tax impact may differ from the ATNT in industries that are very energy intensive because the levy imposes a higher cost burden on these industries. Moreover, as the political cost of job losses is high, policy-makers might be interested in the tax impact on small firms which are responsible for the bulk of total employment. Finally, the impact of the CCL on competitiveness may be particularly high for firms in sectors with high import penetration, as foreign competition prevents them from passing compliance cost on to their customers through higher output prices. To shed light on this, we estimate the impact of the CCL separately: (i) for plants with more vs. less than 250 employees, (ii) for plants with high vs. low energy intensity and (iii) for plants with high vs. low trade intensity.37 The first two columns of Table 5 report the IV coefficients for the split by energy intensity, defined as the share of energy expenditures in gross output. Results for the low- and high-intensity groups are reported in the odd and even-numbered columns, respectively. The IV point estimates for energy intensity and energy expenditures indicate that the average effects reported in Table 4 are due to a strong response by plants in energy intensive sectors. The point estimates in this group are − 0.195 for energy intensity and − 0.154 for energy expenditures, both are statistically significant at 5%. In contrast, the point estimates for the low-intensity group lack statistical significance. The point estimates for electricity consumption are similar in magnitude across groups but lack statistical significance in the low-intensity group. In columns 3 and 4 of Table 5 we split the sample according to the trade intensity in 4-digit NACE sectors, which is computed as the value of imports and exports to non-EU countries over the total market size within the EU27.38 This measure has been used by the EU Commission to gauge the competitiveness impact of the EU ETS on manufacturing firms. To the extent that trade intensity measures the degree of competition from non-regulated countries, it picks up the (lack of) ability of firms to pass on the cost of the CCL to their customers. The point estimates for the ARD variables obtained in the trade intensive group closely follow those obtained in the full ARD sample. In contrast, the impact on energy intensity is not statistically significant in the low-intensity group. We do not find any significant impact on employment in either of the two groups. This gives rise to two interpretations: first, that trade intensity might not be a good criterion for identifying adverse effects on competitiveness; or second, that the hypothesis which states that there are no such effects should not be rejected. The last two columns of Table 5 report the results for the employment split. While the point estimates for energy expenditures in small plants and electricity use in large plants are negative and statistically significant at the 10% level, no clear pattern emerges from this comparison across size groups. Aggregate effects of a carbon tax ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ While the micro-level approach allows for better identification of the causal impacts of the tax, from a policy point-of-view the aggregate implications of the tax matter. In this section, we compute the effect of a counterfactual carbon tax similar to the CCL but without the reduced tax rate. This exercise allows us to compare our results to studies assessing the impact of energy price changes on fuel consumption at the aggregate level. Both numbers are at the upper end of elasticity estimates obtained in comparable studies. For example, Bjorner and Jensen (2002) estimate the energy price elasticity at 1.37 in the pooled cross-section and 0.50 in a fixed-effects specification.41 The reader should bear in mind, however, that we recover an estimate of a tax-induced price elasticity. Davis and Kilian (2011) argue that this is structurally different from elasticity estimates based on other kinds of price variation because taxes may be perceived as more persistent and hence induce larger behavioral changes. They also point to a possible additional effect of media coverage that accompanies the introduction of such taxes. Since the CCL was promoted as the UK's flagship regulation for mitigating climate change, there was ample scope for such an effect of the CCL, and our comparatively large estimates do not speak against this possibility. The CCL and plant exit ~~~~~~~~~~~~~~~~~~~~~~ The analysis so far has focused on how paying the full rate of the CCL affects various outcome variables in surviving plants. Rather than adjusting energy use and production at the intensive margin, there is a concern that firms might respond to the CCL by closing down plants altogether or by re-locating to non-regulated countries (“pollution havens”). After all, the substantial tax rebates granted under the CCA are intended to prevent such extensive-margin adjustments by energy intensive firms.42 This allows for fixed differences in the exit propensity between small and large plants and, since employment size and treatment status are strongly correlated (see Table 2), SMALL may also, to a large extent, control for fixed heterogeneity between treatment and control groups. Moreover, we use the interaction of SMALLi with a post-treatment dummy I{t > 2000} to instrument for CCLit. The idea behind this is (i) to use the fact that size influenced the decision to participate in a CCA and (ii) to rely on variation in size prior to our sample period so as to preserve the exogeneity of the instrument. The estimated coefficient α has the interpretation of a local average treatment effect (LATE). Since all the information needed to estimate Eq. (6) is available from the IDBR, we implement these regressions at the local unit level (see footnote 23 above). Table 6 reports the results from probit and IV probit models, along with the corresponding reduced-form (RF) and first-stage (FS) results. In each of the exit regressions, the coefficient on SMALL is positive and significant, confirming the already well-documented empirical regularity that smaller firms are more likely to exit. The simple probit model yields a positive and significant coefficient estimate on CCL which implies a 5.9% increase of the exit probability at the average CCL plant. Notice that this effect is not necessarily causal. In fact, the positive coefficient is consistent with a reverse-causality explanation according to which, plants that anticipate to exit in the near future do not sign a CCA because the tax savings this generates over the remaining lifetime of the plant do not cover the fixed costs of certification to be paid upfront. Once we instrument for CCL status, the point estimate becomes statistically insignificant, as foreshadowed by the insignificant coefficient estimate on the instrument obtained in the reduced form. The first-stage regression coefficients show that our instrument is strongly correlated with CCL status. In sum, we find no evidence that the CCL had an impact on plant exit decisions. This finding is robust to the inclusion of industry controls and to splitting the sample by either energy or trade intensity as in Section 6.1 above.44 Our analysis has focused on exit decisions at the local unit level whereas the bulk of the variables used in Section 5 are only available at a slightly higher level of aggregation (the ‘reporting unit’ or ‘plant’). Since employment (and only employment) is available at both levels of aggregation, we re-estimate a version of Eq. (2) using employment data at the local unit level in order to verify that the results obtained at the reporting unit level are robust. Table 7 reports estimates of the CCL impact on employment in the full sample and when the sample is split according to energy and trade intensities, or size (defined as above at the reporting unit level). Our preferred specification includes a trend coefficient for the treatment group (NEPER × year diff) because we find it to be statistically significant for the high trade intensity group.45 As before, we do not find evidence of a detrimental effect of the CCL on employment, regardless of which way the data are cut.","There is a growing consensus that climate policy should aim to regulate GHG emissions efficiently across a broad range of economic sectors. While curbing industrial emissions must be an integral part of any such policy, there is surprisingly little empirical evidence on the impacts of large-scale regulations of industrial GHG emissions — let alone using market-based instruments. In this paper we have provided the first micro-econometric evaluation of a carbon tax on the manufacturing sector. Unlike simulation-based evaluations, our approach does not require making assumptions about counterfactual – “baseline” – trends in the outcome variable of interest. Instead, we compare changes in outcomes both over time and between plants that were subject to different tax rates. The “baseline” is hence given by the contemporaneous outcomes of plants that faced lower tax rates by virtue of being in a CCA. Our estimates of the impact of the CCL are thus purged of confounding factors that affect plant performance at the level of the economy, the region and the sector. Since we also control for self-selection into CCAs by exploiting exogenous variation in CCA eligibility rules, we interpret our estimates as the causal effect of the CCL on plant outcomes. We find robust evidence that the price incentive provided by the CCL led to larger reductions in energy intensity and electricity use than the energy efficiency or consumption targets agreed under the CCA. The tax discount granted to CCA plants has been justified as a means of preventing energy intensive firms from losing competitiveness in international product markets due to the unilateral implementation of the tax and to the lack of international harmonization. Although this has been widely argued, we find no discernible impact on employment, gross output or productivity across groups, and we cannot reject the hypothesis that the CCL had no impact on plant exit. Our results show that the introduction of a moderate tax on energy encourages electricity conservation and helps to reduce energy intensity in the manufacturing sector. This is in contrast to previous research that attributed substantial carbon savings to the CCA scheme on the basis of comparisons with counterfactual baseline emissions (Ekins and Etheridge, 2006; Barker et al., 2007; AEAT, 2004).46 While our research design arguably produces a more credible estimate of the effect of the CCL, it is clear that this effect is additional to any effect the CCA targets may have had on firm behavior. Our study constitutes a first step towards building an evidence base that informs policymakers about the impacts of climate change policies on industry. As more such policies are being implemented across countries, and as business microdata are becoming more abundant and easier to access, we expect that researchers will exploit the variation in policies and institutional settings to make important contributions to this evidence base. In the context of climate change policy in the UK, there are several issues that deserve attention in future research. First, it seems important to gain a better understanding of how plants achieved the substantial reductions in energy use that we measure. This will require gathering more qualitative information on the key drivers of energy conservation — be they technical, economic or managerial. This information could lead to the design of more sophisticated policy instruments. From a political economy point-of-view, an analysis of the bargaining over CCA targets and of compliance behavior of individual CCA facilities will provide valuable insights regarding the design of negotiated agreements. Finally, given the long-term nature of climate change, an important open question is whether a moderate energy tax such as the CCL can stimulate much-needed innovation to bring about substantial carbon reductions in the future."],["A burgeoning literature in economics has started examining the role of social norms in explaining economic behavior. Surprisingly, the vast majority of this literature has studied social norms in asocial decision settings, where individuals are observed to act in isolation from each other. In this paper we use a large-scale dictator game experiment (N = 850) to show that “peers” can have a profound influence on individuals’ perceptions of norms of fair sharing, which we elicit in an incentive compatible way. However, in contrast to these strong peer effects in social norms of fair sharing, we find limited evidence of the influence of norms and peers on actual sharing behavior. We discuss how these results can be explained by heterogeneity in normative views as well as in willingness to comply with norms. --------------------------------------------------------------------------------","We study the driving forces underlying one of the fundamental principles of human social behavior: fair sharing. While earlier explanations have focused on the role of other- regarding preferences and preferences for equality (see, e.g., Camerer, 2003, Chap. 2), we investigate a more recent account of fair sharing that relies on the concept of norm compliance: many people have an intrinsic preference to conform to what is collectively perceived as “socially appropriate” and are willing to sacrifice material gain in order to comply with such norms.1 In fact, social norms are thought to drive behavior in a variety of social contexts (e.g., Elster, 1989; Bicchieri, 2006; López-Pérez, 2008; Krupka and Weber, 2013). A number of recent experimental studies use a norm compliance framework to explain behavior across several settings, including dictator games (Krupka and Weber, 2013; Krupka et al., 2017; Kimbrough and Vostroknutov, 2016), third-party allocator games (Barr et al., 2015), gift-exchange games (Gächter et al., 2013), oligopoly games (Krupka et al., 2017), public good, trust and ultimatum games (Kimbrough and Vostroknutov, 2016). However, nearly all of these studies of social norms focus on tightly controlled, but surprisingly asocial decision environments, where individuals face neutral and abstract decision situations, under full anonymity, and in complete isolation from other decision- makers. While the use of contextually sterile decision environments is one of the hallmarks of experimental control, we also notice that contextual variables – from the framing of the decision task to the presence and behavior of other decision-makers in the decision setting – play a crucial role in nearly every conceptual account of social norms. Minimal variations in the context can profoundly change individuals’ perception of the nature of the decision situation and the underlying norms of conduct (Bicchieri, 2006). This highlights the importance of studying the interaction between contextual variables and norm compliance. In this paper we take a step in this direction by systematically studying the influence on norm compliance in fair sharing of one specific contextual variable: the presence of “peers”, i.e. other decision-makers, in the decision setting faced by an individual. We believe that understanding the influence of peers on individual decision-making is important for a number of reasons. First, information about peer behavior is typically available in many natural social settings, where individuals do not act in social isolation. On the contrary, people often have the opportunity to interact with others and observe their choices before making a decision. Thus, studying the influence of peers on individual decision-making is inherently relevant for understanding the general dynamics of human social interactions. Second, the study of peer influence is of theoretical interest because peers are an important determinant of norm-driven behavior in most conceptual accounts of norm compliance across the social sciences. For instance, in economics, Sugden (1998) argues that observing instances of norm-compliance or norm- breaking can reinforce or weaken the expectations that the norm ought to be followed. In social psychology, Cialdini et al. (1990) contend that the behavior of peers exerts normative influence on individual behavior by shaping what individuals perceive as typical or normal behavior in a given situation (the “descriptive norm”). In philosophy, Bicchieri (2006) proposes that whether or not a norm will be followed depends partly on “normative expectations” (whether the individual expects that sufficiently many others expect him or her to comply), and partly on “empirical expectations” (whether the individual expects that sufficiently many others will comply). Sociologists Lindenberg and Steg (2013) argue that the behavior of others can shift the weights that individuals place on the normative- goal (following social norms) relative to the more self-centered hedonic and gain goals (need satisfaction and resource accumulation). Despite the large theoretical literature on the importance of peers for norm-driven behavior, the empirical evidence is scant. In many of the settings where peer effects have been documented empirically (e.g., Keizer et al., 2008; Shang and Croson, 2009; Bicchieri and Xiao, 2009; Krupka and Weber, 2009; Gächter et al., 2012; Falk et al., 2013; Thöni and Gächter, 2015), other behavioral forces may explain the correlations between individuals’ and peers’ actions observed in the experiments.2 Even in settings where the observed data patterns are difficult to reconcile with alternative explanations (e.g. McDonald et al., 2013) and results are strongly suggestive that the presence of peers affects norms, the lack of direct data on how peers affect normative considerations makes it difficult to identify whether the observed impact of peers’ actions on behavior is mediated by corresponding shifts in the normative evaluation of actions. In this paper we present a new set of dictator game experiments that measure the influence of peers on both actual sharing and norms of sharing using the incentive-compatible norm-elicitation task by Krupka and Weber (2013).3 Our experiments set us apart from the existing literature on peer effects mentioned above, in that we are able to explicitly identify the linkages between peers’ actions, normative views, and individual sharing behavior. In this aspect our paper is related to Gächter et al. (2013), who, however, study peer effects in norms and behavior in a gift exchange game. They find that peer effects in norms do not explain the observed peer effects in actual gift exchange. While these results cast some doubt on the importance of norms for peer effects, it would be premature to base judgment on the importance of norm following solely on the study of one specific decision setting and one specific social norm. It is indeed unclear whether the results from the gift exchange game may also extend to other settings and norms, as it may be the case that the influence of peer behavior is more decisive for norms of fair sharing than for reciprocal gift exchange. Moreover, all the experiments reported in Gächter et al. (2013) are based on gift exchange games where the decision- makers observe the decisions of a peer before making their own choices. In this sense, it is not obvious that their experiments allow assessing the causal impact that the presence of peers may have on norms and behavior, because their study lacks a treatment without peers. In this paper, we study settings where the decision-maker is exposed to the influence of a peer as well as settings where the decision-maker acts in isolation from peers. This allows us to examine the causal influence that peers have on norms and behavior. Specifically, in our Peer treatment subjects play a sequential three-person dictator game, where two dictators can transfer money to one recipient. The dictators move sequentially and thus the second dictator can observe the transfer made by the first dictator (the “peer”) before making her own transfer decision. In contrast, our NoPeer treatment is based on a two-person dictator game where there is no peer and her role is replaced with Nature: in this game, Nature moves first and randomly determines an endowment for the recipient; the dictator observes this endowment and then transfers money to the recipient. The crucial difference between the two treatments is thus that, while in the Peer treatment the recipient's wealth (prior to the dictator's transfer) is determined by a peer, in the NoPeer treatment it is determined by chance and there is no decision- maker other than the dictator present in the decision context. Furthermore, to systematically investigate the extent to which the influence of peers on normative considerations and behavior depends on the nature of the underlying norms, our study examines two payoff-equivalent, but differently framed, versions of the dictator game. In one version the dictator can give money to another player, while in the other version the dictator can also take money from the other player. Krupka and Weber (2013) have used similar versions of the dictator game to measure the influence of norms on dictator's behavior.4 They have shown that these “give” and “take” versions of the dictator game produce stark differences in the amounts of money that dictators share with recipients. Moreover, they explain these differences by the fact that the norm that governs behavior in the “give” version of the game is substantially different from the norm that applies to the “take” game. Hence, we use give/take framing to study the extent to which the influence of peers depends on the nature of the norm (norm of giving vs. norm of taking). To summarize, our study is based on four treatments, using a 2 × 2 factorial design where we vary the frame of the game (Give vs. Take) and whether a peer is present or absent (Peer vs. NoPeer). For each treatment, we conduct two types of experiments, a norm- elicitation experiment and a behavioral experiment. In the norm-elicitation experiment, we follow Krupka and Weber (2013) and measure in an incentive compatible way the extent to which the peer's behavior affects the perception of what constitutes socially appropriate behavior. In the behavioral experiment, we check how these variations in perceptions of social appropriateness translate into actual decisions. A total of 850 subjects participated in our experiments. Our norm-elicitation experiments reveal that the presence of peers has a systematic and strong influence on the perceptions of social appropriateness. In the Peer treatment, ungenerous monetary transfers to the recipient are viewed as relatively more appropriate when the peer is also ungenerous towards the recipient. However, when the same levels of recipient's wealth have been determined by chance (NoPeer treatment), the relation between recipient's wealth and appropriateness is reversed: ungenerous transfers are viewed as relatively more appropriate when the recipient is wealthier (i.e. when the recipient has randomly received a larger endowment). Interestingly, we also find that the strength of these effects varies considerably across our two versions of the dictator game. The norm that governs behavior in the Take game is much more stable and resilient to peer influence than the norm in the Give game. Based on the results of the norm-elicitation experiment, we should expect to observe systematic differences in the influence of peers’ actions (and hence recipient's wealth) on dictator's actual behavior across our experimental conditions. In particular, we should expect a positive relation between dictator transfers and recipient wealth in the Peer treatment, while a negative relation should emerge in the NoPeer treatment. Moreover, these treatment differences should be more pronounced in the Give than in the Take game. The results of our behavioral experiments are only partially in line with these expectations. While we observe that dictators in the NoPeer treatment significantly reduce their transfers when the recipient possesses larger endowments, there is, on average, no relation between dictator and peer transfers in the Peer treatment. Moreover, we do not detect any differences in the magnitude of these effects between the Give and Take conditions. The absence of a peer effect in the Peer treatment is consistent with the findings reported by Panchanathan et al. (2013). They also conduct a three-person dictator game experiment where two dictators decide sequentially how much to give to a recipient. They find that, on average, the amount given by the first dictator does not affect the second dictator's giving. At the individual level, they observe substantial heterogeneity in the second dictator's responses: while some dictators increase their giving in the amount given by the peer, others give less when the peer gives more, and others do not vary their giving with the peer's giving. We observe similar heterogeneity in our experiment. This suggests that a potential explanation for the limited support of the norm compliance model in our experiments may lie in the existence of conflicting views about what constitutes a norm in our setting. In Section 5 we examine this possibility in detail and show that there is considerable heterogeneity in the extent to which participants agree on what a norm is in our experiments as well as in the extent to which they are prepared to comply with it.","All our treatments are based on dictator game experiments. The Peer treatment is based on a three-person sequential dictator game where two dictators (D1 and D2) are matched with one recipient (R). Dictators move sequentially: D1 moves first and chooses a monetary transfer for the recipient; D2 observes the transfer chosen by D1 and then chooses a transfer. In the Give version of the game, D1 and D2 receive an initial endowment of £12 each, while the recipient is endowed with £0. Each dictator can then transfer an amount gi∈{D1, D2} ∈ {£0, £1, £2, £3, £4} from her endowment to the recipient. Monetary payoffs are computed as πi = £12 – gi for a dictator, and πR = £0 + gD1 + gD2 for the recipient.7 We study how D2’s behavior is affected by information about their peer's (D1) behavior, by comparing choices made in the Peer treatment with choices made in the NoPeer treatment, where the role of D1 is replaced with Nature. Thus, the NoPeer treatment is based on a two-person dictator game, where one dictator is matched with one recipient. In the Give version of the game, the dictator receives an endowment of £12 while the recipient's endowment, E = {£0, £1, £2, £3, £4}, is randomly determined by Nature. After observing the value of the recipient's endowment, the dictator transfers an amount g ∈ {£0, £1, £2, £3, £4} to the recipient. Payoffs are computed as πD = £12 – g for the dictator, and πR = E + g for the recipient. Note that in both treatments we observe decisions by dictators facing the same five possible situations, each corresponding to a different level of initial wealth of the recipient (£0, £1, £2, £3, or £4). The difference between the two treatments is that in the Peer treatment the recipient's wealth (prior to the dictator's transfer) is determined by the donation of another dictator, whereas in NoPeer the peer is absent and the recipient's wealth is determined at random.8 The corresponding Take versions of the games are analogously defined, except that the initial distributions of endowments differ relative to the Give version. In the Peer/Take game, D1 and D2 are endowed with £9 each, while the recipient is endowed with £6. Each dictator can give/take an amount ti∈{D1, D2} ∈ {−£3, −£2, −£1, £0, £1} to/from the recipient. Payoffs are computed as πi = £9 – ti for a dictator, and πR = £6 + tD1 + tD2 for the recipient. Analogously, in the NoPeer/Take game the dictator is endowed with £9, while the recipient's endowment is randomly determined from the set E = {£3, £4, £5, £6, £7}. The dictator transfers an amount t ∈ {−£3, −£2, −£1, £0, £1} to the recipient, and payoffs are computed as πD = £9 – t for the dictator, and πR = E + t for the recipient. Thus, in both the Give and Take version of the games, dictators can implement exactly the same final payoff allocations between themselves and recipients. However, the Give and Take games differ in whether these allocations can be obtained through “giving to” or “taking from” the recipient. For each treatment and each version of the game, we conducted two types of experiments: a norm- elicitation experiment and a behavioral experiment. The norm-elicitation experiment is based on the task introduced by KW. Subjects were given a description of the five possible situations faced by either D2 in the Peer treatment or the dictator in the NoPeer treatment. We conducted separate sessions for the Give and Take versions of the games. In each case, subjects had to evaluate, for each of the five situations, the appropriateness of each of the five actions that were available to the dictator. For example, subjects in the Peer/Give condition read a description of a situation where D2 observes that D1 has given £0 to the recipient and must decide whether to give £0, £1, £2, £,3 or £4. For each of the possible five actions available to D2, subjects were asked to rate, on a six-point scale, whether that action was “socially appropriate” and “consistent with what most people expect [a dictator] ought to do”, or “socially inappropriate” and “inconsistent with what most people expect [a dictator] ought to do”.9 Similarly, subjects rated the appropriateness of each of the five dictator actions in the other four situations where D1 had given £1, £2, £3 and £4 to the recipient.10 Similar to KW, subjects received a monetary reward if their appropriateness judgments matched the judgments provided by other subjects in their session. In particular, they were told that one of five possible situations, and one of the five actions available to the dictator in that situation, would be selected at random at the end of the session. Subjects were paid £7 (in addition to a £5 show-up fee) if their appropriateness rating for the selected action matched the rating of one other randomly selected subject in the session.11 Thus, as in KW, subjects were given incentives to reveal what they perceived to be the collectively-shared judgment of appropriateness of the actions they evaluated, and not their own personal judgment. Hence, a subject in the norm-elicitation experiment plays 25 coordination games over appropriateness ratings (with no feedback between games) with another randomly selected participant.12 We conducted the behavioral experiments with subjects who had not participated in the norm-elicitation task. Subjects were randomly assigned to either the Peer or NoPeer treatment. In each treatment, half of the subjects participated in the Give game, and the other half in the Take game. In all cases, we paid subjects a £2 show-up fee in addition to any earnings made in the experiment.13 At the beginning of the experiment we matched subjects randomly into groups and assigned a role. In the Peer treatment subjects were matched in three-person groups and assigned the role of D1, D2, or Recipient. In the NoPeer treatment, subjects were matched in two-person groups and assigned either the role of dictator or recipient. Subjects then played a one-shot version of the dictator game, either in the Give or Take frame. We elicited subjects’ choices using the strategy method (Selten, 1967). That is, dictators in the role of D2 in the Peer treatment and dictators in the NoPeer treatment were asked to make one decision for each of the five possible sub-games of the game, corresponding to situations where D1 or Nature had endowed the recipient with £0, £1, £2, £3, or £4 (£3, £4, £5, £6, or £7 in the Take game).14 In total, we conducted 44 sessions with 850 subjects, recruited using ORSEE (Greiner, 2015). All sessions were conducted at the University of Nottingham using z-Tree (Fischbacher, 2007). Sessions lasted between 40 and 60 min. Table 1 summarizes the experiment design and reports the number of subjects who participated in each treatment and version of the game.","We start by presenting the data from the norm-elicitation experiments, to examine whether the behavior of peers influences the norms of fair sharing in our setting. We then turn to the behavioral data, and examine whether any differences in norms across conditions translates into differences in sharing behavior. Norm-elicitation experiments: the influence of peers on norms of fair sharing ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Fig. 1 reports the average ratings of appropriateness collected in the norm-elicitation experiments. We report the full distributions of appropriateness ratings in Online Appendix B and an analysis of the variation of ratings in Section 5.1 (Fig. 3 in particular). The average social appropriateness ratings of dictator transfers in the Peer treatment are shown in the top-left (Give game) and bottom-left (Take game) panels of the figure. The ratings of the NoPeer treatment are shown in the right panels of the figure. In each panel, we show ratings for each of the five possible situations faced by a dictator, corresponding to the five possible levels of wealth of the recipient determined either by D1’s transfers (Peer) or by chance (NoPeer).15 Several interesting patterns can be observed. First, in all five situations and in all treatments and versions of the game, the appropriateness of transfers increases in their generosity: sharing the highest amount available (“give £4″ in Give; “give £1″ in Take) is always considered the most appropriate option. Similarly, in all cases, the least appropriate choice is the level of sharing that maximizes the dictator's payoff (“give £0″ in Give; “take £3″ in Take).16 Second, the level of the recipient's wealth generally influences the perception of what constitutes an appropriate level of sharing. These differences are, however, much more marked in the Give than in the Take game. Thus, the norms of fair sharing in the Give game seem much more malleable than the corresponding norms in the Take game. Third, and most importantly, the levels of the recipient's wealth influence ratings of appropriateness differently depending on whether these levels have been determined by the transfers of another dictator (Peer treatment) or by chance (NoPeer treatment). In the Peer treatment giving little to the recipient is generally viewed as less appropriate when the recipient's wealth is large (i.e., when the peer has been generous) than when a recipient's wealth is small (i.e., when the peer has also given little).17 However, in the NoPeer treatment the relation between appropriateness and recipient's wealth is reversed: giving little to the recipient is viewed as more appropriate when the recipient's wealth is large (i.e. when Nature selects a large endowment) than when it is small.18 We examine these patterns more formally using OLS regressions, reported in Table 2. In Model I we use data from the Peer treatment only, whereas in Model II we use data from the NoPeer treatment only. In both regressions, the dependent variable measures the appropriateness of the dictator's transfers in the five different situations. We regress this on the amount that the dictator transfers to the recipient (“Amount transferred by Dictator”), the amount that the peer (Peer treatment) or Nature (NoPeer treatment) transfers to the recipient (“Amount transferred by Peer/Nature”), and an interaction between these two variables. Moreover, to gauge the extent to which the influence of peers varies across the Give and Take games, we also include a dummy variable taking value 1 for observations in the Take game, and an interaction between the Take dummy and the “Amount transferred by Peer/Nature” variable. The regressions reveal that in both the Peer and the NoPeer treatments more generous transfers by the dictator are viewed as more appropriate than ungenerous transfers. The effect of increasing the dictator's transfer on its evaluation of appropriateness is 0.359 + 0.019 * “Amount transferred by Peer” in the Peer treatment and 0.411 − 0.009 * “Amount transferred by Nature” in the NoPeer treatment. In both cases, the effect is positive for any possible amount transferred by the peer or Nature. To gauge how changes in the recipient's wealth affect the judgments of appropriateness of the dictator's transfers, we need to inspect the coefficients of the variable “Amount transferred by Peer/Nature” and the interaction term “Amount transferred by Dictator * Amount transferred by Peer/Nature” (as well as the interaction with the Take dummy, for the Take game). In the Peer treatment, the peer's generosity negatively influences the judgments of appropriateness of the dictator's transfers. This effect is particularly marked for ungenerous dictator's transfers, while the influence of peers wanes for more generous dictator transfers, as indicated by the positive and significant coefficient of the interaction term between the “Amount transferred by Dictator” and “Amount transferred by Peer/Nature” variables. In contrast, in the NoPeer treatment the judgments of appropriateness of the dictator's transfers become more lenient the higher is the endowment that Nature transfers to the recipient. Again, this effect is particularly marked for ungenerous dictator transfers and it diminishes as dictators transfer more money to the recipient, as indicated by the negative and significant coefficient of the interaction term. Finally, in both treatments, the impact of the recipient's wealth on norms is significantly weaker in the Take than in the Give game. This can be seen by noticing that, in both the Peer and the NoPeer treatments, the coefficient of the interaction term “Amount transferred by Peer/Nature * Take” takes an opposite sign relative to the “Amount transferred by Peer/Nature” variable. In both cases the effect is significant at least at the 5% level. To account for the ordinal nature of the norms data, we ran additional ordinal probit regressions. The results are similar to those reported in Table 2. Moreover, we complement the regression analysis from Table 2 by a further specification in which we pool the data from the Peer and NoPeer treatment and include a Peer treatment dummy as well as all relevant interactions. The results show that the differences between the Peer and NoPeer treatment discussed above are highly significant. Both supplementary regression tables are presented in Online Appendix D. Taken together, these results show that the behavior of peers can have a strong, systematic influence on the perception of what constitutes a norm of fair sharing in our setting. What are the behavioral implications of these results? Assume that, as in the model sketched in Section 2, individuals trade off monetary payoff and norm-compliance utility, whereby individuals gain utility from choosing actions that are viewed as socially appropriate and suffer a disutility from choosing socially inappropriate actions. Within this framework, one would expect a negative effect of the recipient's endowment on giving in the NoPeer treatment: norm-compliant dictators should be more generous when the recipient possesses a small endowment because then ungenerous transfers are more inappropriate (and hence result in stronger disutility) than when the recipient has a large endowment. In contrast, one would expect a positive relation between the peer's and the dictator's transfers in the Peer treatment. In this case, ungenerous transfers are more appropriate when the recipient is poorer than when the recipient receives a larger transfer from the peer. Moreover, we would expect these effects to be stronger in the Give than in the Take version of the game. We summarize these behavioral predictions as follows: Hypothesis 1: In the NoPeer treatment, dictator's transfers correlate negatively with the recipient's initial wealth. Hypothesis 2: In the Peer treatment, dictator's transfers correlate positively with the amount that the recipient received from the peer. Hypothesis 3: These effects are stronger in Give than in Take games. In the next sub-section we present the data from our behavioral experiments to examine the extent to which the observed variations in social appropriateness of transfers translate in differences in behavior. Behavioral experiments: the influence of peers on sharing behavior ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Fig. 2 shows the average monetary transfers made by dictators in the Peer (left panel) and NoPeer (right panel) treatments across the five possible sub-games of the game. In each panel the figure reports the average transfers made in the Give (dark bars) and Take (light bars) versions of the games. In the Take game, transfers have been rescaled to give a score between £0 and £4, to ease comparability with the Give game.19 The figure shows that there is on average no clear relation between the dictator's transfers and the recipient's wealth in the Peer treatment, both in the Give and Take versions of the games. Thus, whether or not the peer is generous with the recipient does not seem to affect the dictator's sharing decisions. In contrast, a negative relation between dictator's sharing and recipient's wealth seems to emerge in the NoPeer treatment, in both versions of the game. Thus, dictators seem to behave less generously towards recipients that have randomly received larger endowments. Table 3 reports OLS regressions of dictator's transfers on a variable measuring the amount that the peer (Peer treatment) or Nature (NoPeer treatment) transfers to the recipient, a dummy variable taking value 1 for observations in the Take game, and an interaction between the two variables. Similar to Table 2, we run separate regressions for the Peer treatment (Model I) and the NoPeer treatment (Model II). Model I confirms that there is on average no evidence of peer effects in the Give version of the Peer treatment: the amount transferred by the peer has no significant influence on the amount transferred by the dictator (p = 0.914). This similarly holds in the Take version, as indicated by the insignificant coefficient of the interaction term “Amount transferred by Peer/Nature * Take” (p = 0.727). In contrast, the recipient's wealth is negatively related to the dictator's transfers in the Give version of the NoPeer treatment. Model II shows that increasing the recipient's wealth by £1 reduces the dictator's giving by about £0.30, and the effect is significant at the 1% level. This negative relation between recipient's wealth and giving is not different across the Give and Take versions of the game, as indicated by the insignificant coefficient of the interaction term (p = 0.858). These results are only partially in line with the results of the norm-elicitation experiment. The negative relation between recipient's wealth and dictator's transfers in the NoPeer treatment is consistent with Hypothesis 1. However, the results of the norm- elicitation experiment also suggest that we should observe a positive relation between recipient's wealth and dictator's transfers in the Peer treatment (Hypothesis 2). Our data do not support this conjecture. Moreover, the norm-elicitation experiment suggests that the norm of fair sharing may be more malleable in the giving than taking setting (Hypothesis 3). However, we do not observe any difference between Give and Take games in the extent to which the recipient's wealth affects dictator's sharing. More generally, we see only small differences in dictator's behavior between the Give and Take games, and only in some subgames of the Peer treatment. This is interesting because KW have shown that using give/take frames in dictator games can produce strong differences in behavior. However, we cannot replicate this result: in our NoPeer treatment, which is most similar to the games used by KW, we do not observe any difference in dictator sharing between Give and Take games, despite the existence of differences in the norms that apply to these games (see Online Appendix C for further detail).","What can explain the observed discrepancies between the norm-elicitation and behavioral experiments? One striking aspect of the behavioral data is that we observe substantial heterogeneity at the individual level in the extent to which dictators are influenced by the level of wealth of recipients (see Online Appendix D for more details). About half of the dictators are not affected by the recipient's wealth and opt for the same monetary transfer across all five sub-games. A third of dictators reduce their transfer as the recipient's wealth increases, whereas about a tenth of dictators respond positively to increases in the recipient's wealth. Our findings are similar to those reported by Panchanathan et al. (2013) in a three-person dictator game that is closely related to our Peer/Give treatment. They find that about half of dictators do not respond to variations in the peer's behavior, a third give more when the peer gives less, and thirteen-percent give more when the peer gives more. This suggests that there may be substantial heterogeneity in the extent to which dictators are willing to comply with norms of fair sharing, or in the extent to which they recognize these norms as applicable. Alternatively, (at least some) dictators may be driven by other types of considerations (e.g. inequity aversion; guilt aversion), that may conflict with normative considerations and pull behavior away from compliance with norms of fair sharing. The next sub-sections investigate these potential explanations. Norm ambiguity ~~~~~~~~~~~~~~ A first possible explanation for our experimental results is that there may be substantial disagreement among subjects about what constitutes a norm of appropriate behavior in our experiments. As we discussed earlier (Section 4.1), in the absence of peers, individuals seem to apply a Rawlsian norm of fair sharing in our experiments, whereby the appropriateness of giving depends in part on the level of need of the recipient. When the peer is present, a different normative consideration is introduced as individuals recognize that the appropriateness of giving also depends on the peer's behavior (what Cialdini, 2001 refers to as the \"principle of social proof\"). Our norm experiments show that on average the principle of social proof overrides the Rawlsian norm of sharing in the Peer treatments (see Fig. 1). Nevertheless, it is conceivable that both norms remain active in our experiments, exerting divergent influences on behavior and potentially explaining the weak support for the norms model in the Peer treatments. To examine this, we take a closer look at the norms data. Recall that in the norm-elicitation experiment subjects could rate the appropriateness of actions on a scale with three levels of “inappropriateness” (very inappropriate, somewhat inappropriate, inappropriate) and three levels of “appropriateness” (very appropriate, somewhat appropriate, appropriate). Fig. 3 shows the percentage of subjects disagreeing with the majority view about the appropriateness of each action across the various situations that they rated.20 We say that a majority of subjects rate an action as appropriate (inappropriate) if the sum of the relative frequencies of the ratings “very appropriate”, “somewhat appropriate” and “appropriate” is greater (lower) than 50%. The light (red) bars indicate that there is a minority of subjects assigning one of the three levels of “inappropriateness” to an action, while the majority rated the action as appropriate. The dark (blue) bars show disagreement in the opposite direction (the majority view the action as inappropriate and a minority rates it as appropriate). For instance, the first dark bar in the top left panel of the figure shows that in the Peer/Give treatment 12% of subjects rated the action “give £0″ as appropriate in the scenario where the peer also gives £0, indicating that the remaining 88% of subjects rated it as inappropriate.21 To assess the presence of norm ambiguity, consider first the NoPeer/Give treatment (top right panel). In most cases, relatively few subjects (less than 20%) disagree on the social appropriateness of actions. The main source of disagreement among subjects is the action “give £2″, which between one- fifth and one-half of subjects view as inappropriate in contrast with the majoritarian view that the action is appropriate. Nevertheless, apart from this action, the general picture emerging from the NoPeer/Give treatment is that there is a reasonably low degree of ambiguity about the social norm in this setting. Consider now the Peer/Give treatment (top left panel). As in the NoPeer/Give treatment, there is little disagreement about the actions “give £0″ and “give £4″. Also as in NoPeer/Give, subjects tend to disagree on how to rate the action “give £2″. However, relative to the NoPeer/Give treatment, subjects also disagree more on how to rate the actions “give £1″ and “give £3″. For both actions there are at least some scenarios where about 40% of subjects disagree with the majority view. Moreover, the source of disagreement seems to be related to the behavior of the peer. For example, when the peer gives £1 most subjects view the dictator action “give £1″ as appropriate, presumably following the principle of social proof. However, 41% of subjects disagree and rate it as inappropriate, presumably following a Rawlsian norm similar to the one that subjects recognize in the NoPeer treatment. As another example, the dictator action “give £3″ is generally viewed as appropriate by a majority of subjects. However, when the peer gives £4, 42% of subjects rate this action as inappropriate, again presumably because this action compares unfavorably with the peer's action. Overall, the observed patterns of disagreement suggest that observing what a peer has decided to do may introduce some ambiguity about the social norm. Finally, Fig. 3 corroborates our previous observation that the norm in the Take treatment (bottom panels) is substantially less malleable than the norm in Give. For all actions and in both the Peer and NoPeer condition, very few subjects disagree with the majoritarian view about the appropriateness or inappropriateness of actions. The degree of agreement seems somewhat stronger in the NoPeer condition, but again the differences are small. To summarize, this qualitative analysis suggests that disagreement among subjects about what constitutes a norm of appropriate behavior can go some way in explaining the lack of support for the norm model in our experiments. Heterogeneity in norm compliance ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Another explanation for our experimental results is that there may be heterogeneity in preferences for norm compliance in the population of dictators we sampled for our experiment. Thus, even if norms of fair sharing were prominent and clear in the population, not all dictators would be willing to follow these norms. Moreover, the dictators’ willingness to follow norms may itself vary across treatment conditions. To explore these possibilities, we follow the econometric methodology used by KW and related papers and investigate the extent to which elicited norms can predict actual behavior in our experiments. Differently from previous papers, we use a mixed logit model (see, e.g., Train, 2003) that allows for heterogeneity in the concerns for norm compliance and allows us to estimate, for each treatment, the share of dictators that are in fact guided by a desire to follow social norms. Table 4 presents the results of the estimation. We estimate four different models, one for each treatment/game combination.22 In all models, the coefficient on own payoff is positive and highly significant, indicating that dictators are more likely to choose transfers that yield higher own payoffs. Turning to norm compliance, Table 4 reports the mean and standard deviation of the norm rating coefficients. Looking first at the estimates of the mean, the regressions confirm the limited success of the norms compliance model in explaining the behavioral data. In the Peer treatment (Models I and II) the average effect of norm ratings on the choice of monetary transfers is not significantly different from zero: on average, dictators do not choose transfers that are deemed more socially appropriate more often. In the NoPeer treatment (Models III and IV) the effect is positive and significant in the Give game, indicating that the average dictator is more likely to choose transfers that are more socially appropriate. The effect is, however, not significantly different from zero in the Take game. Other behavioral explanations: inequity aversion and guilt aversion ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ So far we have considered explanations related to the existence of heterogeneity in norm compliance or in the understanding of what constitutes a norm. However, it is also possible that participants are motivated by other types of behavioral considerations instead of (or in addition to) normative concerns. Here we consider two popular behavioral motives that may have particular bite in the context of our dictator games. First, distributional preferences may play a role, especially because our Peer and NoPeer treatments are based on three-person and two-person games respectively, and this affects the implications that choices have for the redistribution of payoffs across players. For example, when the recipient's wealth is £3 in the NoPeer/Give treatment, there is no action by the dictator that can equalize payoffs between the two players (dictator and recipient). However, when the recipient's wealth is £3 in the Peer/Give treatment, giving £3 to the recipient equalizes earnings between the dictator and the peer. Previous studies have found that payoffs of third parties have strong influences on sharing behavior, even in settings where the payoff of the third party is completely exogenous and cannot be affected by players’ decisions (e.g., McDonald et al., 2013). Thus, payoff comparison considerations may explain some of the differences between the Peer and NoPeer treatments. Can the Fehr and Schmidt model explain the patterns of choices in the behavioral experiments? It turns out that the model does not predict behavior in either of the treatments. In the NoPeer treatment, the model predicts no relation between the recipient's wealth and dictator's giving. This is because in our games the dictator is always at least as well off as the recipient, at all levels of the recipient's wealth and for all the actions available to the dictator. This implies that the model predicts that the dictator either gives nothing (if βi < 1/2) or gives £4 (if βi ≥ 1/2), regardless of the wealth of the recipient. In contrast with this prediction, our data from the NoPeer treatments show that dictators reduce their giving as the recipient's wealth increases. As for the Peer treatment, the Fehr and Schmidt model predicts that, if D2 gives any money to the recipient (which occurs when βi ≥ 2/3), the amount given is positively correlated with the peer's giving. This is because, in the three-person Peer games, D2 compares her payoff not only with the recipient but also with the peer. Thus, because of disadvantageous inequality aversion, D2 is willing to give money to the recipient only to the extent that the peer also gives money, so that her payoff does not fall behind the peer's payoff. Our data do not support this prediction and show no relation between the two dictators’ actions in the Peer treatment.23 A second potential motive that may play a role in our setting is guilt aversion (e.g., Charness and Dufwenberg, 2006). Guilt averse dictators suffer a disutility if they leave the recipient with less money than what the recipient expects to receive. Guilt aversion may predict differences in behavior between our treatments because in the Peer treatments dictators may adjust their beliefs about what the recipient expects to receive based on the giving of their peer. For instance, observing that the peer gives £4 to the recipient may induce dictators to adjust their beliefs upwards as they may interpret the peer's actions as a signal that the peer thinks that the recipient expects £4 from a dictator. This signal is instead unavailable to dictators in the NoPeer treatment. In order to test models of guilt aversion, one needs second-order beliefs of dictators about what recipients expect to receive. This is particularly important if one wishes to test whether these models are observationally different from models of norm compliance (see, for example, Krupka et al., 2017). Because our design is already quite complex, we have not elicited beliefs and so we cannot perform a formal test of guilt aversion as a potentially distinct explanation of our data. Nevertheless, if one plausibly assumes a positive correlation between dictators’ second- order beliefs and peer's giving (along the lines discussed above), then a positive relation between peer's and dictator's giving should emerge in the Peer treatments. At the aggregate level our data do not support this prediction. In this sense, we think that guilt aversion is an unlikely explanation of our behavioral results.","Our study shows that the behavior of others can have important effects on the way individuals perceive what constitutes socially appropriate behavior in a given situation. In our dictator game experiments, whether or not an action is viewed as socially appropriate partly depends on the extent to which another dictator (the “peer”) is willing to take it. These strong effects of peer behavior on norms do not translate, however, into corresponding effects in actual behavior in the aggregate. In particular, we do not observe a positive correlation between the dictator's and peer's generosity in the treatment where dictators receive information about peer behavior. Thus, generous peers do not breed more generosity, despite the strong impact of peer behavior on the average social acceptability of generous and ungenerous behavior.24 We discuss a number of possible explanations for the discrepancies between normative considerations and actual behavior observed in our experiments. We find evidence of heterogeneity in normative views that is related to the presence of peers: the peer's behavior introduces normative cues that are in contrast with the notion of fair sharing that subjects seem to hold when peers are absent (see McDonald et al., 2013 for related evidence). This conflict in normative views can explain why we find a large fraction of subjects unwilling to comply with the average view of appropriateness and why dictators fail to follow the example of peers. Thus, our results suggest that the extent to which peers reinforce or counteract pre- existing notions of appropriateness may be an important determinant of the strength of peer effects. Our results raise a number of interesting questions regarding the existing approaches to norm compliance (e.g., Krupka and Weber, 2013). The current focus on normative consensus (the average or most frequent notion of what is appropriate) may be limiting in contexts where there are conflicting normative views: understanding the interplay between heterogeneous norms and norm compliance seems crucial in order to explain behavior in such situations.25 In this sense, the use of within-subject experimental designs, where normative views and behavior are collected from the same subjects, may prove a useful research tool for further research in this area, since they would allow to correlate at the individual level behavior and beliefs about what constitutes a norm in a given situation.26 Another interesting question relates to the role of sanctions for norm compliance. Recent research has shown that individuals are willing to use direct and indirect punishment to enforce social norms at a cost to themselves even in one-shot interaction with strangers, and this can help explain why norms are adhered to (Balafoutas and Nikiforakis, 2012; Balafoutas et al., 2014). Punishment opportunities may also play a role in resolving norm heterogeneity, for instance if subjects are willing to enforce only some of the conflicting normative views that are present in the population, but not others. In our setting there was no possibility of norm enforcement and so we cannot test this hypothesis in our data, but this could be an interesting avenue for further research."],["Understanding what determines the truth-telling of economic agents towards their regulator is of major economic importance from banking to the management of common-pool resources such as European fisheries. By enacting a discard-ban on unwanted fish-catches without increasing monitoring activities, the European Union (EU) depends on fishermen's truth-telling. Using a coin-tossing task in an artefactual mail field experiment with 120 German commercial fishermen, we test whether truth-telling in a baseline setting differs from behavior in two treatments that exploit fishermen's widespread ill-regard of their regulator, the EU. We find, first, that fishermen misreport coin tosses more strongly to their advantage in a treatment where they are faced with the EU flag, and, second, that misreporting is consistent with behavior in other hidden tasks. We also find some supportive evidence for our first result in a conceptual replication with 1200 UK citizens who voted ‘leave’ in the Brexit referendum. Our findings imply that lying is more extensive towards an ill-regarded regulator and that policy needs to account for this endogenously eroding honesty base. --------------------------------------------------------------------------------","Although honesty is regarded as a virtue or even a moral duty (Kant, 1785), lying and deception permeate economic life (Gneezy, 2005). Studying truth-telling has accordingly become a focus of inquiry for economics.1 An area of particular public economic importance is the truth-telling of economic agents towards their regulating authorities—from the banking industry (Cohn et al., 2014), and tax reporting (Jacobsen and Piovesan, 2016; Kleven et al., 2011) to environmental regulation (Duflo et al., 2013). The case where the German car manufacturer Volkswagen systematically lied about cars’ emissions is but one prominent example. Faced with uncertainty about how honest economic agents are, regulators need to decide how much to invest in monitoring and how to devise appropriate sanctioning schemes for misbehavior. Appropriate monitoring and sanctioning mechanisms are especially crucial for the management of common pool resources (Ostrom et al., 1992; Rustagi et al., 2010), with the fishery as a prime example (Wilen, 2000; Stavins, 2011). Fishery management comes in many different forms around the globe. It ranges from stringent restrictions on fish catches using individual transferable quotas—as in New Zealand (Newell et al., 2005) or Iceland (Arnason, 2005)—to largely unregulated open-access fishing, as it is still the case for most high-seas fisheries. The costs of illegal, unreported and unregulated fishing are substantial and amount to US$ 10 to 23 billion per year (Global Ocean Commission, 2013). Due to its economic importance and the heterogeneity of its regulatory structures, the fishery has gained substantial interest in experimental economic work.2 This paper extends the scope of previous studies and investigates to what extent regulator framing affects truth-telling. Our study therefore adds a new dimension to effective regulatory policy. We present evidence from an artefactual mail field experiment that examines truth-telling of German commercial fishermen. German commercial fishing is regulated by the European Union (EU), which is the world's fourth largest producer of fish, under the European Common Fisheries Policy. The EU has recently enacted a ban on returning unwanted fish catches to the sea (also called “discard ban” or “landing obligation”), as the practice of discarding ensues substantial costs to the public.3 The change in legislation has, as of yet, not been combined with more stringent monitoring. The regulator, and scientists assessing the status of fish stocks upon which recommendations for fishery management are based, thus depend on fishermen's truth- telling. Continuing to discard unwanted fish catches remains the individually optimal choice for fishermen in the present regulatory regime unless the regulator enforces the new policy. This, however, would require costly monitoring and sanctioning mechanisms.4 This trade-off for the regulator between more costly monitoring and reliance on regulatee's honesty is not only relevant in the fishery for the newly enacted European “discard ban” or compliance with fishing quotas, but is present more generally, including the previously discussed cases of banking, tax reporting and environmental regulation. For studying to what extent fishermen might tell the truth towards their regulator, we conduct a coin-tossing game in a mail field experiment targeting all commercial fishermen in Germany. Adapting the 4-coin toss game of Abeler et al. (2014), we ask fishermen to toss a coin 4 times and report back their number of tail tosses. For each reported tail toss, they receive five Euros. In a between-subjects design, we test whether truth-telling in a baseline setting differs from truth-telling in two further treatments with different EU framings, where, first, the EU flag is made salient on the instruction sheet, and, second, a framing that states additionally that the European Commission has funded the research. Based on a simple model of reporting behavior of fishermen that considers internal Nash bargaining among a pay-off maximizing ‘selfish self’ and a ‘moral self’, we hypothesize that the salience of the EU regulator may increase the bargaining power of the ‘selfish self’ vis-à-vis the ‘moral self’ and thus decrease overall lying costs if the EU is ill- regarded. The fishery is an ideal test case for studying how truth-telling behavior may be affected by regulatory framing, as there is well-documented and wide-spread contempt among fishermen concerning stricter EU fishing regulation. We confirm the almost entirely negative view of the EU prevalent among European fishermen (documented for UK fishers by McAngus, 2016) for our field experimental setting in Germany: Besides ample anecdotal evidence, our survey results indicate that the vast majority of participating fishermen have a low trust in the EU, while this is only the case for about a third of a student control group. If regulator framing impacts truth-telling, we will therefore expect an almost uniform direction of the effect. To study the robustness of our findings, we conduct a similar experiment with a population that is similar to the fishermen with respect to the (negative) stance towards the EU: Brexiteers. Using a large sample of UK citizens we conduct the experiment with 1200 individuals who reported to have voted ‘leave’ in the Brexit referendum in previous questionnaires. We find that fishermen misreport coin tosses to their advantage, albeit to a lesser extent than standard theory predicts. As hypothesized, misreporting is larger among fishermen who are faced with the EU flag. Our main effect is supported by further regression analyses. We also find some support for our main result in the conceptual replication with Brexiteers. Furthermore, we find that misreporting by fishermen is consistent with behavior in other hidden tasks involving a sent-along coin and the possibility to cheat in a competition task, suggesting some more general validity of the coin toss findings. Overall, our results imply that lying is more extensive towards an ill-regarded regulator. We close by discussing further policy relevance of our results.","The fishery has economic relevance in the German coastal regions at both the North Sea and Baltic Sea. According to the European Union's Common Fisheries Policy (CFP), the Council of Ministers of the European Union and the European Parliament set fishing quotas for the German fisheries. The German Federal Office for Agriculture and Food distributes the national catch quotas to fishing organizations or individual fishermen. Monitoring and enforcement of compliance are the duty of EU member states, and ultimately of the federal states in the case of Germany. A total of 896 commercial fishermen, owning 1465 fishing vessels (German Fishery Association, 2015), are registered at the German Federal Office for Agriculture and Food as holders of catch permits for the North Sea or Baltic Sea. Cutter type trawlers and coastal vessels constitute the core of the fleet with 300 boats. Small coastal fishing with passive gear such as gill nets and fish traps on vessels of less than 12 m length, composed of 1139 vessels, is predominantly operated at the Baltic coast. The German fishing fleet also includes seven deep-sea trawlers and two special vessels for pelagic fishing that operate in long distant waters, and 46 shell- and other special boats. Fig. 1 depicts a map of Germany's coastal regions, where the red balloons indicate the zip codes of fishermen who have participated in the experiment. The recent economic literature on honesty and lying has substantially advanced our understanding on what determines when and to what extent individuals lie. Abeler et al. (2019) conduct a large-scale meta-analysis of studies using coin-tossing and die-rolling tasks. This meta- analysis shows that, on average, individuals lie to some, but not to an exhaustive, extent and that the extent of lying does not seem to increase with the stakes. This paper contributes a new dimension to the analysis of truth-telling behavior: by studying how the salience of the regulator, who depends on truth-telling behavior in the policy context, affects the behavior of those being regulated. To this end, we adapt the 4-coin-tossing game of Abeler et al. (2014) for our mail field experiment. The fishermens’ task was to toss a one Euro coin exactly 4 times, and report their result in a table printed on the instructions sheet. For each instance they reported that the winning toss “tails” (in German “Zahl”, meaning “number”) laid on top, they received 5 €. A key feature of this task is that lying can be detected on aggregate when examining the distribution of decisions, but not on the individual level. Thus, depending on luck and honesty, each fisherman received between 0 and 20 € for this task. Besides the different sample pool, a major difference to the previous study is that Abeler et al. (2014) conducted their 4-coin experiments via telephone or in the lab and the decision whether to report truthfully or to cheat was immediate, while our subjects had several weeks to decide on whether to report honestly or to lie. In addition to previously studied effects, we hypothesize that the salience of the regulator affects individual lying costs. Salience of the regulator, in this case the EU, may decrease (increase) the ‘selfish self's’ bargaining power αi if the EU is well (ill) regarded. In our experiment we take advantage of the well-documented and wide-spread contempt among fishermen concerning stricter EU fishing regulation over the past decade.8 That is, we unambiguously predict an increase in the ‘selfish self's’ bargaining power αi if the salience of the regulator matters for truth-telling.9 In order to test our prediction, we sent out three versions of the instructions in a between- subjects design: (i) a baseline setting (‘Baseline’) in which only the logos of our two institutions present on the letterhead, (ii) a version where the EU flag is made salient in the letterhead of the instruction sheet (‘EU_Flag’), and (iii) an additional treatment where the framing states that this research has been funded by the European Commission (‘EU_Flag_Funding’). These framings were included on all three experimental sheets.10 Fig. 2 depicts the three letterheads and Appendix A includes the experimental instructions. Based on the insights from previous studies on lying behavior summarized in Abeler et al. (2019) and our treatments regarding the new regulatory dimension, we test three main hypotheses: Fishermen report greater tail-tosses than the truthful distribution, but do not fully misreport in the Baseline treatment. The standard economic hypothesis of pure selfishness is that fishermen report their own payoff-maximizing option, i.e. every fisherman would report 4 times tails. This hypothesis has been called into question by recent empirical evidence on various lying costs (e.g. Fischbacher and Föllmi-Heusi, 2013; Gneezy et al., 2018; Abeler et al., 2019). We therefore expect that fishermen, on average, report coin toss results in between the expected outcome of 2 times tails if all fishermen reported truthfully and the payoff-maximizing outcome of 4 times tails. Explanations for not reporting four winning tail tosses may include individual lying costs and internalized reputational costs for the profession. It may also mirror fishermen's professional behavior of misreporting somewhat instead of lying to the full extent, for example declaring some part but not all of their bycatch. Fishermen report less truthfully in the EU_Flag treatment compared to the Baseline treatment. As documented above, there is evidence for a widespread antipathy towards the EU among German fishermen, as most of new regulations by the EU have been regarded as burdensome for the fishermen. This makes the context of our study very useful to test Hypothesis 2, compared to cases in which the attitude towards the regulator is ambiguous. We therefore hypothesize that the presence of the EU flag will increase the bargaining power of the ‘selfish self’ relative to the ‘moral self’ thus decreasing lying costs and that fishermen in this treatment will thus report less truthfully out of ill-regard towards their regulator. Fishermen may also perceive the difference in the Baseline and the EU_Flag treatment as a difference in wealth of the specific institutions and the research institutions being backed by the EU. This may affect truth-telling, as previous research has shown that costs to others matter for lying behavior (e.g. Gneezy, 2005). In an attempt to disentangle this effect from the direct effect of a particular attitude towards their regulator, we include the third EU_Flag_Funding treatment. Fishermen report even less truthfully in EU_Flag_Funding compared to the EU_Flag treatment. We hypothesize that fishermen may regard the additional informational cue as an indication that there is plenty of funding available to those conducting the study. This may reduce the moral cost of lying, reducing the ‘misreporting aversion’ of the ‘moral self’, and lead fishermen to report less truthfully. Fishermen may also regard the provided information as an opportunity to acquire some of the EU's funds to compensate for the regulatory burdens imposed on them, thus giving more bargaining power to the ‘selfish self’, and leading fishermen to report less truthfully as well. We show below that the experiment rejects hypothesis 3. We come back to this issue and possible hind-sight explanations in the discussion in Section 5. To examine truth-telling of fishermen towards their regulator, we targeted all commercial fishermen in Germany in a mail field experiment. Due to rigorous data protection by the German Federal Office for Agriculture and Food, the address data of fishermen were not available to us. For the purpose of our study, the Thünen Institute of Sea Fisheries, the national fishery research institute responsible for carrying out fishery surveys, sent out the study documents to all 896 fishermen on our behalf. We prepared the envelopes with the survey materials, including stamped return-envelopes, at the University of Kiel. We then delivered the envelopes to the Thünen Institute and were present when the address data was added. The envelopes were sent out on Friday, December 4, 2015, and the closing date for the experiment was January 31, 2016. We assigned anonymous ID numbers to 1200 prepared surveys, which were numbered according to their treatment cell. After having randomly shuffled all envelopes, 896 of these envelopes were sent out to fishermen by the Thünen Institute.11 The experiment material consisted of 7 pages, including a cover letter, three experimental tasks with one page each, a two-page questionnaire and a sheet for payment information. Appendix A contains an English translation of the material. Besides the coin- tossing task, it includes an experimental task to elicit fishermen's risk preferences, and an experimental task on competitiveness.12 Fishermen were told that the payment for participating in the study was limited to 100 €, with an expected payoff of 50 € for around 30 min of work. Payment was made via bank transfer or by check via regular mail. To ensure availability of a coin to toss, we enclosed a 1 € coin that we stuck on the page of the task (see Appendix B). We complement our main study with three additional data collection exercises to examine its relationship to the economic literature on lying and its robustness. First, regarding the relationship to the literature, we examine how our subject pool of interest—professional fishermen—relate to the conventional subject pool of university students. We therefore ran the Baseline treatment of our mail experiment also with 50 business and economics undergraduate students at the University of Kiel at the same time, 44 of whom participated.13 Second, to examine the robustness of our results regarding possible attrition issues we ran an additional online experiment with 717 student subjects from the University of Kiel subsequent to the main study including two treatments: positive and negative framing regarding the fishery policy of the European Union and its effect on trust towards the EU. The reason is that mail field experiments may suffer from substantial attrition and we are not able to rule out by design that attrition in our fishermen study may depend on the treatment. The additional online experiment allows us to investigate the potential importance of attrition for the treatments in our main study and offers secondary causal insights into the effect of positive and negative framing of the policy on trust in the European Union. Third, to conceptually replicate and further investigate our main hypotheses and potential confounds, we conducted another online experiment with 1200 Brexit voters. We describe the details of this additional experiment in Section 4.","We received 136 responses by fishermen, amounting to an overall response rate of 15%.14 Of those, 120 responses included results for the coin-tossing task (see Table 1 for descriptive statistics),15 which were provided by fishermen located in 58 different ports.16 Aggregating all of our three treatments, we find that overall reporting by fishermen differs significantly from the truthful distribution as well as from payoff- maximization: fishermen report to have tossed 2.46 winning tails on average. This indicates substantial lying costs in line with the previous literature. A fraction of 10.83% of fishermen report that they have obtained four times tails, and 42.50% report three times tails. The distribution of reported outcomes is statistically highly distinguishable from both the payoff-maximizing outcome as well as from the truthful distribution. Two-sided binomial tests of the expected truthful against the observed frequency for 3 tails and for the payoff maximizing decision of 4 reported tails yield p = 0.000 and p = 0.055, respectively. In particular, we find reporting of 3 tail tosses at the expense of reporting 0 or 1 coin toss. The latter differs from the truthful distribution significantly (two-sided chi-squared test on combined 0 and 1 reports: p = 0.000). We therefore confirm Hypothesis 1 and previous findings in the literature. Next, we analyze the effects of our treatments on truth-telling.17 Fig. 3 shows the mean reported tail tosses across the three treatments. In the Baseline treatment, as well as in the EU_Flag_Funding, fishermen report an average coin toss result of 2.38 winning tails. In the EU_Flag treatment the average coin toss result was 2.64 tails. While this qualitatively suggests that fishermen over-report tail tosses in the EU_Flag treatment as compared to the Baseline, this is not statistically significant (two-sided t-test: p = 0.191). There is virtually no difference in mean reported tail tosses between Baseline and EU_Flag_Funding. Next, we examine tail toss response across treatments more closely. As Fig. 4 shows, no fisherman in the EU_Flag treatment reported 0 tail tosses, fewer fishermen reported 1 tail tosses compared to the Baseline treatment (8.33% vs. 11.90%) and more fishermen reported 4 tail tosses (16.67% vs. 4.76%). We find that fishermen in EU_Flag over-report tail tosses that yield the highest payoff (two-sided chi-squared test: p = 0.084). This finding provides some confirmation for Hypothesis 2: The salience of the regulator does seem to play a role for truth-telling and the wide-spread ill-regard for the EU seems to translate into stronger over-reporting of tail tosses. We find no material and significant differences between the EU_Flag and the EU_Flag_Funding treatments in terms of 4 tails reporting (two-sided chi-squared test: p = 0.547). However, we find that fishermen in the EU_Flag_Funding treatment report significantly more 0 and 1 tail tosses (combined: 23.81% in the EU_Flag_Funding vs. 8.33% in the EU_Flag treatment; two-sided chi-squared test: p = 0.067). These findings reject Hypothesis 3. Specifically, we do not find support for the ‘wealth-of-funding-institutions’ or ‘taking-back from the EU’ hypotheses as fishermen in the EU_Flag_Funding do not report more 4 or combined 3 and 4 tail tosses. We discuss this finding in Section 5. To further scrutinize hypotheses 2 and 3, we run a regression analysis of overall tail toss reporting with the two treatments and additional controls (see Table 2). Specification (1) includes all eight covariates that are significantly correlated with tail toss reporting in univariate regressions. Specification (2) also includes the standard demographic variables income and education level as additional controls. Since many fishermen did not respond the question on the probability of their income increasing over the next five years, which causes the low number of observations in specifications (1) and (2), we run an additional regression specification (3) dropping this variable. The regression analysis shows that across the three specifications EU_Flag_Funding changes sign. Indeed, we find that it is very far from significantly explaning different overall tail toss reporting compared to the baseline, with p = 0.808, p = 0.543 and p = 0.784, respectively. We find that the only consistently significant explanatory of variable of tail toss reporting is the EU_Flag, with a p-value in specification (1) of p = 0.022, in specification (2) of p = 0.005, and in specification (3) of p = 0.074. This provides further support for our Hypothesis 2. We summarize Fishermen over-report more severely when facing the EU flag. Our study also included two novel ‘hidden’ tasks that offer the possibility to underscore truth-telling or lying behavior. First, we deliberately left the ownership about the one Euro coin that we included on the coin tossing decision page unclear. A related aspect of fishermen's fidelity is thus whether they sent back the coin with their decision sheets. We find that the 22 fishermen who sent back the coin report a coin toss result of 2.23 tails on average, compared to 2.51 tails for those who did not send back the coin (see Fig. C.1 in Appendix C). This difference is not significant (two-sided t-test: p = 0.205), yet tentatively suggests consistent behavior between the coin-tossing task and this hidden measure and therefore some external validity. Second, we conducted a separate task to measure fishermen's competitiveness using a real production task where fishermen have to produce paper shreds by hand from an A7-sized (74 × 105 mm) piece of paper. Fishermen decided on whether they want to be paid 0.05 € per piece, or whether they want to play competitively and receive 0.15 € per piece if they perform better than a randomly drawn other participant. As the A7-sized paper we sent the fishermen was of standard white format, dishonest fishermen could add additional alien paper shreds to increase their payoffs. To control for this possibility to cheat, we measured the weight of the returned paper shreds on an analytical scale from the physical chemistry lab. We find that the 10 heaviest envelopes with paper shreds, i.e. those where paper shreds have been added most likely to unduly increase payoff, report a mean coin toss result of 3.00 tails, compared to 2.41 tails for the rest (t-test: p = 0.062).18 We summarize these two indicative findings as: (Mis-)Reporting in the coin-tossing task is consistent with behavior in hidden truth-telling tasks. These findings are in line with growing and distinct evidence on the external validity of experimental lab measures of truth-telling in the literature (Cohn and Maréchal, 2018; Cohn et al., 2015; Dai et al., 2018; Gächter and Schulz, 2016; Potters and Stoop, 2016). Our indicative finding, together with the mounting evidence in the literature, therefore suggests that the coin-toss truth-telling measure is informative of fishermen's honesty behavior in the field. Next, we compare fishermen in the Baseline treatment with our student sample that faced the exactly same study design as the fishermen (see Fig. 5). As outlined in the previous section, we collected this data to see how our fishermen subject pool relates to the conventional student subject pool. While fishermen reported to have tossed 2.38 tails on average, students report 3.14 tails on average, which is significantly higher (two-sided t- test: p = 0.000).19 This result points to the importance of using our subject pool of fishermen to answer our research questions regarding fishery policy. Inspired by the quote of McAngus (2016) at the beginning of our paper, we were also interested to see how the reported trust in the European Union (specifically the European Commission), measured on a Likert scale from 1 to 9, compares between our fishermen and students. Our findings reflect the spirit of McAngus’ quote: the average trust in the Baseline treatment of the fishermen in the EU is 2.29 compared to 5.05 of the students. We find no significant differences in the reported trust levels among fishermen across the three treatments (pairwise t-test: p > 0.60 for all cases). Fig. 6 depicts the distribution of trust towards the EU for all fishermen versus the students. We see that almost half of the fishermen report a trust level of ‘1’, the lowest possible answer on the scale. The two distributions are highly different from each other (two-sided Kolmogorov–Smirnov: p = 0.000). From the distributions in Fig. 6, it is evident that there is heterogeneity of students’ trust in the EU (while the fishermen's trust is skewed). We took advantage of this observation to run an additional online experiment with students from the same university (as outlined in Section 2). This online experiment serves the purpose to investigate the impact of trust in the EU on attrition. As noted in the experimental design, mail field experiments may suffer from attrition. As we cannot rule out effects of our treatments on the extensive margin by design, i.e. that participants selected on responding not independently from the random treatment assignment, an alternative explanation for our findings would be that there is a fixed proportion of honest and dishonest fishermen and that the honest participants were less likely to send in the study when being confronted with the EU flag. To study whether the EU framing may induce attrition, we collected additional data from 717 student subjects in the online experiment. We randomly assigned the subjects into two treatments: positive EU framing and negative EU framing. In both treatments we provided subjects with true information regarding the EU fishing policy. The difference was that in the positive EU framing treatment the regulatory efforts of the EU fishery policy were framed as a success and in the negative EU framing treatment the regulatory efforts of the EU fishery policy were framed as a failure.20 To reinforce the strength of the framing, we subsequently asked subjects to report positive (negative) own experiences with EU regulations. Thereafter we asked subjects to report their trust (a) in the European Commission, (b) the federal German government and (c) the government of their German state.21 This exercise yields two key insights. First, we find that subjects in the positive EU framing treatment report to trust the EU significantly more than subjects in the negative EU framing treatment (two-sided Kolmogorov-Smirnov test: p = 0.017).22 Fig. 7 depicts these results. Thus, the framing seems to have worked and we can examine how it may impact attrition. Second, we found overall attrition of only 1.25% (from 717 to 708 participants). Crucially, there is no significant difference between the attrition in the two treatments (two-sided chi-squared test: p > 0.1). While this finding provides suggestive evidence that attrition does not depend on the sort of framing (positive or negative) regarding the EU, it does not preclude the possibility that attrition in our main experiment with fishermen has been treatment-specific.","To study the robustness of our results and possible mechanisms for regulator framing, we conducted a second online experiment. As it is difficult to get access to addresses of fishermen for studies and as we had already contacted all commercial fishermen in Germany for our main experiment described above, the aim was to find a group of individuals who, like fishermen in Germany, are known for not being fond of the EU as a regulator. We identified Brexiteers as such a candidate and accessible population, who may not have the same close regulatory experience with the EU as fishermen but are certainly know for not being fond of the EU. The platform Prolific (www.prolific.ac) offers the opportunity to run online studies with a fairly large and diverse UK population (N>20,000) and the platform pre-collected information on subjects’ self-reported Brexit votes (‘leave’, ‘remain’ or no vote). This gave us the opportunity to run a conceptual replication and extension of our initial study with up to 4064 potential subjects who had previously stated that they voted ‘leave’ in the Brexit referendum. The subsections below describe the experimental design and the results of our conceptual replication. Design ~~~~~~ In our conceptual replication with Brexit voters we used the same 4-coin-tossing task of Abeler et al. (2014) that we also used in our main experiment with fishermen. For each instance they reported that the winning toss “tails” laid on top, they received 1 GBP. Thus, each subject received between 0 and 4 GBP for this task, in addition to a 2 GBP participation fee. Following a few introductory questions, we used similar headers for the screen of the coin tossing task on the website in the new Brexiteers_Baseline and Brexiteers_EU_Flag treatments as in the Fisher_Baseline and Fisher_EU_Flag treatments with the fishermen, except that the headers were slightly larger in relation to increase the visibility on the computer screen. Based on inconclusive evidence regarding the Fisher_EU_Flag_Funding treatment in our main study, we included a third Brexiteers_Funding treatment, in which we displayed a text informing the participants that the study is funded by public research funds – yet the Brexiteers_Funding treatment does neither include any mentioning of the EU nor any EU flag. The Brexiteers_Funding treatment aims at disentangling the effects of providing information on research funding from the EU flag effect. Figure E.1 in Appendix E displays the three screen headers of the coin tossing task. We included further follow-up questions to investigate potential mechanisms that might drive treatment effects. First, our replication includes three variables on negative reciprocity adapted from Falk et al. (2018). We included these variables to reveal possible mechanisms for why individuals who dislike or mistrust the EU are reporting higher tail tosses in the EU_Flag treatment (and are therefore more likely to lie in this treatment). Second, as a potential confound that might drive treatment differences we investigate experimenter demand effects (Zizzo, 2010; De Quidt et al., 2018). To this end, we included the following question in as a follow-up: “How strongly do you feel that the researchers of this study wanted you to report in a particular way in the coin task?” and elicited responses on a 9-point Likert scale from 1 (not at all) to 9 (very strongly). Finally, even though Prolific allowed us to pre-screen for Brexit voters, we asked participants of our replication study whether they voted ‘leave’ in the referendum and whether they would still vote for some form of ‘leave’ (Hard Brexit or PM Theresa May's Leave Deal). We based the number of observations (N) in this conceptual replication with Brexit voters on a power calculation that is informed by the effect size of the EU_Flag versus Baseline treatment on coin toss results in our main experiment with fishermen. The power calculation informed us that it would be necessary to collect 173 observations per treatment in order to replicate the fishermen's effect.23 With three treatments, this would yield N = 519. Due to the differences between our fishermen experiment and this replication, we decided to increase N by spending up to our budget constraint of 8000 Euro, and to pre-specify hypotheses and use one-sided tests for those pre-specified hypotheses (we use two-sided test for analyses that we did not pre-specify). In particular, we expected that participants in such an online experiment participate in studies more often and may take our study less seriously than the commercial fishermen. For this reason, there are likely attention differences between our print-out mail experiment, for which fishermen had a long time to make decisions, and the online screen experiment for which participation times are usually very short and inattention may be an issue. Thus, in our pre-analysis plan we set the number of observations to be collected to N = 1200 complete responses and set this total number on the Prolific platform. The data collection started on March 6, 2019. Subjects’ average completion time was 3.7 min (222 s) and they received an average payment of 4.32 GBP, amounting to an hourly-equivalent payment of 70.14 GBP. For the collection of our data, we employed the platform Social Science Survey (www.soscisurvey.de). Results ~~~~~~~ Randomization, which was computerized through the Social Science Survey platform, yielded almost exactly 400 observations for each of our three treatments (399, 401 and 400, respectively). Kruskal–Wallis tests and chi-squared tests over all three treatments do not report significant differences for any of the collected variables.24 This evidence suggests that the randomization process worked well. We do not find significant differences between our three treatments for the reported number of tail tosses using all data (see Fig. 8). The average reported number of tail tosses is higher in Brexiteers_EU_Flag than in Brexiteers_Baseline, and is therefore qualitatively as expected and formulated in our pre-analysis plan. Yet, this difference is not significant (p = 0.267, one-sided t-test, as formulated in our pre-analysis plan). As stated in our pre- analysis plan, we anticipated that the attention and involvement of participants of our online experiment may be heterogeneous and that completion time can be a proxy for (in)attention or fast-clicking. As attention and involvement may be crucial for regulator framing to work, we stated in our pre-analysis plan that we will analyze the data for a sub-sample that suggests a relative high level of attention and involvement. We expected that the sample of Prolific may contain participants who just fill out items and finish studies as fast as they can. Likewise, there may be participants who leave the screen to do other things for some time. These participants may not have paid the required attention to the screens and added noise to our dataset. For this reason, we proceed to analyze the treatment effects for a sample that excludes the 10% fastest and the 10% slowest participants, which we denote as ‘inattention truncation’.25 Fig. 8 depicts the mean reported tail tosses between treatments both for the full sample and the inattention truncation sample. In the next step, we conduct the (one-sided) t-tests for effects across treatments. We find marginally significant evidence for a higher mean number of reported tail tosses in Brexiteers_EU_Flag compared to Brexiteers_Baseline (p = 0.093). We therefore find suggestive evidence that points into the same direction and is consistent with Result 1 in our main experiment with fishermen. Our hypothesis in the pre-analysis plan regarding the difference in the mean number of reported tail tosses between Brexiteers_Funding and Brexiteers_Baseline is that there is a lower mean number in Brexiteers_Funding. The one-sided t-test clearly rejects our hypothesis (p = 0.959). Rather, there is evidence pointing into the opposite direction (p = 0.083, for a two-sided t-test). As such, this does not support the tentative conclusion from the main fishery experiment that the funding information might curb over-reporting. A third hypothesis of ours in the pre-analysis plan was that the mean number of reported tail tosses is higher in Brexiteers_EU_Flag than in Brexiteers_Funding. A t-test rejects this hypothesis (p = 0.656). To examine the robustness of these results and to examine the potential mechanisms discussed above (trust in the EU, experimenter demand effects and negative reciprocity), we run a regression analysis of overall tail toss reporting including control variables. Table 3 reports the estimation coefficients. Specification (i) includes the treatment dummies and the three covariates (‘age’, ‘relative income’ and ‘number of moves’) that are pairwise significantly correlated with tail toss reporting over all three treatments. Specification (ii), the full model, also includes all other variables collected from the participants. The estimations show that the effect size in the multivariate regressions for the EU_Flag treatment is much smaller in the Brexiteers sample as compared to the fishermen. For the full model – including all variables collected from participants in our online experiment – we find additional marginally significant evidence pointing again into the direction that the average number of reported tail tosses is higher in Brexiteers_EU_Flag than in Brexiteers_Baseline and also higher in Brexiteers_Funding than in Brexiteers_Baseline.26 Overall, our replication using an online experiment with Brexiteers provides some supportive evidence for Result 1 from our mail experiment with fishermen. Our replication study did, however, not provide the expected explanation for the treatment effect of Brexiteers_Funding compared to Brexiteers_Baseline. We come back to this result in Section 5. The replication with Brexiteers also allows us to include additional questions before and after the coin tossing task and to examine possible mechanisms explaining mean coin toss reporting in our three treatments. Before the Brexiteers encountered the coin-tossing task, in all treatments they were asked for their trust in the EU. After the coin tossing task, they were further asked to state to what extent they felt they were pushed to report the coin tossing in a certain manner (to study experimenter demand effects) and, also after the coin tossing task, the Brexiteers were asked to answer questions on negative reciprocity (inspired by Falk et al., 2018). For the purpose of examining possible mechanism of our treatment effects, we dissect the mean number of tail tosses by reporting pairwise correlation matrices for the relevant five variables by treatment (see Tables E.2a–c in Appendix E). For Brexiteers_Baseline we do not find any significant pairwise correlations of the five variables with the reported number of tail tosses (Table E.2a). In our Brexiteers_EU_Flag treatment we find that one variable correlates (marginally) positively with the number of reported tail tosses: the answer to “Willingness to punish someone who treats me unfairly”. Hence, we get a weak hint that negative reciprocity might be a motivation for lying in Brexiteers_EU_Flag. Finally, in Brexiteers_Funding, we observe that all three negative reciprocity answers correlate (marginally) negatively with tail toss reporting. A preference for reciprocity in combination with a generally careful view on public funds might curb some lying. This exercise neither yields any correlation between the number of reported tail tosses and trust in the EU nor between tail tosses and perceived experimenter demand. Finally, we briefly address attrition in our online experiment with Brexiteers. Randomization took place when entering the treatment page and prolific only counted participants that finished the study and came back to prolific afterwards. Of the 66 incomplete responses that are included in the dataset, 29 dropped out immediately, 31 looked at the introduction page but did not proceed to the next page with preliminary questions, which only six individuals entered and four of them completed. Two of the four individuals proceeded to the treatment page. They were allocated to the EU_Flag and Funding treatments, respectively. These two individuals did not report a coin toss result and stopped the survey at this point. We therefore find that treatment induced attrition might be relevant for less than 0.2 percent of the sample, and based on the small sample we find no indication for it. This corroborates the findings from the additional student sample on the likely absence of treatment-specific attrition in the setting of online experiments. However, this does not imply that treatment-specific attrition is no concern in the main fishermen mail field experiment, as both the populations and response times differ substantially. Indeed, we cannot rule out the possibility that some treatment differences can be driven by treatment-specific attrition. The potential magnitude of such attrition biases would have to be tested in future mail field experiments.","This paper presents field experimental evidence on truth-telling of German commercial fishermen who are regulated by the European Union (EU). To our knowledge, this is the first artefactual field experiment with professional common-pool resource users on truth- telling.27 Examining truth-telling of German fishermen is of direct relevance, as the member states of the European Union stand to decide on how much costs to incur to monitor a recently enacted ban on discarding unwanted fish catches to the sea. The regulator thus currently depends on fishermen's honesty, while standard economic theory predicts substantial lying behavior. This paper not only studies fishermen's overall degree of dishonesty but extends the scope of previous studies by asking how regulator framing affects truth-telling—a dimension that is relevant for the effective and efficient design monitoring and sanctioning mechanisms. Our results are therefore not only relevant for the specific fishery context, but crucial for a broader understanding of truth-telling, the management of common pool resources around the world, and for regulatory policy more generally. To examine the more general applicability of our findings, we also conducted a conceptual replication and extension of our initial study with 1200 UK citizens who voted ‘leave’ in the Brexit referendum. Adapting an established coin-tossing game (Abeler et al., 2014), where subjects have to toss a coin 4 times and receive 5 € for each of the 0 to 4 reported tail tosses, we test whether truth-telling in a baseline setting differs from behavior in two treatments with different EU framings. The fishery is an ideal test case for studying how truth-telling behavior may be affected by regulatory framing, as there is almost uniform contempt among fishermen concerning stricter EU fishing regulation. We therefore hypothesized that if regulatory framing affects truth-telling, it would lower lying costs and thus result in higher misreporting among the treated fishermen. We find overall that fishermen misreport coin tosses to their advantage, albeit to a significantly lesser extent than standard theory would predict. Specifically, we find an average reported tail toss result of 2.46, while the expected truthful distribution would result in 2 and the payoff-maximizing choice in 4 reported tail tosses. Fishermen thus do not lie to their maximum advantage, but partial misreporting is prevalent among fishermen, in line with recent evidence by Abeler et al. (2019) and Gneezy et al. (2018). Qualitatively, we find the same to be true for Brexiteers. Crucially, we find that misreporting is larger among fishermen who are faced with the EU flag compared to the control sample without the EU flag on the instruction sheet. Furthermore, the replication with Brexiteers provides some supporting evidence for this main finding. This result confirms our hypothesis according to which many fishermen (and Brexiteers) have lower moral lying costs towards the EU, which they dislike. This indicates that previously elicited degrees of truth-telling may not be appropriate for principal-agent relationships, where the principal or regulator is ill-regarded by the economic agents. Although we have found consistent evidence across our main experiment and the conceptual replication, our central result is only borderline significant. The main study with fishermen may be underpowered, and the replication study with Brexiteers may have suffered from lack of attention. Future research should therefore be devoted to confirm the generalizability of this result. In particular it would be important and interesting to test whether the results also extend to other settings. The experiment with fishermen did not confirm our initial hypothesis 3 that fishermen report less truthfully in EU_Flag_Funding compared to the EU_Flag treatment. Our expectation was that fishermen may regard the additional informational cue as an indication that there is plenty of funding available to those conducting the study, and that this may reduce the moral cost of lying, such that fishermen would report less truthfully. This effect is not apparent in the experiment with fishermen. In the regression analysis, the EU_Flag_Funding treatment variable does not explain any difference in reported tail tosses compared to the baseline treatment. The conceptual replication with Brexiteers includes a treatment where we show the information that this study is funded by public funds, but without the EU flag, i.e. omitting reference to the regulator. This treatment leads to increased misreporting in a similar way as the EU_Flag treatment. Indeed, it might be that the additional information box about funding may make the wealth of the funding institution more salient and taking money may appear permissible to the Brexiteers. This would be in line with our initial hypothesis 3 formulated for the experiment with fishermen. However, the conceptual replication did not provide an explanation why fishermen were more honest in the EU_Flag_Funding treatment than in the EU_Flag treatment. One mechanism that may seem plausible is that fishermen may have considered the joint information on research funding and the EU flag as information that the EU is using money to survey fishermen. As this was a mail experiment, fishermen may even have taken the opportunity to discuss these issues with relatives or family members, which may have reinforced such a view. This may lead to more support for the regulator, rendering the responses statistically indistinguishable from the baseline treatment. However, this is just one possible explanation and there may be others. Our findings show that the ill-regard of the regulator is not the only effect present, and that other mechanisms may offset the dislike effect. It is an interesting question for future research if and how the regulator can approach regulated individuals in such a different kind of way to offset the negative attitude and thus increase truth- telling behavior. Moreover, we find evidence suggesting some consistency of behavior between the coin-tossing task and two other measures of truth-telling or lying behavior. This finding is based on two hidden tasks in the experiment—leaving the ownership of a coin to flip ambiguous and using an additional task in which it was possible to provide more material than was supplied—that may be of use for experimental methodology beyond our specific context to investigate the external validity of standard lying tasks. Overall, our findings imply that regulators not only have to consider some exogenous degree of dishonesty among the regulated, but also take into account that truth-telling may erode in reaction to the regulatory policy. Faced with a variable degree of dishonesty, the regulator can act strategically in adopting its regulatory approach, such as shifting part of the regulatory work to bodies that are closer to the regulated, thus considering how the regulated will adapt their behavior. Whereas the substantial number of fishermen who likely report honestly might suggest that softer monitoring approaches could be sufficient, the strategic aspect of regulatory experience calls for a more deliberate approach. One possible solution to coping with this strategic dimension of dishonesty would be to choose the ‘corner solution’ and comprehensive control.28 In practice, this would mean a monitoring scheme relying on-board observers or camera systems. However, instead of directly incurring the high costs to the regulator and fishermen of comprehensive control, our recommended approach would be to introduce monitoring of different degrees of stringency selectively to study the effects of monitoring on honesty. Overall, our findings imply that lying is more extensive towards an ill-regarded regulator and that policy needs to account for this endogenously eroding honesty base. Studying this new dimension of truth-telling in further detail is a promising avenue for future research."],["The UK change of government in 2010 provoked a large structural change in the English education landscape. Unexpectedly, the new government offered primary schools the chance to have ‘the freedom and the power to take control of their own destiny’ with better performing schools given a green light to fast track convert to become an academy school. In England, schools that become academies have more freedom over many ways in which they operate, including curriculum design, budgets, staffing issues and the shape of the academic year. However, the change to allow primary school academisation has been controversial. This paper reports estimates of the causal effect of academy enrolment on primary school pupils. While the international literature provides growing evidence on the effect of school autonomy in a variety of contexts, little is known about the effect of autonomy on primary schools (which are typically much smaller than secondary schools) and in contexts where the converting school is not deemed to be failing or disadvantaged. The key findings are that English primary schools did change their mode of operation after the exogenous policy change, utilising more autonomy and changing spending behaviour, but this did not lead to improved pupil performance. --------------------------------------------------------------------------------","Since 2010, the educational landscape in England has radically altered. By 2017, nearly two-thirds of secondary schools and over a fifth of primary schools are academies. Academy schools are granted considerable operational autonomy by government and have a battery of freedoms they can use that standard state schools cannot. As Michael Gove, the Minister then responsible for education, put it by enabling academisation these schools have been ‘given the freedom and the power to take control of their own destiny’.1 Although academies were present before then – principally as a school improvement policy for underperforming secondary schools since 2002 – the programme was radically altered and significantly expanded following the election of the new UK government in May 2010. It became a school structure to which all schools were invited to aspire as enabling legislation – the Academies Act of 2010 – was rapidly put in place two months after the election of the new government.2 For the first time, and through this completely unexpected policy change,3 primary schools were invited to become academies, with better performing schools being given priority to convert. The first batch of such schools converted in the school year beginning in September 2010. This paper reports estimates of the impact of primary school conversion to academy status on their operation and on the performance of enrolled pupils. This introduction of primary academies took place in an international context where publicly-funded autonomous schools have become a familiar form of school improvement policy, most notably through charter schools in the US and free schools in Sweden. Research on the former tends to find achievement gains associated with charter status and with the ‘injection’ of charter school features to public schools, particularly in urban settings where the schools typically enrol disadvantaged students.4 In the Swedish context, there is some evidence of positive short and long term effects of the free school program, but these are found to work primarily through competition (see Bolhmark and Lindahl, 2015). The policy studied here differs from most others in the literature in three important respects. Firstly, it involves conversion of existing schools rather than the creation of new schools.5 Secondly, it is about the voluntary conversion of better performing schools and not the forced conversion of failing schools. These better performing schools very clearly have a lower proportion of children from disadvantaged backgrounds. Thirdly, the focus is on young children (aged 7–11) who attend primary schools, which are much smaller than secondary schools.6 Although there have been studies of elementary schools in the charter school context, these are less prevalent than studies of middle and high schools. Similarly, studies of autonomy in the context of the English education system have focused on particular subsets of secondary schools; specifically, advantaged secondary schools voluntarily gaining greater autonomy (Clark, 2009), disadvantaged secondary schools (Eyles and Machin, 2015), and secondary schools in relatively disadvantaged local authorities (e.g. Birmingham in the case of Bertoni et al., 2017). Upon conversion, academy schools gain autonomy over many process and personnel decisions. This greater freedom may have positive effects on student outcomes because of superior information held by local decision makers (Hanushek and Woessmann, 2011). Indeed, the first secondary schools in England to become academies (in the early 2000s) did seem to deliver positive effects on student outcomes (Eyles and Machin, 2015; and Eyles et al., 2016a, 2016b). However, the context was one in which a couple of hundred (previously significantly underperforming) secondary schools became academies. It is not necessarily the case that these positive effects carry through to better performing schools and/or to (much smaller) primary schools. If the autonomy offered within the academies model was unambiguously advantageous for schools, one would imagine that all schools would want to become academies. However, recently the UK government has had to back out of a policy to force all schools in England to become academies by the end of 2022 because of fierce hostility to this by the educational establishment (although the current government vision is still to encourage all schools to become academies). Whether such radical upheaval is in the interests of students is an empirical question. Most schools yet to convert are primary schools, which represent the vast majority of schools in England. One might hypothesise that schools which volunteered to convert to academy status early-on are those that were most amenable to academy status, anticipating positive benefits. If effects are not found for such schools, one might question whether it is such a good idea to extend it to schools that are less enthusiastic. An important feature of the policy being studied here is that it was in no way anticipated by schools or parents. This gives leverage to identify causal effects since the conversion was exogenous to pupils already enrolled in the school. Thus, the sample studied is restricted to these “legacy enrolled” pupils who can be observed before and after academisation takes place. The importance of estimating effects for pupils who were already enrolled in the school prior to conversion emerges because student mobility post-conversion is potentially endogenous to the policy itself. For example, parents may be attracted by the idea of academy status and be more likely to enrol their children to newly converted primary schools. Exit from the school post- conversion might also be non-random (for example, if schools change policies in a way that is less attractive to certain students or their parents). However, in the empirical work discussed below, a very strong first stage estimate (of the effect of pre-conversion enrolment on the probability of attending an academy) suggests that a causal effect of academy attendance is identified for the majority of eligible pupils in the school. In practical terms, the empirical strategy adopted in this paper first involves selection of treatment and control groups of schools. The treatment group consists of primary schools that converted to academy status between 2010/11 and 2014/15. In each case, the control groups are those that converted in later academic years, but before 2016/17. Under certain conditions, these treatment and control schools are shown to have similar pre-trends in outcome variables. Further, enrolment in the primary school prior to conversion is used as an instrument for actual attendance in the academy in grade 6 when national tests in reading and maths take place. The legacy enrolment strategy mirrors that used in Eyles and Machin (2015) in their study of the first underperforming English secondary schools to become academies in the early 2000s. It also draws on Fryer (2014) who looks at the effect of injecting charter school practices into traditional public schools and Abdulkadiroglu et al. (2016), who study school takeovers in New Orleans, referring to pupils who stay in converting schools as ‘grand-fathered’ pupils. The rest of the paper is structured as follows. Section 2 describes primary education in England and offers a discussion of the institutional features characterising the introduction of academy schools. Section 3 describes the data and research strategy. Section 4 reports results from the first part of the empirical analysis, looking at whether primary schools that became academies did in fact change their modes of operation upon conversion. Section 5 reports the legacy enrolment results looking at causal effects of academy conversion on pupil performance. Conclusions are given in Section 6. Primary education ~~~~~~~~~~~~~~~~~ In England, children start school in the September after they reach the age of 4. Most children attend a primary school up to age 11, after which they go to secondary school.7 Schooling in England is organised into Key Stages. At the end of Key Stage 1 (age 7), pupils are assessed by their teachers in English and maths according to national guidelines. At the end of Key Stage 2 (age 11), they undertake national tests in English and maths.8 These tests are used to construct Performance Tables for primary schools, which are publicly available. There is next to no grade repetition within the system. Up until the introduction of academies in 2010, schooling had been organised at the local level into Local Education Authorities (LEAs). There are 152 LEAs in England and around 15,000 primary schools. The LEA's main functions in relation to primary schools are in building and maintaining schools, providing support services (e.g. for children with special needs), and acting in an advisory role to the head teacher regarding school performance and implementation of government initiatives. LEAs also have an important role in the funding allocations of schools. The bulk of schools funding comes from the dedicated schools grant which is given to LEAs and then distributed according to the LEA's own funding formula. The funding allocated to the LEA is based on a historically determined formula which is mainly driven by the numbers of pupils, ‘additional educational needs’ and local conditions. These local conditions include population sparsity, measures of deprivation, and wage costs in the area (Roberts and Bolton, 2017). As well as allocating funding, the LEA also appoints one or two representatives on to a school's governing body – a group of parents, teachers and community representatives that provides governance to the school. LEAs typically offer a number of administrative and management functions including training, personnel and financial services. Up until the 2010/11 school year, the majority of primary aged pupils (67%) attended community schools in which LEAs are the statutory employer of school staff, owner of the buildings and the authority that manages student admissions.9 Most other state primary schools are faith schools (which have greater autonomy from the LEA). Although parents can apply to send their child to any primary school (i.e. there are no strict catchment areas), popular schools are often oversubscribed and places are rationed according to a Schools Admissions Code.10 Academy schools ~~~~~~~~~~~~~~~ When a school becomes an academy, it is governed outside the LEA and is overseen and funded directly by central government. An academy school is run in many ways like a company, where governors are classed as trustees or directors and the principal/head teacher is the chief executive. Strong financial management and governance at the level of the individual academy are very important (National Audit Office, 2012), especially given that oversight is no longer provided by the LEA. Unlike Community Schools (i.e. most state primary schools), academies manage their own admissions. While they still have to adhere to the Schools Admissions Code, they may choose to run their admissions policy differently than in the past. Although academies are required to teach a broad and balanced curriculum, including English, maths, science and religious education, they are not legally required to use the national curriculum. They have the ability to set their own pay and conditions for staff and more freedom in their hiring decisions (e.g. they may hire unqualified teachers).11 Although academies are supposed to be funded on an equal basis with non-academies, they do get extra funds to cover the services that the LEA provides freely to other state maintained schools12; therefore, they have greater freedom on how to use the budget allocation. They also have the responsibility of organising payroll functions, insurance and accountancy functions in-house or by contracting this out. Academies also have the ability to change the length of the school day and the shape of the academic year (through term times). In the interests of minimising risk following the Academies Act of 2010, the Department of Education adopted a phased approach to the criteria for schools wishing to convert (National Audit Office, 2012), prioritising and giving a green light in the conversion process to better performing schools. A key component of this prioritising decision featured the rating in the reports of the Schools Inspectorate (Ofsted) that visits schools every 3–5 years and rates schools on a four point scale ranging from ‘outstanding’ to ‘unsatisfactory’. At the time, about 20% of schools were rated as ‘outstanding’ and 50% as ‘good’. The coalition government initially prioritised schools rated as outstanding and fast-tracked their applications for conversion. The first such schools were converted to academy status in September 2010. In November of the same year, this fast track route was extended to all good schools with outstanding features. At the same time, recognising the potential for economies of scale, academies were also encouraged to convert in chains or undertake some post-conversion collaborative arrangement with other schools. This option was made available for any school (irrespective of Ofsted grade) if it joined an academy trust with an outstanding school or an education partner with a strong record of improvement. In April 2011, the criteria was further widened to include schools that were ‘performing well’, which included consideration of the last three years' exam results, the latest Ofsted inspections, and financial management. As shown in Fig. 1, the initial take-up rate for primary academies in the first possible academic year (2010/11) was modest. This is unsurprising given the unexpected nature of the announcement, with legislation being rapidly passed by receiving royal assent in June 2010 and the fact that schools are likely to take time before making the decision to take on extra responsibilities (especially given the small size of primary schools in England). However, after that, there was a huge rise in the number of primary school academies in England between 2010/11 and 2016/17, with nearly a quarter of the sector being academy schools by 2016/17. Following the way in which numbers were constructed for Fig. 1, schools are said to convert in a given academic school year (September to August) if they are running as an academy by December of that academic year. Thus, for example, a school is classed as converting to academy status in 2014/15 if it converts at some point during the 2014 calendar year. The number of schools in the sample of converter academies studied in this paper, by year of academy conversion, is given in Table 1. There are a number of reasons for there being some discrepancy in numbers between Fig. 1 and Table 1. Firstly, because of the research design that is adopted, schools are only included in the sample if they have students enrolled in grades 2 and 6 in each academic year between 2006/07 and 2014/15. Secondly, the analysis focusses only on academies that voluntarily convert to academy status (around 30% of primary academies are sponsored academies that typically convert as a result of government intervention). Thirdly, schools that participated in the KS2 strike of 2009/10, and who therefore have missing outcome data in that year, are excluded. One further institutional detail of interest is that, in the post-May 2010 phase of academisation, schools have also been encouraged to convert in a chain or partnership. The Department for Education has stated ‘this can enable schools to support one another once they are academies, share resources, experience and ideas. Such an approach is particularly valuable to small primary schools where working together allows economies of scale to be achieved’ (Department for Education, 2013). The most prevalent model of collaboration is the multi- academy trust (MAT) wherein all schools within the MAT are governed by one trust and board of directors. MATs perform a role similar to that which would otherwise be played by the LEA in that they hire/fire teachers and are responsible for negotiating every aspect of teacher contracts - the disciplinary process and redundancy pay amongst other things - with the exception of pensions. MATs can also substitute for local educational authorities (LEAs) in that they top-slice funds allocated to schools under their trust and use this to supply central services previously provided by the LEA. In 2016/17, about 80% of primary academies were in a multi-academy trust. In the sample studied here, there are slightly fewer schools in MATS, with around 70% operating under this organisational structure. Data ~~~~ The National Pupil Database (NPD) is a census of all pupils in the state system in England. NPD includes basic demographic details of pupils – such as ethnicity, free school meal eligibility (FSM), gender, and whether or not English is their first language. The school attended by pupils can be linked to other school-level information such as the date of conversion to an academy school and the date and grade of Ofsted inspections (which are publicly available data). The data is longitudinal and tracks students as they progress through the state school system. As discussed in Section 2, the national curriculum in England is organised around Key Stages, the first two undertaken in primary school (in grades 1 to 6) and the second two in secondary school (in grades 7 to 11). Head teachers have a statutory duty to ensure that their teachers comply with all aspects of the Key Stage assessment and reporting arrangements. During primary school, this corresponds to Key Stage 1 and 2 which respectively cover grades 1–2 and 3–6. Local Authorities (and other recognised bodies) are responsible for moderation of schools. Thus, although teachers make their own assessments of students (and therefore could be susceptible to potential bias), there is a process in place to ensure that there is a meaningful assessment that is standardised over all of England. At the end of grade 2 in Key Stage 1, students are given a ‘level’ (i.e. there is no test score as such). However, following standard practice, National Curriculum levels achieved in Key Stage 1 assessments are transformed into point scores using Department for Education point scales and these scores are used in the empirical work reported on below.13 At the end of primary school in grade 6 (or the end of the Key Stage 2 phase of education), pupils take national tests in reading and maths, which are externally set and marked on a scale of 1–100. The final dataset used in this paper consists of multiple cross sections of grade 6 pupils linked to their school, demographic information, and test scores for the academic year 2006/07 to 2014/15. Test scores – both baseline KS1 and the outcome KS2 - are standardised, within the sample, at the grade/year/subject level. Other data sources are utilised in parts of the empirical analysis. The School Workforce Census is school level data that is available from the 2010/11 school year and provides a snapshot of each maintained school's workforce composition. Also studied is publicly available information on the income and expenditure of maintained schools and academies that is available from the 2009/10 school year. Finally, some results from a survey conducted by the Department for Education regarding the use of academy freedoms are presented (the source of these survey data is Cirin, 2014). This survey, which covers 25% of the 2919 academies that had opened by 1st May 2013, pertains to the freedoms exercised by schools once they gain academy status. Methodology ~~~~~~~~~~~ The main research question of interest is to identify the effect of academy conversion on pupil achievement in the grade 6 national Key Stage 2 tests taken by pupils at the end of primary school. Administrative data that follows pupils through their school careers is used to estimate the impact of academy enrolment on Key Stage 2 performance. In order to study this, a research design where instrumental variables are combined with difference- in-differences is implemented. In this design, outcomes of individuals in academies are compared with those who attend schools that later become academies, but do so after they sit their KS2 exams. In (1) i denotes pupil, s denotes the legacy enrolment school and t denotes school calendar year. Thus, αs is a legacy enrolment school fixed effect14 and αt is a time effect for the academic year in which the pupil is in grade 6. The vector X is a set of control variables, and the binary Academy variable takes value 1 if pupil i who was legacy enrolled in school s sits their end of primary school KS2 examination in an academy school. Finally, v1 is an error term. Despite already restricting to the legacy enrolment sample so as to avoid endogeneity concerns it may still be problematic to estimate Eq. (1) by ordinary least squares because legacy enrolled pupils may leave the school before the end of Key Stage 2. To allow for students to (potentially) sort into schools non-randomly as a result of the school obtaining academy status, an instrument for academy attendance is therefore used. This is whether or not the pupil was already enrolled in the school in the year prior to conversion in grades 2–5. Those for whom this variable takes a value of one are referred to as being intention-to-treat (ITT). Attention is focussed on those pre- enrolled in grades 2 to 5 as grade 5 is the penultimate year of primary education and grade 2 is when the KS1 assessment takes place, thus ensuring that KS1 assessment does not take place in an academy for the ITT pupils.15 It is important to note that pupils enrolling in the school after conversion are not included in the analysis. To ensure that the control group and treatment group are selected in the same way, a slightly different control group is used for each cohort of academy converters. For those converting in 2010/11 for example, the control group consists of pupils who are in grades 2–5 in 2009/10 at schools that convert between 2011/12 and 2016/17, but are not expected to sit their exams in an academy.16 Control groups are defined similarly for all schools converting up to and including 2014/15.17 Because of these restrictions, the event study on pupil performance has to be limited to a maximum of four years post-conversion, including the year of conversion itself. This is because there are 4 remaining years of primary school after the Key Stage 1 assessment. Thus, pupils affected by conversion in grade 2 of primary school (when KS1 assessments are taken) could have up to four post-conversion years of education in the academy. Similarly children affected by conversion when enrolled in the predecessor school in grade 3 could have up to three conversion years, and so on for children in grades 4 and 5 in the predecessor school. In the first stage, Eq. (2), estimates of θ2 show the proportion of the ITT group that stay in the academy and take their KS2 tests there. Eq. (3) is the reduced form regression of KS2 on the instrument. A two stage least squares estimate (2SLS) estimate can then be obtained as the ratio of the reduced form coefficient to the first stage coefficient, θ3/θ2. The main specifications that are estimated - Eqs. (2) and (3) - are based on pooled data for the five cohorts of academy conversions already described. Extending this to an event study framework enables separate estimates for the number of years a pupil is exposed to being in an academy post conversion (up to a maximum of four including conversion year) to be obtained. In this case, there are four instruments for whether a pupil is expected to sit their exams in the year of conversion (those in grade 5 in the year prior to conversion to instrument one year of exposure), the next year (those in grade 4 in the year prior to conversion to instrument two years of exposure) and so on up to the maximum of four years exposure (for those legacy enrolled in grade 2). It should be noted that, because of the data that we have, not all cohorts of converters contribute to the exposure estimates for later years. For instance, we can only identify the effect of four years exposure for those who are pre-enrolled in grade 2 in the first two cohorts of conversions. Comparison schools ~~~~~~~~~~~~~~~~~~ A naive comparison between primary academies and all other state-maintained schools is likely to suffer from significant selection bias, since (as discussed above) conversion to an academy was done on a voluntary basis and better-performing schools were prioritised and actively encouraged to convert.18 One might expect schools seeking to become academies to have common unobservable characteristics such as having a school ethos more in line with the academy model. To account for this, pupils attending future converters are used as a control group in a difference-in-differences setting. Thus the data structure that is utilised is a balanced panel of schools for the school years 2006/07 to 2014/15 with repeated cross-sections of grade 6 pupils. Balancing tests ~~~~~~~~~~~~~~~ This approach can be legitimised first through covariate balancing tests between treatment and controls in the baseline academic year (2006/07). Second, and probably more importantly, the empirical analysis shows there to be no evidence of differential pre- conversion trends in outcomes between pupils in treatment and control schools. On the former of these, Table 2 shows the extent to which treatment and control groups are balanced at baseline (2006/07) for the full sample of treatment and control schools, and separately for outstanding and non-outstanding schools. In terms of the full sample of schools, there is a significant difference with respect to KS2 scores prior to the policy, with treatment schools being better performing in maths. The workforce in treatment schools also appears to be both larger and, on average, younger and there are more pupils enrolled in the treatment schools. The above differences are not so surprising when it is acknowledged that the government prioritised better performing schools for conversion to academy status. For instance, within the sample of schools studied here, over 80% of the first cohort of conversions were deemed outstanding by Ofsted. This proportion declines monotonically to 11% for the 2016/17 cohort of conversions.19 For this reason it is necessary to look within Ofsted grades (as defined by the latest Ofsted grade awarded prior to 2010/11) when comparing treatment and control schools. When this is done, the schools look much more balanced on observables.20 In fact, as Table 2 shows, within categories of outstanding and non-outstanding schools, there are few statistically significant differences at baseline between treatment and control schools. In fact, for outstanding schools there are no statistically significant differences (at the 5% level) for any baseline characteristics including KS2 and KS1 scores.21 Thus, regressions are estimated for schools within each Ofsted grade, as well as for the pooled sample.22","Before looking at the effect of primary academies on pupil performance, evidence is presented on whether changes in the mode of operation occurred at primary schools that became academies prior to or during the 2014/15 academic year. Four aspects of this are considered. First, whether primary schools took up the option to exercise the many academy freedoms that became available from increased autonomy. Second, whether patterns of expenditure changed. Third, whether there were changes in workforce composition. Fourth, whether academies altered their pupil intake. Use of academy freedoms ~~~~~~~~~~~~~~~~~~~~~~~ There have been various investigations into whether schools actually use their academy freedoms upon conversion (e.g. Academies Commission, 2013; Cirin, 2014). The existing descriptive evidence confirms that they mostly do, but with some degree of variation. The Academies Commission (2013) conclude that take-up of freedoms had been ‘piecemeal rather than comprehensive’, in part because changes can take time to implement and sometimes require consultation. Surveys of recent converters by Bassett et al. (2012) and Cirin (2014) found financial motives to be important in the decision to convert. In the former study, over 75% of respondents cited it as one of their reasons for converting and two- fifths as their primary reason. Cirin (2014) found that the desire ‘to gain greater freedom to use funding as you see fit’ was the most commonly cited reason for conversion (cited by 83% of respondents). The vast majority (almost 9 in 10) also moved to procure services themselves. Importantly, Cirin (2014) breaks down results by primary and secondary status. This shows that the majority of academies do exercise freedoms, but this is more common in secondary than in primary schools. This is shown in Table 3, taken from his survey of 720 academies which were open on 1 May 2013. The numbers in the Table show that most schools report a use of academy freedoms, but that the percentage of primary schools making a particular change is smaller than it is for secondary schools. Furthermore, Cirin (2014) reports that almost all schools surveyed made at least one change (702 out of 720), implying that at least 95% of primary converters (262 primary converters were surveyed) exercised at least one freedom, with two-thirds believing that the changes improved attainment. Changes in expenditure patterns ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ Studies cited above on the use of academy freedoms suggest that the financial motive to convert was important. Table 4 shows numbers on income and expenditure before and after conversion in treatment and control schools using administrative data on school income and expenditure. Changes between the 2009/10 and the 2014/15 school years are reported. There are some data issues that need to be highlighted upfront before discussing these numbers. First, the timing of reporting changed after conversion, with academies reporting in the September–August school year as opposed to the April–March financial year.23 The latter is in line with local authority financial statements and was the practice in schools before they converted to become an academy and in control schools throughout the period of the analysis. Secondly, accounts for schools that do not convert in the period (the control schools) do not include the value of LEA provided services; however, information is available on how much extra income is given to academies to cover the value of these services (the Education Services Grant - ESG). To make the numbers comparable this is removed from both the grant income and expenditure for academies in column (2) of Table 4. Columns (1) and (4) of Table 4 show that per pupil income and expenditure was similar for the treatment and control schools before conversion. For example, as shown in Panel A, total income in all treatment and control schools was £3974 and £4156 per pupil respectively. Total expenditure was £3966 and £4154 per pupil in treatment and control schools respectively. As shown in Panels B to C of the Table, these pre-conversion numbers are also closely aligned for the comparisons undertaken within the outstanding and non- outstanding groups of schools. It is evident, however, that converting primary schools both received more money and spent more money post-conversion, even once the extra money given for LEA provided services is accounted for. The Table also shows the income and expenditure per pupil after conversion and a difference-in-difference estimate in the final column. This shows significant income and expenditure gaps arising after conversion relative to what happened in the control schools. The differences in total income and expenditure are estimated as £296 and £522 per pupil per year. The increases are clearly driven by the relative increase in grant income. A similar qualitative pattern is shown for schools classified as outstanding and non-outstanding, but with higher income and expenditure shown for the latter schools, most likely reflecting a higher proportion of disadvantaged students in this group. Table 5 shows the change in categories of expenditure per pupil before and after conversion.24 There are three Panels, which differ according to assumptions made about which services the academies procure post-conversion given that they are no longer provided for them by the local authority. The numbers in the upper Panel A are changes inclusive of the extra money delegated to them. The numbers in the middle Panel B subtract an equal share of the ESG money from each category of expenditure. Finally, those in the lower Panel C remove all of the extra ESG money from the expenditure on non-staff related running costs. In each case it is very clear that, even though primary academies spent more on teaching staff, non-teaching staff, and other running costs after conversion (relative to control schools), the increase was greater for administrative costs (i.e. non-teaching staff and other running costs). This is true for schools in all Ofsted categories. Because the amount of money earmarked for services previously provided by LEAs from expenditure is removed, these shifts cannot be attributed solely to the mechanical shift caused by the school having to take on more administrative tasks post-conversion. It seems that the primary academies studied in this paper did receive more income, but that they spent it disproportionately on day-to-day running operations rather than on ‘frontline services’ such as teaching staff. Changes in workforce ~~~~~~~~~~~~~~~~~~~~ Table 6 reports evidence on changes in the composition of the school workforce between 2010/11 and 2014/15 for schools that became academies in that period relative to schools that became academies in 2015/16 and 2016/17. Changes are shown for all schools and stratified by Ofsted rating. The Table reports difference-in-differences estimates for the total number of teachers employed, the pupil/teacher ratio, the mean teacher salary, the proportion of teachers who are in the leadership group or whether the school changes its head teacher. In general, the results reported in the Table show little evidence of workforce changes resulting from academisation. The one exception is head teacher turnover. For the full sample, there is a statistically significant 6.3 percentage point reduction in head teacher turnover in primaries that became academies. When broken down by Ofsted category, this occurs only in non-outstanding schools, which are 7.2 percentage points less likely to take on a new head teacher. This stands in direct contrast to the finding of Eyles and Machin (2015) who found that the vast majority of the first phase of academy conversions in the 2000s were characterised by new head teachers coming into academies and therefore that changes in managerial structure were a key feature of academy conversion that facilitated increased autonomy. This mechanism appears to be completely absent in the case of primary schools. Changes in intake ~~~~~~~~~~~~~~~~~ Alongside performance effects we look at whether pupil composition changed once a school gained academy status. As data is available prior to 2010/11, the analysis considers year- on-year changes between 2006/07 and 2014/15 in the characteristics of those entering the earliest grade in which the schools enrol pupils. We also include observations of pupils over this period that enter schools which become academies after the sample ends (in 2015/16 and 2016/17). The three outcomes considered are the fraction of the pupil intake who are eligible for free school meals, the fraction with English as a native language, and the total size of the entry year intake (in logs). In each case, school and year effects are included. The results, presented in Table 7, show no evidence that schools alter their intake along these dimensions.25","Taken together, the findings show that primary schools changed some aspects of their operations after becoming academies. In particular, most primary academies began to use freedoms made available to them as a consequence of conversion. They also received more income and altered how their expenditure was allocated across functions. With regard to the latter, the observed spending changes mainly affected administrative functioning and day-to-day operations, because of the removal of such provision from the local authority. At the same time, there was not much change in the school personnel or in composition of the pupil intake.","This section reports the results on pupil performance, starting with the main baseline set of results showing the causal impact of academy conversion on pupil performance. Then, in the light of the previous section's results showing that most, but not all, primary schools altered their modes of operation post-conversion, heterogeneous estimates along a number of dimensions are reported. Main results ~~~~~~~~~~~~ Table 8 shows estimates of the 2SLS specifications studying the impact of academisation on pupil performance in reading and maths tests at the end of primary school. Separate coefficients are shown for each subject, both for the pooled sample and by whether the predecessor school's Ofsted grade was outstanding or not. Columns (1) to (3) show estimates when the treatment is whether the school converts to academy status. Columns (4) to (6) show estimates for years of exposure. As the vast majority of legacy enrolled pupils stay in the school to take their KS2 exams - first stage estimates range from 0.92 to 0.95 - only 2SLS estimates are presented. In all cases, there is no evidence of any performance boost from academisation. The estimates are small in magnitude, sometimes negative, and almost all statistically insignificant. In terms of magnitude, the largest positive estimate is 0.02σ (with standard error 0.03) for reading in outstanding schools as reported in specification (2) of the Table. All of the other 2SLS estimates are lower than this, and nine of the twelve maths and reading estimates (including all six for maths) have negative signs. When considering the average of reading and maths scores it seems that primary age pupils did not benefit from attending an academy school in terms of their performance at the end of primary school.26 As results prove similar whether reading or math marks are used as the outcome of interest, only results based on average points are shown for the rest of the paper.27 One might be concerned about the research design being potentially contaminated by differential pre-policy trends.28 Fig. 2 therefore shows estimates from an event study, for the pooled sample, for pupils attending academies four years prior to academy conversion to three years after. The effects of being in an academy remain numerically small and insignificant (as the c to c + 3 coefficients all overlap with the zero line on the Figure). Moreover, there is no sign of pre-policy trends, nor any gradual improvement in results post-conversion. Table 9 also further generalises the Table 8 baseline results by reporting estimates for legacy enrolled pupils by discrete years of exposure, ranging from one to a maximum of four. Again, there is neither any sign of a positive effect nor any suggestion that benefits might be increasing with years of exposure. If anything, the opposite is the case, as the absolute values of the negative coefficients mostly get larger with more years of exposure. Heterogeneity ~~~~~~~~~~~~~ While there is no evidence of performance effects on average, both in the event study and years of exposure analysis, it may still be the case that academisation has scope to benefit some subsets of students and not others. It is also possible that certain school characteristics may be associated with differential academy effects on pupil performance. Table 10 therefore shows results from an investigation of whether the estimated 2SLS effect size differs in several ways: i) with whether the pupil is eligible for free school meals or not; ii) with an indicator for whether the school is in an urban area or not (given that the charter school literature finds positive effects to be concentrated amongst urban schools); iii) whether it differs with pre-conversion school size (as larger schools may be more adept at managing their extra freedoms); and iv) whether the school joins a multi-academy trust (MAT). The results reported in the Table do little to alter the prior analysis. First, there is little evidence that the effect of academy attendance differs depending on whether one is eligible for free school meals or attends an urban academy. Panel C of Table 10, shows that the same can be said for pupils attending schools of differing sizes. Although performance effects appear to decline with school size, none of these interactions reach statistical significance. The final aspect of heterogeneity considered – whether or not pupils attend an academy that becomes part of a (MAT) or not – does uncover some differences. The most noteworthy is that some of the estimates for not being in a MAT are significantly negative. This is the case for all schools where there is a 0.06σ (0.02) fall – closer investigation shows that this is confined to the non- outstanding schools. This is consistent with the hypothesis that conversion in stand-alone (non-MAT) schools, which are not able to benefit from the economies of scale that a MAT can bring, may have actually proven detrimental to pupils enrolled in previously non- outstanding academies. However, this result should be taken with caution. About 60% of the primary academies considered here are part of a MAT, but it should be acknowledged that whether or not a school is able to join a trust is endogenous to KS2 performance; results showing performance drops could be due to negative selection to the category of non-MAT academy schools (evident only for non-outstanding schools).","The English government has radically restructured its school system under the assumption that academisation delivers benefits to schools and students. This paper studies the unexpected policy change that occurred in 2010 that enabled (and encouraged) primary schools to become academies. It looks at the first primary schools that have become academies in England (between 2010/11 and 2014/15) and finds no evidence of pupil performance improvement resulting from conversion. How should an overall zero effect be interpreted in the light of evidence showing positive effects of autonomy in other contexts? One reason is that schools that converted were already doing well within the system and simply did not require additional autonomy in order to thrive and therefore did not make substantive changes. Indeed the limited changes that are seen – increasing expenditure on non-instructional tasks – do not correspond to the kinds of changes, such as effective discipline and higher quality teaching that have been found to increase test scores in other contexts such as charter schools (Fryer, 2014). In existing research, much of the positive effects of autonomous schools have been shown for disadvantaged students and not so much for advantaged students. While there was scope to improve achievement within these schools, it may be that changes introduced as a result of school autonomy simply do not benefit such students at the margin. However, given the survey evidence reported above and the research into how additional income was used by schools, it would appear that many of these schools did not make changes that affect ‘frontline services’ (as opposed to administrative roles). Another possible reason is that effects are estimated in the short run. It may be that the programme will bear fruit once more schools convert and facilitate greater economies of scale by entering into or deepening collaborative arrangements with each other. In the heterogeneity analysis, we found some evidence of variation by whether or not schools are in a multi-academy trust. Although we do not take the effect to be causal due to the endogenous decision to join a MAT, it is still a worrying finding that performance dips for the non-outstanding primary schools (around 40% of converters) that do not join a multi-academy trust. Finally, one of the key models for some successful urban charters in the US and some secondary schools in England29 – an effective discipline approach for academies and the No Excuses model of charters – is of less relevance to the age range of children enrolled in English primary schools than for secondary age children (since behavioural problems that may lead pupils to be suspended or excluded from school are much more prevalent in the latter).30 In the light of all these factors, it is not surprising that there has been no overall effect on pupil performance. One might argue that if academisation has no average effect on pupil performance, this could still be a reasonable public policy if there are other reasons for why this might be beneficial – for example, if school leaders can more easily make changes that might benefit students (or their parents) and staff. However, the process of restructuring individual schools has been shown to be financially costly and restructuring on a system wide basis would likely prove to be too costly in the long run if it fails to generate gains for students in terms of test scores. Furthermore, risks are also posed by an increasing number of schools becoming academies.31 For example, they are no longer regularly monitored at the local level. Problems might not therefore come to light unless they are flagged up by an Ofsted inspection, which are not regular events. There are potential negative spill-overs on other schools if opting out of Local Authority control undermines services that the Local Authority is able to provide to other schools in the same geographic area (e.g. child psychologists to support children with special needs in many schools). Studying the operational aspects of academies, and the institutional structures in which they function, is an important subject for future research."],["We study the opportunistic political budget cycle in the London Metropolitan Boroughs between 1902 and 1937 under two different suffrage regimes: taxpayer suffrage (1902-1914) and universal suffrage (1921-1937). We argue and find supporting evidence that the political budget cycle operates differently under the two types of suffrage. Taxpayer suffrage, where the right to vote and the obligation to pay local taxes are linked, encourages demands for retrenchment and the political budget cycle manifests itself in election year tax cuts and savings on administration costs. Universal suffrage, where all adult residents can vote irrespective of their taxpayer status, creates demands for productive public services and the political budget cycle manifests itself in election year hikes in capital spending and a reduction in current spending. © 2014 The Author. --------------------------------------------------------------------------------","Suffrage rules regulate who can vote and this, in turn, influences the interests served by elected politicians. While today we associate democracy with equal and universal suffrage, historically the power to elect or appoint representatives was the privilege of narrow elites. Suffrage rules focussed on specific characteristics of the individual such as ownership of property, payment of taxes, residency and gender. The logic behind linking the right to vote to property holdings or tax payments can be traced back to mediaeval Britain and reflected the belief that it restricted the franchise to individuals with a longer-term interest in the welfare of the community, akin to the shareholders of corporations. Economic models in the tradition of Meltzer and Richard (1981) predict a straightforward positive link between demands for public goods and redistribution and extension of the franchise. However, evidence from analyses of historical data show that the impacts were more complex than predicted by theory and were functions of the specific rules that determined who could vote.2 While progress has been made in understanding the public finance consequences of franchise extension, little is known about the influence these rules have on the incentive to manipulate tax and spending patterns prior to elections in the quest for votes. A well-established literature, drawing on evidence from modern democracies and surveyed by Paldam (1997), Alesina et al. (1997) and most recently by Drazen (2008), offers a strong argument for the existence of opportunistic political budget cycles in both national and local elections.3 The construction of cross-country datasets (of OECD countries and more recently of developing countries) and of rich datasets for local governments (municipalities or states) from the modern period has tended to draw attention to the experience of the late 20th and early 21st centuries at the expense of earlier periods. Consequently the focus has been on opportunistic political budget cycles operating under universal suffrage; quite how the cycle might manifest itself in polities with economic and social restrictions on who could vote has been completely overlooked. The purpose of this paper is to draw upon the historical experience of early 20th century London to study the nature of the political budget cycle under two different suffrage regimes: taxpayer suffrage, where the right to vote is linked to specific tax payments; and universal suffrage, where all adults can vote (with minor qualifications), irrespective of their economic status. While the identity of the “pivotal voter” differs systematically under the two suffrage rules, electorally-motivated politicians can be expected to be equally determined to manipulate fiscal policy before elections to win support from the pivotal voter. We, therefore, conjecture that an opportunistic political budget cycle will be present in both regimes but that its nature will vary systematically with the suffrage rules.4 The setting for our study is the London Metropolitan Boroughs (LMBs) before and after the First World War. The 28 LMBs were established in 1899 and had powers to levy local property taxes, to decide on the provision of local services (sewer connections, bathhouses, parks, libraries, dairies and milk shops, etc.) and to take out loans to finance capital expenses on the security of future property taxes. Within the statutory boundaries, the LMBs had significant fiscal autonomy and the elected representatives of the councils could decide on the level, composition and the timing of key fiscal variables. All councillors were elected every three years. The franchise before the First World War was based on property tax payment and restricted to men; we refer to it as taxpayer suffrage. The Representation of the People Act (sometimes referred to as the Fourth Reform Act) in 1918 eradicated the tax payment requirement at all levels of government (including for the LMBs) and introduced almost equal and universal suffrage.5 This quasi-natural experiment allows us to study the opportunistic political budget cycle under two different suffrage regimes.6 Besides adding new historical evidence to the debate on the opportunistic political budget cycle, our study contributes directly to two more specific strands of literature.7 Firstly, it significantly enhances our understanding of fiscal retrenchment and taxpayer democracy in Britain. Until 1918, voting rights in local elections linked representation to the prompt payment of the local property tax (known in Britain as the rate) such that only local taxpayers had the right to vote. This had intriguing implications for the relationship between the size of the electorate and local public finance. In particular when the balance of power shifted to small-scale, middle class taxpayer-voters, demands were made for retrenchment and economy rather than fiscal expansion, despite apparently large social returns on public investment in local public goods (Hennock, 1963, 1973; Wohl, 1983; Szreter, 1988, 1997). As documented by Aidt et al. (2010), this generated a negative relationship between spending on local public goods and the extension of the franchise.8 We add to this by studying how opportunistic political budget cycles operate in an environment with taxpayer-voters. This restricted franchise is compared to the regime of universal suffrage, where the pivotal voter often does not contribute much to the local tax base. Secondly, our study contributes to the fast expanding research on the conditional political budget cycle initiated by Persson and Tabellini (2003), Brender and Drazen (2005), Shi and Svensson (2006), Alt and Lassen (2006a,b) and Alt and Rose (2007) amongst others, and recently surveyed by de Haan and Klomp (2013). The general point here is that the size and nature of the political budget cycle are conditional on the political and economic environment. They depend, amongst other factors, on economic conditions (e.g., the level of income), the institutional framework (e.g., the level of corruption, the type of election or political system), and the monitoring framework (e.g., fiscal transparency and quality of the press). We add an important dimension to this conditionality by showing that the opportunistic political budget cycle is influenced by the details of the franchise. We find the following results. Under taxpayer suffrage (1902–1914), the opportunistic political budget cycle materializes as tax cuts and in reduced spending on administration in election years. Under universal suffrage (1921–1937), we find that expenditures in election years are shifted towards productive public goods (capital spending) and away from other types of (current) spending, with no effect on tax income. The LMBs operated under a balanced budget rule which limited their ability to deficit finance election year tax cuts or spending booms, yet we find evidence of smaller surpluses in election years under both suffrage regimes. We interpret these findings in the light of the different incentives that variations in the suffrage rules generate for politicians to engineer opportunistic cycles. Building on Lohmann (1998), Shi and Svensson (2006), and Aidt et al. (2010), we provide a formal rational choice model that illustrates the logic. Under a restricted taxpayer suffrage that explicitly disenfranchises non-taxpayers and enfranchises owners of property in the locality who can reside elsewhere, taxpayer-voters often demand retrenchment and economy. Politicians respond to this by cutting taxes and reducing spending on administration in election years, as we observe in the data. In contrast, under universal suffrage all adult residents hold the right to vote, including many poorer residents who contribute little in terms of property tax payments to the funding of spending. This generates demand for fiscal expansion. Politicians, therefore, aim to engineer additional electoral support by adjusting the portfolio of spending towards productive public services which benefit the pivotal voter and away from other spending without necessarily increasing taxes. The rest of the paper is organized as follows. In Section 2, we introduce the institutional setting of our study and the particularities of the suffrage rules governing elections to the councils of the LMBs before and after the First World War. In Section 3, we develop the theoretical foundation for our empirical investigation. To this end, we sketch a rational choice model and provide an online supplementary appendix with technical details. In Section 4, we present the data and discuss some stylized facts about local public finance in London between 1902 and 1937. In Section 5, we consider the evidence of an opportunistic political budget cycle. In Section 6, we lay out our empirical strategy. We present the main findings in Section 7 and in Section 8 we discuss alternative interpretations and robustness checks. The concluding remarks in Section 9 recapitulate our findings in the context of conditional political budget cycles.","The 28 LMBs were established by the London Government Act of 1899 and they took office in November 1900 (Robson, 1939, chapter 10; Young and Garside, 1982).9 LMBs were created from the largest of the existing Vestries and District Boards of Works and by combining smaller Vestries and Boards into bigger and fiscally more viable units.10 As with the Vestries and District Boards, the main responsibility of the boroughs was the provision of local urban amenities. This included construction and maintenance of local streets, refuse collection, provision of public lighting (by 1912, 15 LMBs were generating their own electricity for street lighting), sewers and drainage, burial grounds, libraries, parks, baths and washhouses, and the employment of health officers. They could also purchase land and build public sector housing (White, 2001). Other services, such as schools, infectious disease hospitals, policing and major roads and infrastructure projects fell outside their jurisdiction and were handled by a variety of city-wide authorities, but the bulk of spending on sanitation and health-related public services was undertaken by the boroughs.11 The responsibilities stayed constant over the period from 1901 to 1937 (in fact to the 1960s) and there were no substantial changes in fiscal federalism over the period.12 The main source of LMB revenue was receipts from the rate – the local property tax – which often contributed around 90% of total income. User charges for specific services were also important and some equalization funds were available, though poorer boroughs complained about the iniquity of the redistribution (Booth, 2009). From this base, the boroughs provided local public goods and financed the administrative cost of running the council. They also collected taxes on behalf of other local authorities (e.g., the School Board for London, London County Council, the Boards of Guardians, and the Metropolitan Police). Within these institutional and fiscal constraints, the elected councillors had freedom to allocate public monies as they saw fit and to raise the tax resources they deemed necessary to fund required expenditures. While they could borrow funds for the purpose of capital investment, they were not allowed to do so to finance current spending and they effectively operated under a balanced budget rule which, however, did not preclude surpluses. The LMBs were governed by a council, consisting of a mayor, aldermen and councillors.13 These were elected in competitive elections. Unlike prior to 1901, where elections for the Vestries took place each year for a third of the vestrymen, all the LMBs adopted an election cycle in which the entire council was elected every three years.14 The rules governing the electoral franchise for the LMBs between 1901 and 1918 were codified in the Local Government Act of 1894. Voters consisted of two groups of men (and a limited number of widows and spinsters): the Parochial Electors and the Parliamentary Electors were entitled to vote under the Parliamentary Reform Act of 1884 and the Registration Act of 1885 (Keith-Lucas, 1952, p. 233). Both groups were required to occupy a property in the borough for a sufficient time period (ranging from 6 to 12 months), but permanent residence in the borough was not necessary. Some boroughs, therefore, had a significant number of absentee voters. Most importantly, however, eligibility to vote for the council was linked directly to payment of the rate. Provided that the occupancy requirement was satisfied, the right to vote was conferred on occupiers of property worth at least £10 and had been subject to 12 months' rating with the rate paid in full. This implied that the right to vote was restricted to the taxpayers of the borough who had paid their dues on time and in full. This disenfranchised many poorer inhabitants. Since the fraction of the total stock of property rated in each borough varied (slum areas were sometimes not rated) as did the diligence of tax collection, the fraction of males aged 20 and above that could vote varied greatly. In 1909, for example, about 37% of adult males in Stepney and 78% of adult males in Battersea were eligible.15 The average extension of the franchise across the boroughs between 1902 and 1914 was about 60%. We refer to this as the taxpayer suffrage. Taxpayer suffrage was abolished by the Representation of the People Act of 1918 which established one standard franchise for all general and local elections in Great Britain. For men the requirement was six months' occupation of land or premises in the area (i.e., no tax payment requirement). The condition for women was six months' occupation of land or premises in the area or as the wife of a man so qualified, on account of premises in which they both resided, if she was 30 years old (Keith-Lucas, 1952, p. 235). The Act also abolished the disenfranchisement of paupers for all local government purposes. While owners of land or buildings within the borough previously were entitled to vote whether they lived in the borough or not, after 1918 they qualified to be elected as a borough councillor, but not to vote. Although some women had to wait until 1928 to get the right to vote, we refer to this post-1918 situation as universal suffrage.","We consider a borough populated by capitalists (C) and workers (L) during two periods, t = 1,2. Each capitalist is endowed with capital (k) and two units of housing. Workers are endowed with one unit of labour, which is supplied in-elastically to a competitive labour market, and nothing else. There are nc capitalists and more workers than that. Each period, the capitalists combine their capital endowment with hired labour to produce output using a CRTS technology. The market clearing wage and profit income, wt⁎ and πt⁎, are both strictly increasing in total factor productivity. The capitalists “consume” one unit of housing privately and pay the property tax levied on it directly. The other unit is supplied to a competitive market as rental accommodation for workers. Under the assumption that the supply of houses is fixed, the incidence of the property tax levied on rented accommodation, if any, falls on the capitalists and workers therefore do not pay the local property tax. The capitalist-politician generates a rational political budget cycle. The capitalist-politician wants more rents and more spending on the non-productive public good than do capitalist-voters. Under taxpayer suffrage, the capitalist-politician cuts spending on the non-productive public good and rents to convince capitalist-voters of his quality. Since all capitalists agree on the optimal level of the productive public good, there is no pre-election distortion in that item. The combined consequence is that the tax rate falls. In short, the rational political budget cycle manifests itself as pre- election cuts in rents, less spending on non-productive public goods and lower taxes. We call this the retrenchment hypothesis. Under universal suffrage, worker-voters want more spending on the productive public good than the capitalist-politician. They are not concerned with the other budget items because the incidence of the property tax is passed on and because they do not benefit from non-productive public goods. Consequently, the capitalist-politician delivers the utility target UUS by spending more on the productive public good. The maximum rent is then extracted and spending on the non-productive public good is cut. The reason for the latter is that the increase in spending on g decreases net profit income, making it optimal to reduce spending on q. The net effect on the tax rate is ambiguous. In short, under universal suffrage the rational political budget cycle manifests itself as a pre-election hike in spending on productive public goods and a cut in non-productive services, with an uncertain effect on taxes. We call this the expenditure switching hypothesis. The LMBs could not run deficits but surpluses could be accumulated for precautionary reasons.18 We can capture this by assuming that the incumbent politician has a surplus target in non-election years but may deviate from this in election years (at a cost). Under taxpayer suffrage, the benefit is that more rents can be retained. Under universal suffrage, some of the increase in spending on productive public goods can be financed by suspending the surplus target. Surpluses may, therefore, be lower in election than in non-election years irrespective of the suffrage rules. This is the third hypothesis we test.","Table 1 lists the 28 London Metropolitan Boroughs and the ID number used to identify each of them in the maps shown below. Information on the LMBs' accounts is published in the Local Taxation Returns (1901–1914) and in the Local Government Financial Statistics (1920–1938). These sources contain detailed information on income, expenditures (current and capital) and debt for each borough. The format of the accounts, however, changed significantly after the First World War, when the responsibility for collecting and reporting local government public finance data moved from the Local Government Board to the Ministry of Health. After this change, a greater emphasis was put on recording information related to public health. This makes it impossible to match disaggregated budget items between the two sources and we consider two separate samples, corresponding to the two suffrage regimes. We stress, however, that a careful reading of the notes to the accounts gives us no reason to believe that there were any substantial alterations to LMB accounting practises that could account for systematic differences in the nature of the political budget cycle before and after the change in suffrage rules. The fiscal year runs from April 1 to March 31 throughout and we use the convention to refer to a fiscal year by the calendar year in which it ends. The taxpayer suffrage sample runs from 1902 to 1914. We cannot use the data for 1901 because the accounts only refer to a part of the year (November 1900 to March 1901) and 1914 is the last fiscal year available since systematic reporting was suspended during much of the War. The first accounts after the War for the fiscal year 1920 were incomplete and are excluded from the analysis. The universal suffrage sample, therefore, starts with the fiscal year ending in 1921 and runs to 1937. This gives a total of 364 observations for the taxpayer suffrage sample and 476 for the universal suffrage sample.19 The fiscal data is converted into real values using the Sauerbeck-Statisk price index from Mitchell (1988) with base year 1871 and expressed in per 1000 capita terms. The seven particular fiscal outcomes that we study are listed and defined in Table 2. Elections took place every three years: 1900, 1903, 1906, 1909 and 1912 before the War; and 1919, 1922, 1925, 1928, 1931, 1934, and 1937 after. The potential manipulation of the budget would occur before the election and would therefore fall in the fiscal year spanning the November election. We define the dummy variable election as being equal to one if fiscal year t is an election year and zero otherwise. Table 3 reports descriptive statistics separately for the two samples and Figs. 1–6 show the average trends for the fiscal outcome variables for the fiscal years 1902–14 and 1921–37, respectively. We notice a number of important facts. First, both current income and current expenditure increase in real terms from around £15 per 1000 capita (in 1871 prices) under taxpayer suffrage to £25 per 1000 capita under universal suffrage (see Table 3). The increase in capital expenditure (and capital income) is less pronounced. Secondly, there were no particular trends in current expenditure or in spending on administration under taxpayer suffrage (Fig. 1). Likewise, current income and rate income are stable in this period (Fig. 2). A similar characterization applies to the trends under universal suffrage (Figs. 4 and 5) and we note that, on average, the LMBs' spending and taxation levels were comparable in 1914 and 1921, despite the interruption of the War and the franchise change. We do, however, observe a decline in capital expenditure (and capital income) under taxpayer suffrage in the years before the War (Figs. 1 and 2). The spike in capital expenditure in 1905 is entirely attributed to a large investment in electricity in St. Marylebone and is (more than) matched by a large increase in capital income (a big loan). Thirdly, around 1930, a marked level shift upwards in current expenditure and in rate income but not in capital expenditure, takes place. A disaggregated analysis of the data [not reported] suggests that this reflects increases in spending on streets as well as increases in wage costs. Gillespie (1989) documents how some boroughs in the 1920s used resources for public relief work and we conjecture that this endeavour was intensified during the recession years. Fourthly, we observe substantial year-on-year variation in the average current deficit (Figs. 3 and 6). Mostly the LMBs were close to balancing the books and, on average, they ran a small surplus both before and after the change in the franchise (see Table 3). This suggests that the balanced budget rule mattered, but, at the same time, allowed some flexibility for fiscal manipulations. The average trends hide substantial cross sectional variation: some boroughs spent, taxed and borrowed much more than others. The dispersion is particularly large with regard to capital expenditures (and income) where the standard deviation is about twice as large as the mean values (see Table 3). We visualize this dispersion in Maps 1–3. Each map consists of two panels, one for the pre-war and for the post-war period, and colour codes the spatial distribution of rate income (Map 1), current expenditure (Map 2) and capital expenditure (Map 3). There is a consistent spatial pattern of high-tax–high-current-spending in north-west London, including Westminster, Holborn, St. Marylebone, and Hampstead, and Woolwich in the south- east. These are also the areas with high levels of capital expenditure under taxpayer suffrage. After the change to universal suffrage in 1918 there is a marked shift in capital expenditure to the east and south-east of London, with Poplar, Bermondsey and Greenwich standing out as big spenders. Much political debate was generated after the First World War about the high-rating and -spending policies of east end Labour councils such as Poplar. Leaders of the Labour Party in London were worried that this approach would alienate potential middle-class support in other parts of the capital (Gillespie, 1989). We also collect demographic data – total population, population growth (absolute change in number of inhabitants), population density (inhabitants per house) and age structure (proportion of the population below 20) – from the decennial Censuses.20 We do not have income or GDP data for the boroughs, but we record the average value of properties subject to taxation in each borough each year and use the variable wealth (defined as taxable value per 1000 houses) to proxy for income or wealth effects.21 We record information on the stock of outstanding loans at the end of each fiscal year and use the variable debt (defined as outstanding real debt per capita), as a proxy for accumulated spending on public services. Finally, for the taxpayer suffrage sample, we have collected information on the number of registered voters in each borough. We normalize this with the size of the adult male population to get the variable franchise extension which we use to control for variations in the size of the electorate. We have collected a number of additional variables used for robustness checks. We introduce these in Section 8.","The fiscal outcome variables defined in Table 2 are selected to facilitate tests of the retrenchment and expenditure switching hypotheses. Retrenchment effects would primarily show up as election year tax cuts. This is captured by rate income and current income where the latter, in addition to property tax revenue, includes income from user charges for local public services, but excludes revenues raised on behalf of other local authorities. Expenditure switching involves increasing spending on productive public goods that benefits all and cuts in non-productive spending. We presume that the outputs generated by capital expenditures represent productive public spending. In contrast, many current spending items are non-productive. Therefore, we use the variables capital expenditure and current expenditure to test the expenditure switching hypothesis. In addition, we use the variable capital income to test if there is a tendency to take out loans in election years. If this is the case, the need to increase the yield from property taxes to fund the pre-election spending hike in capital spending anticipated under universal suffrage would be reduced and we might expect to see a fall in tax income to match the expected fall in current spending. We use expenditure on administration as a proxy for bureaucratic spending with the rationale that the taxpayer-voter might find such outlays particularly wasteful.22 Finally, we use current deficit, defined as total current expenditure minus current income, to test election cycles in the fiscal balance. Before we turn to the formal statistical analysis, we present some descriptive evidence on the nature of the opportunistic political budget cycle in London between 1902 and 1937. Figs. 7–10 show plots of the seven fiscal outcome variables in “event time”. That is, each figure shows the average of the relevant fiscal outcome in election years, one year before an election and one year after an election. A “V” or an inverted “V” shape indicates a political budget cycle. We observe a clear revenue pattern under taxpayer suffrage: lower current income and lower rate income in election years than in other years (Fig. 7). We note a fall in spending for both administration and current expenditure (Fig. 8). The pattern is noticeably different under universal suffrage (Figs. 9 and 10). The budget cycle in rate income has gone. Instead, we observe a clear election year increase in capital expenditure with a hint of a cycle in current income. Under both franchise regimes, surpluses are lower in election years. Altogether, it appears that the political budget cycle differed before and after the expansion of the franchise in ways that are consistent with the retrenchment and expenditure switching hypotheses. It is clear, of course, that many other factors than the differences in the suffrage rules could be behind this, including the political ideology of the councils' governing parties and macro- economic trends such as the Great Depression. We consider these and other potential influences in a later section. First, however, we turn to a systematic analysis of the data.","In an attempt to balance various econometric issues with the data at hand, our model uses two different estimators. The first is a fixed effect estimator. We cluster the standard errors at the borough level to take into account the fact that autocorrelation in a fixed effect model may inflate the z-statistics and cause invalid inference (Bertrand et al., 2004). The lagged dependent variable may, however, cause a Nickell bias (Nickell, 1981), since our two samples have only 12 and 16 years of observations, respectively.24 Our second estimator takes this into account. The GMM estimator (Blundell and Bond, 1998) used by, e.g., Shi and Svensson (2006), is not ideal in our case as it requires many more cross sectional units to yield consistent estimates than we obtained. For this reason, we use the bias-corrected least-squares dummy variable (LSDV) estimator. It performs better than the GMM estimator in panels with a small cross section (Bruno, 2005a,b).","The main results for the taxpayer suffrage sample are recorded in Table 4 (revenue outcomes) and Table 5 (expenditure outcomes). The corresponding results for the universal suffrage sample are reported in Tables 6 and 7. For each fiscal outcome, we report both the estimates obtained with the fixed effects and the LSDV estimator, which mostly yield similar results for the election year indicator variable. We find strong evidence of an opportunistic political budget cycle in both samples but the nature of the cycle is conditional on the suffrage rules. Under taxpayer suffrage, the political budget cycle shows up as a reduction in current income and in rate income in the election year (Table 4). Spending on administration is cut in election years. There is no detectable impact on the relative composition of capital and current spending (Table 5). The cut in administration is insufficient to balance the books and we find that election years are associated with lower surpluses (Table 4). In other words, the election year tax cut is partly funded by cutting back on bureaucracy and partly by running a smaller surplus (or a small “unplanned” deficit). The magnitude of the tax cut is about £0.6 per 1000 capita which should be compared to the total income (net of precept) of £15 per 1000 capita. The reduction in administration corresponds to about a 1.5 per cent cut in the election year. These are sizable effects which are consistent with the retrenchment hypothesis. We observe a different pattern under universal suffrage. Most notable are election year increases in capital expenditure and reductions in current expenditures (Table 7). The increase in capital expenditure is £0.93 per 1000 capita with the average capital expenditure being about £5.2 per 1000 capita. The reduction in current expenditure is somewhat smaller (£0.55 per 1000 capita with average expenditure being £25). This suggests that the LMBs systematically moved large-scale capital projects to the election year, while cutting back on current spending. On the revenue side, we find no evidence of a political budget cycle in tax income or in capital income. There was an election year drop, however, in current income (Table 6) and a tendency to run smaller surpluses or larger deficits in election years. Since rate income is unaffected, the fall in current income can be attributed to election year reductions in user chargers. These findings are consistent with the expenditure switching hypothesis. The estimations yield some additional results which are of independent interest. Firstly, the variable wealth is positively related to current spending and revenues in both samples. This is consistent with Wagner's Law that relates the size of government to income and wealth (Wagner, 1883). Secondly, insofar as the variable population captures scale effects, we notice that the negative point estimate on this variable in the estimations with current (and sometimes also with capital) expenditure is consistent with decreasing returns to scale in the production of these services. Millward and Sheard (1995), in their study of the local finances of 25 provincial municipalities in England and Wales from 1870 to 1914, also find evidence of (moderate) diminishing returns. Population growth correlates negatively with revenue and expenditure outcomes, but is only significant in the universal suffrage sample. In the taxpayer suffrage sample, we control for the size of the electorate with the variable franchise extension. The boroughs which experienced an extension of the suffrage (due, for example, to changes in the fraction of property rated for tax purposes) tended to collect more tax income and to run smaller surpluses (Table 4).","In this section, we discuss a number of robustness checks and evaluate some alternative interpretations of our findings. Partisan cycles ~~~~~~~~~~~~~~~ The opportunistic political budget cycles that we have emphasised above do not make a distinction between the ideologies of the political parties in power but simply assume that all parties are primarily interested in getting re-elected. There exists, however, a well-established literature, beginning with the classical work by Hibbs (1977), Chappell and Keech (1986), and Alesina (1987), which takes the view that partisan cycles in economic and fiscal outcomes can emerge because parties have different views on appropriate policies and their hold on power fluctuates. In the context of budget cycles, the relevant distinction is between parties on the left which support higher spending and, with a balanced budget rule, higher taxes; and parties on the right which support lower spending and taxation. Both before and after the First World War, elections in London were fought along partisan lines (White, 2001). Before the War, the Progressive Party and the Moderate Party were the two dominant parties in local elections in London. At the time, the Progressive Party consisted of a mixture of Liberals, Fabian socialists and radicals and it generally favoured high (local) government spending. The Moderate Party and, from 1906, The Municipal Reform Party, created by Conservatives and Unionists, were the dominant right-wing parties. The political landscape changed after the First World War with the increased prominence of the Labour Party. This undoubtedly had its source in the Representation of the People Act of 1918 which enfranchised the working-class and gave it a strong voter base. Consequently, the Labour Party replaced the Progressive Party as the dominant left-wing party. The right-wing opposition formed a secret anti-Labour pact in 1922 in response to the popularity of the Labour Party in the first election after the War. To investigate whether partisan cycles were important, we code the dummy variable left for years in which a notionally left-wing party (the Moderate Party, the Labour Party or, in one case, the Socialist Party) holds the majority in the borough council and zero otherwise.25 We conjecture that, if anything, spending and taxation levels should be higher during the term of a left-wing party. The results are reported in Panel A of Tables 8 and 9 for the taxpayer suffrage and universal suffrage sample, respectively. The variable left is not significant except for one of the fiscal outcomes: for the taxpayer suffrage sample, left has a positive effect on current expenditures (at the 10 per cent level of significance). The evidence for partisan cycles is clearly weak, justifying our focus on opportunistic cycles. Importantly, controlling for ideology does not affect the evidence for opportunistic cycles at all: the rise of the Labour Party after the First World War cannot by itself explain the observed difference in the opportunistic political budget cycle. Absentee owners ~~~~~~~~~~~~~~~ Under taxpayer suffrage, owners of rated property in a borough were eligible to vote even if they did not reside in that borough. These absentee voters did not enjoy the benefits of better local public services to the same extent as resident voters. Accordingly, they might have been particularly inclined to support retrenchment and economy. It is, therefore, possible that variations in the fraction of absentee owners could by itself affect fiscal outcomes under taxpayer suffrage. Unfortunately, we do not have information on the number of absentee owners/voters, so we use data from the Censuses of 1901, 1911 and 1921 on the number of uninhabited houses as a proxy. The number of uninhabited houses in the boroughs ranged from 199 to 3283. We have no way of testing how strong the correlation between empty property and absentee owners is. Nonetheless we see from Panel B of Table 8 that the variable absentee owners is insignificant for all fiscal outcomes and that evidence of the opportunistic political budget cycle is as before. The Great Depression ~~~~~~~~~~~~~~~~~~~~ The economic climate in the 1930s was very different to that in the first decade of the century and in the 1920s. It is possible, therefore, that the Great Depression, and not the change from taxpayer to universal suffrage, could explain the differences in the nature of the budget cycle that we observe between the two sample periods. It should be mentioned here that London and the South East fared comparatively well during the Great Depression (White, 2001). There were pockets of extremely high unemployment in the capital but the overall unemployment rate in London was low (Marriott, 1991). As Booth and Glynn (1975) observed, when national insured unemployment peaked in 1932 at 22.1%, it was 13.5% in London, 28.5% in the North East, and 36.5% in Wales. Similarly, Hatton (2003) noted that across the period 1923–1938, average regional rates of unemployment varied from 8% in London and the South East to around 22% in parts of Wales and Northern Ireland. It is therefore unsurprising that White has written that the “1920s and 1930s consolidated the rise in the standard of life of the London working class that the First World War had so unexpectedly fan-fared” (White, 2001, p. 226). We recall from Figs. 4 and 5 that current spending and tax income shift upwards around 1930, suggesting that the depression years triggered a fiscal expansion amongst the LMBs. To investigate if the Great Depression and the associated jump in spending and taxation contributed to shaping the political budget cycle during the interwar years, we have re-estimated the model for the universal suffrage sample with the inclusion of a dummy variable, Great Depression, coded one for depression years from 1929 to 1937 and zero otherwise. The results are reported in Panel B of Table 9. The dummy variable is positive and significant, as expected, for current spending, current income and tax income. The depression years were, on average, associated with larger surpluses. More importantly, the evidence for the political budget cycle from Tables 6 and 7 is robust after controlling for Great Depression. We conclude from this that the Great Depression did exert some influence on the public finances of the LMBs but was not itself responsible for the difference in the nature of the political budget cycle before and after the extension of the franchise. Heterogeneous election year effects ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ The baseline results concern the average election year effect across the 28 LMBs in the two samples. This may mask important heterogeneity. To investigate this, we have re- estimated the baseline specification with a set of 28 borough-specific election year dummy variables. The results are summarized in Tables 10 and 11 which report the coefficient on the election year dummy for each borough for the two samples. We observe some heterogeneity as one would expect, but there is no indication that the average results are driven by one or two outliers. Table 10, with the results from the taxpayer suffrage, is sorted according to the size of the electorate (franchise extension). While the point estimates on the vast majority of borough-specific election year effects in the current income and tax income regressions are negative and significant, there are a few boroughs where the effect is positive. These are concentrated in boroughs with the most restricted franchise. This is consistent with the finding in Aidt et al. (2010) that retrenchment is most pronounced where penny-conscious middle class voters gain control of the councils. Other robustness checks ~~~~~~~~~~~~~~~~~~~~~~~ We have checked whether the null result for capital expenditure and capital income in the taxpayer suffrage sample can be attributed to the large investment recorded for St. Marylebone in 1905. Excluding this borough from the sample does not affect any of the results [not reported].","The evidence base for the existence of political budget cycles is overwhelming. Politicians use the fiscal levers granted to them to win re-election if they can. How this plays out is, unsurprisingly, a function of the institutional constraints imposed on the elected representatives. As pointed out by Brender and Drazen (2005), Shi and Svensson (2006) and many others, the political budget cycle is conditional. We contribute to the literature on the conditional political budget cycle in two main ways. Firstly, the focus of previous research has been on the period after the Second World War. In contrast, we enlist data from the early part of the 20th century and find that the political budget cycle is by no means a recent phenomenon: it was alive and kicking in London both in the years leading up to the First World War and during the interwar period. Secondly, precisely because of the emphasis on modern data, previous research explored the political budget cycle in the context of universal suffrage.26 Our historical perspective allows us to investigate the nature of the cycle under two different suffrage regimes and we find that it differs in marked but predictable ways. This strengthens the existing evidence base that the political budget cycle is contingent on political institutions and rules. Persson and Tabellini (2003), Streb et al. (2009), and Klomp and de Haan (2012) have previously demonstrated that election rules, regime types and legislative checks and balances affect the nature of political budget cycles in modern democracies. Gonzalez (2002) finds that the political budget cycle in Mexico was magnified as democratic institutions improve in quality while Potrafke (2012) reports that the cycle is stronger under two- rather than under multi-party systems. Others have found that the experience of voters, information flows and fiscal transparency are also important.27 The picture that emerges from this literature has yet to come into sharp focus, but one lesson is clear: context is crucial for the incentive and ability of incumbent politicians to manipulate expenditure and taxes for electoral gain. We have added new evidence to the understanding of the conditional nature of the political budget cycle by demonstrating that the suffrage rules themselves matter."]]