Orply.

Broad Transfers Work Best as Emergency Insurance, Not Stimulus

Valerie RameyHoover InstitutionWednesday, August 19, 20269 min read

Valerie Ramey argues that fiscal stimulus should be judged by the economic gain it produces relative to the debt it adds, not by the amount government spends. She says broad COVID-era payments were defensible as emergency income insurance in spring 2020, when targeted aid was impractical, but later rounds were harder to justify as the economy reopened. For short-run stabilization, she places greater weight on government purchases that can occur quickly and cautions that tax changes and deficit reduction carry their own, often larger, economic effects.

Fiscal action has to clear a higher bar than simply raising spending

Valerie Ramey treats fiscal stimulus as a question of what economic gain a policy produces relative to the debt it leaves behind. The answer depends on the instrument, the timing, how it is financed, and what would have happened without it.

That distinction separates emergency income support from ordinary demand management. Broad transfers may be defensible when a sudden disruption makes targeted aid impossible. But when households are able to save rather than spend, when purchases are constrained for reasons other than income, or when a recovery is already underway, the case for additional deficit-financed payments becomes harder to make.

Broad transfers were defensible when COVID made targeted insurance impossible

Ramey distinguishes pandemic spending with direct health purposes—such as vaccination and testing—from payments intended to stimulate demand. The central question for stimulus policy concerns the large checks sent to households.

She says she has not conducted the historical counterfactual needed to establish the full effect of COVID-era payments. Still, she suspects they stimulated spending even less than ordinary rebate programs. During the early pandemic, households could not freely spend on many services and activities. Rather than immediately translating into consumption, much of the support was saved; household balance sheets rose.

The first broad round of payments in spring 2020 nevertheless had a rationale, in her account. Many people had been ordered to stay home or were unable to work, while the government lacked time to determine exactly which households needed income insurance. In that setting, sending payments widely was a practical response to a sudden and widespread interruption of earnings.

PeriodRamey’s rationaleCentral qualification
Spring 2020Broad payments served as emergency income insurance while many people could not work.Government did not have time to identify precisely who needed support.
Later 2020Later payments made less sense to most economists Ramey knows.They added deficit-financed spending as recovery was under way.
March or April 2021Consumption picked up after the third payment.Ramey says it is unclear whether checks, reopening, or vaccine rollout drove the rise.
Ramey’s distinction between the first COVID payment round and later transfers

The harder case is the second round later in 2020 and the third round under the Biden administration in March or April 2021. Consumption did pick up after the third payment, Ramey says, but timing alone does not establish causation. One possibility is that households had finally received enough money that they began spending more. Another is that vaccinations, which began rolling out in January, and the easing of lockdown-related restrictions allowed people to resume purchases they would have made even without the checks.

Most people thought that the second and third round of stimulus made no sense.
Valerie Ramey

Ramey says economists she knows generally regarded the later rounds as difficult to justify because they were financed through deficits, debt-to-GDP ratios were already high, and the economy was coming out of the COVID recession. The issue was not whether a check could induce some spending at the margin. It was whether that spending justified additional debt when reopening itself could explain much of the return in consumption.

Rebate studies can overstate spending when past recipients become controls

Ramey cautions that even household-level evidence with an apparently random payment schedule can produce a misleading estimate of the marginal propensity to consume.

In the rebate program she describes, the government could not send every check at once. Payments were rolled out from April through subsequent months, with timing determined by the final two digits of recipients’ Social Security numbers. Ramey describes that payment timing as effectively random because those digits have no relationship to relevant household characteristics.

That created what looked like a strong research design. Economists could compare a household receiving a check in June with households that did not receive one that month, while controlling for household-specific characteristics and ordinary differences across calendar months. But Ramey says the control group was not actually made up only of unaffected households.

Consider a May recipient and a June recipient. If the May household spends even a small portion of its payment when it arrives, its consumption rises in May and may fall back in June. The June household’s consumption rises in June. Comparing the June recipient with the May recipient can therefore enlarge the apparent difference: one household is rising with the payment while the other is falling from its earlier increase. The supposedly untreated May recipient has become part of the June control group despite still having behavior shaped by the rebate.

A second problem follows from the fact that each household received only one payment. A household paid in May had zero probability of being paid again in June. The relevant receipt dates were therefore not independent in the way some estimation approaches implicitly assumed.

Ramey presents these as subtle errors rather than evidence that the researchers using the method were careless. Many scholars used similar designs without recognizing the issue. She describes the underlying problem as involving “forbidden counterfactuals”: researchers were treating earlier recipients as a valid representation of what current recipients would have done without receiving a payment, when that was not the comparison the rollout permitted.

The implication is not that rebates have no consumption effect. It is that a high estimated household MPC may partly reflect an invalid counterfactual. Randomized payment timing, by itself, does not ensure that the estimated treatment and control groups remain economically comparable over time.

More household realism did not make the aggregate result more credible

The obvious response is that averages conceal the people most likely to spend a transfer. Resource-constrained, hand-to-mouth households may use much more of an additional payment than households with savings or access to credit.

Ramey says her own macroeconomic work built that distinction directly into the model. Her team used a two-agent New Keynesian model: one group consisted of hand-to-mouth consumers who spent all of their income, while the other consisted of Ricardian, or permanent-income, consumers. By changing the proportion of each type, the researchers could match the household-level consumption estimates in the rebate literature.

They then generated artificial households from the model and had what Ramey calls “virtual econometricians” estimate MPCs from those simulated data. Those estimates reproduced the high MPCs found in the micro evidence, including the estimates reported by Parker and his collaborators.

The problem emerged at the macroeconomic level. Once the household responses were embedded in the broader model, it generated large V-shaped movements that Ramey considered implausible in light of the period being studied. Matching the estimated responses of individual households did not establish that the resulting path for the whole economy was believable.

A fuller heterogeneous-agent model would not, in her view, resolve that issue. It would produce V-shapes that were steeper and lasted longer. More detailed household heterogeneity could reinforce rather than soften the aggregate implication she found unconvincing.

Short-run stimulus depends on whether government can purchase something in time

Ramey puts more weight on government purchases than on broad transfers when the aim is short-run stimulus. Building up the military is one example. Infrastructure could also help, but only if projects are genuinely ready to begin and the lags that typically delay infrastructure spending can be reduced.

The constraint is practical rather than semantic. Calling a program infrastructure does not make it timely stimulus. An appropriation must translate into purchases, construction, and work soon enough to affect the downturn it is meant to address. If projects arrive after the relevant period, their economic value may be separate from their usefulness as short-run stabilization.

Asked about federal employment programs modeled on the New Deal, including the Public Works Administration and Civilian Conservation Corps, Ramey says people have studied their effects during the Great Depression. She does not recall those studies finding high stimulus effects, though she says she would need to check the evidence. One complication is that the federal government remained concerned about deficits and sometimes raised taxes to finance programs, so the historical cases did not consistently amount to pure deficit spending.

She suggests that direct federal hiring may not be necessary if the policy aim is to get a project built: government can choose to build something and private companies can carry out the work. In that framing, the relevant practical question is whether the spending can turn quickly into actual purchases and construction, not simply whether workers are formally federal employees.

Ramey notes that New Deal public works left visible assets; she still sees sidewalks marked with the WPA. But visible construction does not settle the policy assessment. Such programs may also have added to debt, and she says the balance deserves further study.

A different tool, which she calls unconventional fiscal stimulus, is a temporary reduction in a value-added tax. The United States does not really have a VAT, she notes, though many other countries do; a temporary sales-tax reduction is the closer US analogy. A consumer considering a car purchase might buy now if the tax is temporarily lower.

The limitation is that the policy may merely pull purchases forward by a few months. If so, it does little to change the broader path of economic activity. It would be more consequential only if it shifted demand from substantially farther into the future, and Ramey describes the extent of that effect as controversial.

Tax rates and deficit reduction can impose different kinds of contraction

Valerie Ramey says tax changes can have larger effects than spending changes because most tax changes alter rates, not simply transfers. Those rates affect willingness to work, invest, and start businesses.

Tax rate changes affect people's willingness to work, people's willingness to start businesses, people's willingness to invest, and all of those sorts of things.
Valerie Ramey · Source

Her example is a high-income lawyer or doctor deciding whether to take on another client. Ramey says that, in California, combined federal and state income taxes can take more than half of earnings for some higher-income people. The after-tax reward, rather than the gross payment, may affect the decision to do additional work. Lower-wage workers can face related incentive effects when additional earnings are reduced by payroll taxes and interact with programs such as food stamps.

The empirical findings she describes are large. Research by David and Christina Romer used a narrative method to identify US tax legislation from 1946 onward and estimated tax multipliers of roughly minus two to minus three: tax increases were associated with substantially lower GDP. Ramey says follow-up work refined the approach while continuing to find large multipliers. Studies in other countries, including work by an IMF team across a wide range of countries, also found large tax effects.

−2 to −3
Tax multiplier range Ramey attributes to Romer and Romer

That evidence is distinct from Ramey’s own assessment of it. She says the estimated effects are bigger than most theories can explain, and that she is pursuing plausibility studies of the tax multiplier for that reason. She has dug through the Romer and Romer study without finding an issue, she says, and has not identified a decisive problem in some related work. But she remains open to continuing to investigate whether there is one.

The same distinction matters for fiscal consolidation. Ramey points to an older, contentious literature on “expansionary fiscal consolidation,” including work by Giavazzi and Pagano on Ireland and another country, in which GDP appeared to rise as governments consolidated debt. The findings generated substantial debate, and she says other research did not necessarily find the same result.

More recent research over roughly the past 10 to 15 years, she says, generally finds that fiscal consolidations reduce GDP regardless of method—but that tax-based consolidations have a more negative effect than spending-based consolidations. The relevant comparison is not between a painless tax increase and a painful spending reduction. Both are associated with lower GDP in the work she describes; the difference is the size of the decline.

Ramey allows that consolidation can rarely raise GDP. She describes Argentina as a striking case: conditions initially looked negative, but later seemed better, and, she says, she thinks the country has a budget surplus. Her point is not that consolidation is generally expansionary, but that such an outcome is possible and unusual.

The frontier, in your inbox tomorrow at 08:00.

Sign up free. Pick the industry Briefs you want. Tomorrow morning, they land. No credit card.

Sign up free