Econometrics, economics, finance, random rants.

Econometrics, economics, finance, random rants...
Showing posts with label Bayesian. Show all posts
Showing posts with label Bayesian. Show all posts

Monday, March 12, 2018

Sims on Bayes

Here's a complementary and little-known set of slide decks from Chris Sims, deeply insightful as always. Together they address some tensions associated with Bayesian analysis and sketch some resolutions. The titles are nice, and revealing. The first is "Why Econometrics Should Always and Everywhere Be Bayesian". The second is "Limits to Probability Modeling" (with Chris' suggested possible sub-title: "Why are There no Real Bayesians?").

Sunday, December 10, 2017

More on the Problem with Bayesian Model Averaging

I blogged earlier on a problem with Bayesian model averaging (BMA) and gave some links to new work that chips away at it. The interesting thing about that new work is that it stays very close to traditional BMA while acknowledging that all models are misspecified.

But there are also other Bayesian approaches to combining density forecasts, such as prediction pools formed to optimize a predictive score. (See, e.g. Amisano and Geweke, 2017, and the references therein.  Ungated final draft, and code, here.)

Another relevant strand of new work, less familiar to econometricians, is "Bayesian predictive synthesis" (BPS), which builds on the expert opinions analysis literature. The framework, which traces to Lindley et al. (1979), concerns a Bayesian faced with multiple priors coming from multiple experts, and explores how to get a posterior distribution utilizing all of the information available. Earlier work by Genest and Schervish (1985) and West and Crosse (1992) develops the basic theory, and new work (McAlinn and West, 2017), extends it to density forecast combination.

Thanks to Ken McAlinn for reminding me about BPS. Mike West gave a nice presentation at the FRBSL forecasting meeting. [Parts of this post are adapted from private correspondence with Ken.]

Sunday, December 3, 2017

The Problem With Bayesian Model Averaging...

The problem is that one of the models considered is traditionally assumed true (explicitly or implicitly) since the prior model probabilities sum to one. Hence all posterior weight gets placed on a single model asymptotically -- just what you don't want when constructing a portfolio of surely-misspecified models. The earliest paper I know that makes and explores this point is one of mine, here. Recent and ongoing research is starting to address it much more thoroughly, for example here and here. (Thanks to Veronika Rockova for sending.)



Sunday, November 26, 2017

Modeling With Mixed-Frequency Data

Here's a bit more related to the FRB St. Louis conference.

The fully-correct approach to mixed-frequency time-series modeling is: (1) write out the state-space system at the highest available data frequency or higher (e.g., even if your highest frequency is weekly, you might want to write the system daily to account for different numbers of days in different months), (2) appropriately treat most of the lower-frequency data as missing and handle it optimally using the appropriate filter (e.g., the Kalman filter in the linear-Gaussian case).  My favorite example (no surprise) is here.  

Until recently, however, the prescription above was limited in practice to low-dimensional linear-Gaussian environments, and even there it can be tedious to implement if one insists on MLE.  Hence the well-deserved popularity of the MIDAS approach to approximating the prescription, recently also in high-dimensional environments.     

But now the sands are shifting.  Recent work enables exact Bayesian posterior mixed-frequency analysis even in high-dimensional structural models.  I've known Schorfheide-Song (2015, JBES; 2013 working paper version here) for a long time, but I never fully appreciated the breakthrough that it represents -- that is, how straightforward exact mixed-frequency estimation is becoming --  until I saw the stimulating Justiniano presentation at FRBSL (older 2016 version here).  And now it's working its way into important substantive applications, as in Schorfheide-Song-Yaron (2017, forthcoming in Econometrica). 

Sunday, July 23, 2017

On the Origin of "Frequentist" Statistics

Efron and Hastie note that the "frequentist" term "seems to have been suggested by Neyman as a statistical analogue of Richard von Mises' frequentist theory of probability, the connection being made explicit in his 1977 paper, 'Frequentist Probability and Frequentist Statistics'".  It strikes me that I may have always subconsciously assumed that the term originated with one or another Bayesian, in an attempt to steer toward something more neutral than "classical", which could be interpreted as "canonical" or "foundational" or "the first and best".  Quite fascinating that the ultimate "classical" statistician, Neyman, seems to have initiated the switch to "frequentist".

Monday, July 3, 2017

Bayes, Jeffreys, MCMC, Statistics, and Econometrics

In Ch. 3 of their brilliant book, Efron and Tibshirani (ET) assert that:
Jeffreys’ brand of Bayesianism [i.e., "uninformative" Jeffreys priors] had a dubious reputation among Bayesians in the period 1950-1990, with preference going to subjective analysis of the type advocated by Savage and de Finetti. The introduction of Markov chain Monte Carlo methodology was the kind of technological innovation that changes philosophies. MCMC ... being very well suited to Jeffreys-style analysis of Big Data problems, moved Bayesian statistics out of the textbooks and into the world of computer-age applications.
Interestingly, the situation in econometrics strikes me as rather the opposite.  Pre-MCMC, much of the leading work emphasized Jeffreys priors (RIP Arnold Zellner), whereas post-MCMC I see uniform at best (still hardly uninformative as is well known and as noted by ET), and often Gaussian or Wishart or whatever.  MCMC of course still came to dominate modern Bayesian econometrics, but for a different reason: It facilitates calculation of the marginal posteriors of interest, in contrast to the conditional posteriors of old-style analytical calculations. (In an obvious notation and for an obvious normal-gamma regression problem, for example, one wants posterior(beta), not posterior(beta | sigma).) So MCMC has moved us toward marginal posteriors, but moved us away from uninformative priors.

Monday, January 23, 2017

Bayes Stifling Creativity?

Some twenty years ago, a leading Bayesian econometrician startled me during an office visit at Penn. We were discussing Bayesian vs. frequentist approaches to a few things, when all of a sudden he declared that "There must be something about Bayesian analysis that stifles creativity.  It seems that frequentists invent all the great stuff, and Bayesians just trail behind, telling them how to do it right".

His characterization rings true in certain significant respects, which is why it's so funny.  But the intellectually interesting thing is that it doesn't have to be that way.  As Chris Sims notes in a recent communication: 
... frequentists are in the habit of inventing easily computed, intuitively appealing estimators and then deriving their properties without insisting that the method whose properties they derive is optimal.  ... Bayesians are more likely to go from model to optimal inference, [but] they don't have to, and [they] ought to work more on Bayesian analysis of methods based on conveniently calculated statistics.

See Chris' thought-provoking unpublished paper draft, "Understanding Non-Bayesians". 

[As noted on Chris' web site, he wrote that paper for the Oxford University Press Handbook of Bayesian Econometrics, but he "withheld [it] from publication there because of the Draconian copyright agreement that OUP insisted on --- forbidding posting even a late draft like this one on a personal web site."] 

Sunday, January 31, 2016

Shrinking VAR's Toward Theory: Supplanting the Minnesota Prior?


A recent post, On Bayesian DSGE Modeling with Hard and Soft Restrictions, ended with: "A related issue is whether 'theory priors' will supplant others, like the 'Minnesota prior'. I'll save that for a later post." This is that later post. Its title refers to Ingram and Whiteman's 1994 classic, entitled "Supplanting the 'Minnesota' Prior: Forecasting Macroeconomic Time Series Using Real Business Cycle Model Priors."

So, shrinking VAR's using DSGE theory priors improves VAR forecasts. Sounds like a victory for economics, with the headline "Using Economic Theory Improves Economic Forecasts!" We'd all like that. We all want that.

But the "victory" is misleading, and more than a little hollow. Lots of shrinkage directions improve forecasts. Indeed almost all shrinkage directions improve forecasts. Real victory would require theory-inspired priors to deliver clear extra improvement relative to other shrinkage directions, but they usually don't. In particular, the Minnesota prior, centered on a simple vector random walk, remains competitive. (See Del Negro and Schorfheide (2004) and Del Negro and Schorfheide (2007).) Sometimes theory priors beat the Minnesota prior by a little, sometimes they lose by a little. It depends on the dataset, the variable, the forecast horizon, etc.

The bottom line: Theory priors seem to be roughly as good as anything else, including the Minnesota prior, but certainly they've not led us to anything resembling wonderful new forecasting success. This seems at best a small forecasting victory for theory priors, but perhaps a victory nonetheless, particularly given the obvious appeal of using a theory prior for Bayesian VAR forecasting that coheres with the theory model used for policy analysis.

Monday, November 23, 2015

On Bayesian DSGE Modeling with Hard and Soft Restrictions

A theory is essentially a restriction on a reduced form. It can be imposed directly (hard restrictions) or used as as a prior mean in a more flexible Bayesian analysis (soft restrictions). The soft restriction approach -- "theory as a shrinkage direction" -- is appealing: coax parameter configurations toward a prior mean suggested by theory, but also respect the likelihood, and govern the mix by prior precision.

(1) Important macro-econometric DSGE work, dating at least to the classic Ingram and Whiteman (1994) paper, finds that using theory as a VAR shrinkage direction is helpful for forecasting.

(2) But that's not what most Bayesian DSGE work now does. Instead it imposes hard theory restrictions on a VAR, conditioning completely on an assumed DSGE model, using Bayesian methods simply to coax the assumed model's parameters toward "reasonable" values.

It's not at all clear that approach (2) should dominate approach (1) for prediction, and indeed research like Del Negro and Schorfheide (2004) and Del Negro and Schorfheide (2007) indicates that it doesn't.

I like (1) and I think it needs renewed attention.

[A related issue is whether "theory priors" will supplant others, like the "Minnesota prior." I'll save that for a later post.]