Research Notebook

Deriving the Gordon Model

July 26, 2026 by Alex

The Gordon model is a mainstay of MBA classes and motivating examples. The model says that a stock’s current share price will equal its expected dividend next year, \mathbb{E}_t[\text{Div}_{t+1}], times a forward multiple, \big( \frac{1}{\mathrm{r} - \mathrm{g}}\big),

(1)   \begin{equation*}\text{Price}_t = \mathbb{E}_t[\text{Div}_{t+1}] \times \bigg( \frac{1}{\mathrm{r} {-} \mathrm{g}}\bigg)\end{equation*}

\mathrm{r} is the stock’s annual risk-adjusted discount rate. A dollar paid out 4 years from now is worth \frac{\mathdollar 1}{(1+\mathrm{r})^4} today. \mathrm{g} = \big(\frac{\mathbb{E}_t[\text{Div}_{t+h}]}{\text{Div}_t}\big)^{1/h}{-}1 is the company’s anticipated dividend-growth rate at every horizon h \geq 1.

Myron Gordon’s idea was to scale up next period’s dividend forecast by a factor of \big( \frac{1}{\mathrm{r}-\mathrm{g}} \big) to capture the present value of the firm’s dividend stream from year (t{+}2) onward. e.g., suppose a firm has promised to pay \mathdollar 5.00/\mathrm{sh} in dividends next year. If investors apply a 10\% discount rate and anticipate 5\% annual dividend growth, then the Gordon model would price the stock at \mathdollar 5.00/\mathrm{sh} \times \big( \frac{1}{10\%-5\%} \big) = \mathdollar 100/\mathrm{sh}.

There’s nothing remotely complicated about this calculation. Researchers all learn the Gordon pricing formula the first week of their PhD program. We’re all very comfortable reasoning in these terms. So it’s easy to forget how much work goes into producing the result. There’s nothing simple or straightforward about it. The key step in the derivation of the Gordon model isn’t about assuming constant parameters. It’s getting rid of the unknown future resale price, which is only a problem when applying present-value logic to stocks.

This post walks through what it takes to derive the Gordon model. I use the Lean proof assistant to do a proper accounting of all the assumptions and steps involved.

Perpetuities

The Gordon model prices stocks by pretending they are bonds. To see the logic, it’s important to understand why it’s easier to apply present-value logic to fixed-income assets. Let’s start with the simplest one: a perpetuity. This is an asset that will pay the same annual coupon starting next year and continuing on until Kingdom come. The present value of this perpetual stream of coupon payments is

(2)   \begin{equation*}\text{Price}_t \;=\; \sum_{h=1}^{\infty} \frac{\text{Coupon}}{(1 {+} \mathrm{r})^h}\end{equation*}

\mathrm{r} > 0\% denotes the annual discount rate.

This geometric series can be simplified as follows

(3)   \begin{equation*}\text{Price}_t \;=\; \text{Coupon} \times \bigg( \frac{1}{\mathrm{r}} \bigg)\end{equation*}

Doubling the annual coupon payment doubles the price of the perpetuity. Lowering the discount rate makes each dollar that a perpetuity delivers in the future more valuable today, thereby increasing the overall price.

Coupon Bonds

An H-year coupon bond works like a perpetuity for the first (H{-}1) years. Both pay the same coupon each year. However, in the final year H, the coupon bond delivers its coupon payment as well as its face value, \mathrm{FV}. The price of a coupon bond reflects the present value of its payout stream

(4)   \begin{equation*}\text{Price}_t \;=\; \sum_{h=1}^{H} \frac{\text{Coupon}}{(1 {+} \mathrm{r})^h} \;+\; \frac{\mathrm{FV}}{(1{+}\mathrm{r})^H}\end{equation*}

The present-value logic is the same. Only the payout stream has changed.

The face value determines the scale of the bond. The size of the coupon is typically reported as a fraction of this number, \mathrm{Coupon} = \mathrm{c} \cdot \mathrm{FV}. Thus, we can write the price as

(5)   \begin{align*}\text{Price}_t \;&=\; \sum_{h=1}^{H} \frac{\text{Coupon}}{(1 {+} \mathrm{r})^h} \;+\; \frac{\mathrm{FV}}{(1{+}\mathrm{r})^H} \\ &=\; \sum_{h=1}^{H} \frac{\mathrm{c} \cdot \text{FV}}{(1 {+} \mathrm{r})^h} \;+\; \frac{\mathrm{FV}}{(1{+}\mathrm{r})^H} \\ &=\; \mathrm{FV} \times \Bigg\{ \sum_{h=1}^{H} \frac{\mathrm{c}\phantom{i}}{(1 {+} \mathrm{r})^h} \;+\; \frac{1\phantom{n}}{(1{+}\mathrm{r})^H} \Bigg\}\end{align*}

The first term in the curly braces, \sum_{h=1}^{H} \frac{\mathrm{c}\,}{(1 {+} \mathrm{r})^h}, is the present value of the coupons spun off by each dollar of face value. The second term, \frac{1\phantom{n}}{(1{+}\mathrm{r})^H}, is the present value of receiving that dollar when the bond matures.

Par Value

The face value is the relevant reference point for pricing bonds. If \text{Price}_t = \mathrm{FV}, then we say that a bond is “priced at par”. Each dollar of face value spins off a \mathrm{c} \cdot \mathdollar 1 coupon once a year for the next H years. When priced at par, the present value of these coupons exactly offsets the loss from having to wait H years to receive the dollar back

(6)   \begin{equation*}\text{@ par:} \qquad \underbrace{\phantom{\Bigg(}\!\!\!\!\mathdollar 1 - \frac{\mathdollar 1\phantom{m}}{(1{+}\mathrm{r})^H}}_{\substack{\text{Loss\phantom{j}from} \\ \text{waiting}}} \;=\; \underbrace{\sum_{h=1}^{H} \frac{\mathrm{c} \cdot \mathdollar 1}{(1 {+} \mathrm{r})^h}}_{\substack{\text{Gain\phantom{j}from} \\ \text{\phantom{t}coupons\phantom{t}}}}\end{equation*}

These two forces offset when the coupon rate equals the discount rate, which is why par bonds have \mathrm{c} = \mathrm{r}.

Bond traders use par pricing as a reference point when performing back-of-the-envelope calculations. When \mathrm{Price}_t = \mathrm{FV}, it doesn’t matter whether you get paid the face value at time (t{+}H) or continue to collect an infinite stream of coupons from year ([t{+}H]{+}1) onward

(7)   \begin{align*}\text{@ par:} \qquad \text{Price}_t \;&=\; \text{FV} \\ &=\; \text{FV} \times \underbrace{\bigg( \frac{\mathrm{c}}{\mathrm{r}} \bigg)}_{=1} \;=\; \underbrace{\text{Coupon} \times \bigg( \frac{1}{\mathrm{r}} \bigg)}_{\text{Perpetuity formula}}\end{align*}

If \mathrm{c} < \mathrm{r}, then the bond’s priced at a discount (below par). If \mathrm{c} > \mathrm{r}, then it’s priced at a premium (above par).

Core Problem

At first glance, it seems like it should be possible to apply the same present-value logic to pricing stocks. The one-year-ahead pricing rule for stocks looks similar to the pricing formula for a one-year coupon bond

(8)   \begin{align*}\text{bond:} \qquad \text{Price}_t \;&=\; \frac{\;\;\!\mathrm{Coupon}\;\;\!}{1{+}\mathrm{r}} + \frac{\;\;\;\;\;\;\;\!\mathrm{FV}\;\;\;\;\;\;\;\!}{1 {+} \mathrm{r}} \\ \text{stock:} \qquad \text{Price}_t \;&=\; \frac{\mathbb{F}_t[\text{Div}_{t+1}]}{1{+}\mathrm{r}} + \frac{\mathbb{F}_t[\text{Price}_{t+1}]}{1 {+} \mathrm{r}}\end{align*}

\mathbb{F}_t[\cdot] denotes investors’ forecast given time-t information. A forecast is just a number in investors’ heads. Nothing guarantees it obeys the laws of probability, so I reserve \mathbb{E}_t[\cdot] for forecasts that do. This distinction will play a big role later on. Chekhov’s gun applies to both screenplays and academic research.

The stock’s forecasted dividend payment next year is kind of like the bond’s coupon. The stock’s anticipated resale price a year from now is sort of like the bond’s face value. However, there’s a key difference. For the bond, \mathrm{Coupon} and \mathrm{FV} are both known at the time of purchase. In ye olde times, when you bought a bond, you received a big piece of paper with a bunch of tabs on the bottom. Each year, you tore off a tab and mailed it in to receive your coupon. When the bond matured, you sent in the last tab and the big sheet of paper to get paid the face value. You couldn’t do this for a stock. Nobody knows \mathrm{Div}_{t+1} or \mathrm{Price}_{t+1} with certainty when you buy a share at time t.

The core problem with using present-value logic to price equities is the forecasted resale price on the right-hand side, \mathbb{F}_t[\mathrm{Price}_{t+1}]. If you’re trying to figure out the functional form of \mathrm{Price}_t, then how are you supposed to know the right value to plug in for next year’s resale price? It is always possible to write an equation in which a stock’s current price equals the discounted payoff to owning a share next year. But for this equation to mean something, you need to remove the dependency of next year’s payoff on the future resale price. Otherwise, the relationship is circular.

Prior to Myron Gordon, people knew how to price bonds using present-value logic. But they didn’t know how to apply similar logic to assets like stocks where fluctuations in the future resale price represent a significant portion of the future payout. Gordon’s 1959 paper showed how to get around this problem by treating stocks like coupon bonds priced at par. Notice that the pricing rule is just a modified perpetuity formula, which includes an adjustment for a growing coupon. This is a bold claim about how stocks get priced. At the very least, it ain’t how people talk about pricing shares of Nvidia or Tesla.

Full Derivation

When researchers describe the Gordon model, they tend to focus on the fact that both \mathrm{r} and \mathrm{g} are constant. This is the least interesting part of the derivation. Let’s walk through what’s required to get from the one-period-ahead present-value formula to Myron Gordon’s result

(9)   \begin{equation*}\text{Price}_t \;=\; \frac{\mathbb{F}_t[\text{Div}_{t+1}] + \mathbb{F}_t[\text{Price}_{t+1}]}{1 + \mathrm{r}_t} \qquad \rightsquigarrow \qquad \text{Price}_t \;=\; \mathbb{E}_t[\mathrm{Div}_{t+1}] \times \bigg( \frac{1}{\mathrm{r}{-}\mathrm{g}} \bigg)\end{equation*}

The discount rate now carries a time subscript. Nothing in one-period-ahead present-value logic requires investors to apply the same discount rate every year, so from here on I let \mathrm{r}_t vary over time. There are 5 steps. The first 4 are where all the real heavy lifting takes place. \mathrm{r} and \mathrm{g} only lose their time subscripts in step #5 after the main formula has been derived. This is just cosmetic tidying-up.

Step #1: Assume Consistent Pricing

To iterate forward, investors must believe the one-period-ahead pricing rule holds at every future date (t{+}h). The same formula that governs today’s price must also govern the price targets in investors’ heads

(10)   \begin{equation*}\text{Price}_{t+h} = \frac{\mathbb{F}_{t+h}[\text{Div}_{(t+h)+1}] + \mathbb{F}_{t+h}[\text{Price}_{(t+h)+1}]}{1 + \mathrm{r}_{t+h}} \qquad \text{for all } h \geq 0\end{equation*}

\mathrm{r}_{t+h} is the one-period discount rate applied to payouts received at time ([t{+}h]{+}1). i.e., each dollar paid the following year is worth \frac{\mathdollar 1}{(1+\mathrm{r}_{t+h})} at time (t{+}h). One more assumption hides in this notation. Discount rates can differ across years, but the entire path \mathrm{r}_t, \mathrm{r}_{t+1}, \mathrm{r}_{t+2}, \ldots is known at time t.

There’s an important economic distinction between applying the formula today, h{=}0, and applying the formula in future years, h \geq 1. Even if most investors don’t think in present-value terms, you could argue that the invisible hand of the market somehow forces the current price to obey the one-period-ahead present-value rule at time t. But you can’t make the same argument for h \geq 1. The pricing formula for these future dates can only exist in investors’ heads. If they don’t think in present-value terms, then there’s no reason for the formula to hold. The claim is a substantive assumption about how investors think.

Step #2: Assume The Tower Property

The tower property says that today’s forecast of next year’s forecast is just today’s forecast

(11)   \begin{equation*}\mathbb{F}_t\big[\mathbb{F}_{t+1}[\,\cdot\,]\big] = \mathbb{F}_t[\,\cdot\,]\end{equation*}

The same is true if we replace next year, h{=}1, with any other longer horizon. Arbitrary forecasts do not have this feature. The tower property is only satisfied by conditional expectations that stem from a well-posed probability space. It is often referred to as the “law of iterated expectations”. I call it the “tower property” to emphasize the distinction between arbitrary forecasts and coherent expectations. From here on out, I write \mathbb{E}_t[\cdot] rather than \mathbb{F}_t[\cdot].

One last thing. Coherent doesn’t mean correct. The law of iterated expectations can be applied to expectations that aren’t objectively correct. It’s a property of the belief structure, not whether these beliefs match the true data-generating process. Biased subjective expectations are a subset of all possible forms of incorrect beliefs. It’s possible to make incorrect forecasts that violate the laws of probability.

Step #3: Iterate Forward Finite Times

The next step is finite induction. Take the one-period-ahead pricing rule and replace the resale price on the right-hand side with its functional form for the following year. If you do this (H{-}1) times, then you get the following expression

(12)   \begin{equation*}\text{Price}_t \;=\; \underbrace{\sum_{h=1}^{H} \frac{\mathbb{E}_t[\text{Div}_{t+h}]}{\prod_{k=0}^{h-1} (1 {+} \mathrm{r}_{t+k})}}_{\text{PV first H dividends}} \;+\; \underbrace{\frac{\mathbb{E}_t[\text{Price}_{t+H}]}{\prod_{k=0}^{H-1} (1 {+} \mathrm{r}_{t+k})}}_{\text{PV resale price}}\end{equation*}

The company’s current share price reflects its expected discounted dividend payments over the next H years plus the present value of the expected resale price H years from now.

Notice that this step doesn’t purge the future resale price from the right-hand side. The date of reckoning has just been pushed farther into the future. In a sense, this makes the original problem worse. If next year’s resale price was hard to fathom, then why would investors have any idea about the price each share might sell for 20 or 100 years in the future? At this point, it’s not obvious progress has been made.

Step #4: Assume Limit Is Well-Behaved

The payoff to iterating forward only occurs when you take the infinite limit, H \to \infty. We’re looking to remove the dependency of the current price on the expected future resale value. For this to happen, we need two things to be true:

  1. Transversality. The expected discounted resale price must go to zero

    (13)   \begin{equation*}\lim_{H \to \infty} \, \frac{\mathbb{E}_t[\text{Price}_{t+H}]}{\prod_{k=0}^{H-1} (1 {+} \mathrm{r}_{t+k})} \;=\; \mathdollar 0\end{equation*}

    The one-period recursion has infinitely many solutions. A rational bubble also satisfies it. Transversality selects the “correct” price, which reflects expected discounted dividends alone.

  2. Convergence. The infinite sum of the stock’s expected discounted dividends must converge to a single finite number, and the answer cannot depend on the order in which the terms get added up. This second requirement is where the bite is. Adding up the discounted dividends in time order and getting a finite limit follows for free from step #3 plus transversality. Absolute convergence does not. Researchers often focus on transversality and take this second condition for granted. But both are strong assumptions. Convergence is a genuine premise of its own, not merely a footnote.

By making both assumptions, it’s possible to eliminate the future resale price entirely. The resulting pricing rule is known as the Dividend Discount Model (DDM). It says that a company’s share price at time t should reflect the discounted value of its expected future dividend stream from time (t{+}1) onward

(14)   \begin{equation*}\text{Price}_t \;=\; \sum_{h=1}^{\infty} \frac{\mathbb{E}_t[\text{Div}_{t+h}]}{\prod_{k=0}^{h-1} (1 {+} \mathrm{r}_{t+k})}\end{equation*}

We’ve now overcome the main challenge in deriving a present-value pricing rule for stocks.

Step #5: Assume Constant Parameters

All the heavy lifting is already done. This last step is about ease-of-use. Most people don’t have clear views about a company’s likely dividend in 2077. They don’t have nuanced views about whether to apply a higher one-year discount rate in 2077 or 2076. So, to make the formula more practical, let’s assume that the stock’s future dividend grows at a constant annual rate

(15)   \begin{equation*}\mathbb{E}_t[\text{Div}_{t+h}] \;=\; (1 + \mathrm{g})^{h-1} \!\cdot \mathbb{E}_t[\text{Div}_{t+1}] \;=\; (1 + \mathrm{g})^h \cdot \text{Div}_t\end{equation*}

Let’s also assume that the same annual discount rate gets applied to every horizon h \geq 0

(16)   \begin{equation*}{\textstyle \prod_{k=0}^{h-1}} (1 {+} \mathrm{r}_{t+k}) \;=\; (1 + \mathrm{r})^h\end{equation*}

The assumption of constant parameters turns the infinite sum with a telescoping product in the denominator into a simple geometric series

(17)   \begin{align*}\text{Price}_t \;&=\; \sum_{h=1}^{\infty} \frac{\mathbb{E}_t[\text{Div}_{t+h}]}{\prod_{k=0}^{h-1} (1 {+} \mathrm{r}_{t+k})} \\ &=\; \sum_{h=1}^{\infty} \frac{(1{+}\mathrm{g})^{h-1} \cdot \mathbb{E}_t[\text{Div}_{t+1}]}{(1 {+} \mathrm{r})^h} \\ &=\; \mathbb{E}_t[\text{Div}_{t+1}] \times \sum_{h=1}^{\infty} \frac{(1{+}\mathrm{g})^{h-1}}{(1 {+} \mathrm{r})^{h\phantom{-1}}} \\ &=\; \mathbb{E}_t[\text{Div}_{t+1}] \times \bigg(\frac{1}{\mathrm{r} {-} \mathrm{g}}\bigg)\end{align*}

Assuming constant \mathrm{r} and \mathrm{g} makes it possible to express the implications of the DDM in a clean way.

With constant parameters, the transversality and convergence assumptions in step #4 boil down to the requirement that \mathrm{r} > \mathrm{g}. If this condition is violated, \mathrm{r} \leq \mathrm{g}, then the present value of the stock’s expected discounted dividend stream will be infinite. e.g., suppose a stock’s future payout stream gets discounted at \mathrm{r}=3\% annually and the company paid a \mathdollar 1.00/\mathrm{sh} dividend last year. If the firm’s dividend-growth rate is \mathrm{g}=4\%, then next year investors expect \mathbb{E}_t[\mathrm{Div}_{t+1}] = \mathdollar 1.04/\mathrm{sh}. Had the firm maintained the same dividend, this cash flow would only be worth \mathdollar 0.97 today. But they expect an extra \mathdollar 0.04 in dividends next year, and this is more than enough to make up for the valuation drag created by discounting.

Assumption Accounting

I use Lean to properly account for all the different assumptions used in the derivation of the Gordon model. The hard part is getting rid of the price forecast on the right-hand side:

  1. Assume that the one-period-ahead pricing rule holds today as well as at every future date. It governs observed prices and the price forecasts in investors’ heads.
  2. Assume that investors’ price forecasts satisfy the tower property. This requires their subjective beliefs to represent conditional expectations that stem from a well-defined subjective probability measure.
  3. Iterate forward a finite number of times, pushing the unknown future resale price far into the future.
  4. Assume that the infinite limit has the properties needed to eliminate the current price’s dependence on the future resale price. These are transversality (a.k.a., no bubbles) and convergence.

The final step is purely cosmetic. It occurs after the troublesome resale price has already been expunged.

  1. Assume constant \mathrm{r} and \mathrm{g}.

The standard telling treats step #5 as the key assumption behind the Gordon model, but the honest ledger shows it is the last and lightest. The core derivation lives in steps #1-4. In addition to maintaining the ledger, the proof in Lean shows that each of the load-bearing assumptions in these steps is necessary: the tower property, transversality, and convergence. There are explicit counterexamples that satisfy everything else and yet break the conclusion.

The point of running the Gordon model through a proof assistant is not the machinery. Every well-trained economist has seen all these ideas before. The issue is that researchers have gotten so familiar with Gordon logic that they often forget all that it requires. The derivation is neither short nor innocent. Lean forces you to reckon with every required step in the proof.

Filed Under: Uncategorized

Trailing PEs Imply Low Elasticities

July 22, 2026 by Alex

A frictionless mean-variance model predicts an aggregate demand elasticity of \nu = 25. Suppose the level of the stock market rises by 1\% on no fundamental news. It’s now 1\% more expensive to buy stocks, but nothing’s changed to make the anticipated payout next year more desirable. Textbook theory says that investors ought to look at this drop in forecasted returns and dump 25\% of their holdings. The data disagrees. There, the aggregate demand elasticity is much much lower. Gabaix-Koijen estimate \nu \approx 0.2.

To get an elasticity that low, investors need to look at the 1\% increase in today’s price and shrug their shoulders. In this note, I show that this is exactly what happens when investors rely on trailing PE ratios when setting price targets. I show that this one simple observation is able to generate a predicted demand elasticity of \nu \approx 0.8. This is well within spitting distance of the estimated 0.2.

The trailing-PE mechanism is kind of like a dogmatic-learning story. Think about a Bayesian investor who treats the current price level as a very precise signal about next year’s payout. Such an investor would face the same demand curve as a trailing-PE user. But the analogy isn’t perfect. The trailing-PE approach doesn’t force next year’s price target to agree with next year’s dividend forecast in present-value terms. When the current price rises, the target rises with it, but the dividend forecast doesn’t budge.

Demand Elasticity

Suppose an asset’s current price changes a tiny bit for non-fundamental reasons. Suppose an investor’s forecasting and allocation rules remain unchanged. How much will her desired position change in response? The answer to this question is called the demand elasticity

(1)   \begin{equation*}\nu \;=\; - \frac{\partial \log \mathrm{Dmnd}}{\partial \log \mathrm{Price}} \;=\; (1 {-} \theta) \;+\; \bigg( \frac{\mu}{\bar{r}} \bigg) \times \eta\end{equation*}

A change in the current price of an asset affects the investor’s demand in two ways. There’s a rebalancing channel, (1{-}\theta), which creates a difference between stock-level and aggregate demand elasticities. There’s also a belief channel. A change in today’s price can impact the investor’s views about next year’s payoff. This is the \big( \tfrac{\mu}{\bar{r}} \big) \times \eta term, and it pins down the overall level. Here’s where this formula comes from.

Let \theta = \mathrm{Dmnd}_t \times \big\{ \frac{\mathrm{Price}_t}{\mathrm{Wealth}_t} \big\} denote an asset’s share of an investor’s wealth. Hold her forecasted return fixed, which switches off the belief channel and keeps a fixed fraction of her wealth in the asset. Under this assumption, a 1\% price rise means the same number of shares now ties up 1\% more of her wealth. Restoring her target weight means trimming shares and parking the proceeds in the rest of the portfolio.

If the investor already holds the asset, then the trim is partly cancelled. The increase in the current price will also revalue her existing position, raising her wealth by \theta percent

(2)   \begin{equation*}\frac{\partial \log \mathrm{Wealth}_t}{\partial \log \mathrm{Price}_t} \;=\; \theta\end{equation*}

This change lifts her target dollar allocation for the asset by the same \theta percent.

The net sale is (1{-}\theta) percent of her shares, which is exactly the share of her portfolio held in other assets. That is the room she has to rebalance into. For a single stock inside a diversified portfolio, we have \theta \approx 0 and (1{-}\theta) \approx 1. For an investor choosing how much to invest in the market as a whole, we have \theta = 1 and (1{-}\theta) = 0. The whole term vanishes.

The belief channel starts with a mapping from the current price to beliefs about future returns. I write the investor’s forecasts as \mathbb{F}_t[\cdot], rather than \mathbb{E}_t[\cdot], because forecasts don’t need to come from a well-posed probability space. A forecast is just a number that the investor writes down.

A share bought today for \mathrm{Price}_t delivers the dollar payout \mathrm{Payout}_{t+1} next year. The realized gross return on this investment will be

(3)   \begin{equation*}1 + \mathrm{Ret}_{t+1} \;=\; \frac{\mathrm{Payout}_{t+1}}{\mathrm{Price}_t} \;=\; e^{\log \mathrm{Payout}_{t+1} - \log \mathrm{Price}_t}\end{equation*}

If you expand the exponential expression around the steady state, then you get the following first-order approximation for the net return

(4)   \begin{equation*}\mathrm{Ret}_{t+1} \;\approx\; (1 {+} \overline{\mathrm{DY}}) \cdot \big( \log \mathrm{Payout}_{t+1} \,-\, \log \mathrm{Price}_t \big) \,+\, \mathrm{constant}\end{equation*}

I use the Gabaix-Koijen calibration values: a long-run dividend yield \overline{\mathrm{DY}} = 3.7\% and an average forecasted return \bar{r} = 4.4\%. A 1\% rise in the payout raises the gross return by (1+\overline{\mathrm{DY}}) \times 1\% \approx 1.037\%, and a 1\% rise in the current price lowers the return by the same amount.

Assumption A1. The investor forecasts the asset’s future payout with a rule \mathbb{F}_t[\mathrm{Payout}_{t+1}] that is differentiable in \log \mathrm{Price}_t, and her return forecast is given by

(5)   \begin{equation*}\mathbb{F}_t[\mathrm{Ret}_{t+1}] \;=\; (1 {+} \overline{\mathrm{DY}}) \cdot \big( \log \mathbb{F}_t[\mathrm{Payout}_{t+1}] \,-\, \log \mathrm{Price}_t \big) \,+\, \mathrm{constant}\end{equation*}

Notice what A1 does not assume: present-value logic. Nothing forces today’s price to equal her payout forecast discounted at a required return, so her payout forecast can move independently of the multiple, and a shock to \mathbb{F}_t[\mathrm{EPS}_{t+1}] need not have any impact on the PE. A1 does not impose Gordon logic, either directly or approximately as in Campbell-Shiller. A1 pins down the investor’s return forecast for next year as a function of her payout forecast and the current price.

Differentiating gives the belief drag, \mu. This parameter represents the price sensitivity of the investor’s forecasted return for the upcoming year

(6)   \begin{equation*}\mu \;=\; {-}\frac{\partial \, \mathbb{F}_t[\mathrm{Ret}_{t+1}]}{\partial \log \mathrm{Price}_t} \;=\; (1 {+} \overline{\mathrm{DY}}) \times \bigg( 1 \,-\, \frac{\partial \log \mathbb{F}_t[\mathrm{Payout}_{t+1}]}{\partial \log \mathrm{Price}_t} \bigg)\end{equation*}

If the current price goes up by 1\% and nothing else changes, how much will the investor’s return forecast fall in response?

In a frictionless mean-variance model, the investor observes the asset’s current price. But this information doesn’t impact how she values the stock. Her payout forecast is built from fundamentals alone, so \frac{\partial \log \mathbb{F}_t[\mathrm{Payout}_{t+1}]}{\partial \log \mathrm{Price}_t} = 0 and \mu = (1 {+} \overline{\mathrm{DY}}) \times (1{-}0) \approx 1.037. She suffers the full drag.

The belief drag has units of percent per year. \mu is a change in the asset’s anticipated return over the next twelve months. The average forecasted return \bar{r} has the same units. So the ratio \big( \tfrac{\mu}{\bar{r}} \big) is dimensionless, which an elasticity term has to be. The two terms combine to form the elasticity of the forecasted return with respect to the current price.

A textbook investor in a frictionless mean-variance model has belief drag \mu = (1 {+} \overline{\mathrm{DY}}) \approx 1.037. A 1\% increase in the current price level will lower her forecasted return for next year by 1.037\%\mathrm{pt}. When we compare this effect to the long-run average return forecast, \bar{r}=4.4\%, we get a relative change of \big( \tfrac{1.037\%\mathrm{pt}}{4.4\%} \big) \approx 23.6\%. The original 1.037\%\mathrm{pt} drag on the asset’s forecasted return may not sound like much, but it’s a big deal compared to the average return forecast.

\big( \tfrac{\mu}{\bar{r}} \big) measures the percent decline in the forecasted return when the price rises 1\%. The parameter \eta converts this elasticity of forecasted returns into an elasticity of demand. You might think this conversion requires a full-fledged asset-pricing model. It turns out any allocation rule with the following form will do.

Assumption A2. The dollar allocation is \mathrm{Wealth}_t \times w\big(\mathbb{F}_t[\mathrm{Ret}_{t+1}]\big) for a smooth increasing rule w(\cdot).

The number of shares that the investor demands can be written as w times the ratio of her initial wealth and the asset’s share price

(7)   \begin{equation*}\mathrm{Dmnd}_t \;=\; w\big(\mathbb{F}_t[\mathrm{Ret}_{t+1}]\big) \times \bigg\{ \frac{\mathrm{Wealth}_t}{\mathrm{Price}_t}\bigg\}\end{equation*}

e.g., mean-variance preferences deliver the special case w(x) = \big(\frac{1}{\gamma \cdot \sigma^2}\big) \cdot x.

\eta(\bar{r}) represents the elasticity of the investor’s dollar position in the asset with respect to her forecasted return next year evaluated at the asset’s long-run average forecast

(8)   \begin{equation*}\eta(\bar{r}) \;=\; \bar{r} \times \bigg\{ \frac{w'(\bar{r})}{w(\bar{r})} \bigg\}\end{equation*}

e.g., if an asset’s forecasted return improves by 1\%, from \bar{r} = 4.4\% to 4.444\%, the investor scales her dollar allocation in the asset up by \eta(\bar{r}) \times 1\%.

Note that \eta(\bar{r}) \approx 1 to leading order. For mean-variance preferences, we have \eta(\bar{r}) = 1 exactly. To see why, take any smooth rule with w(0) = 0. A Taylor expansion gives w(\bar{r}) = w'(0) \cdot \bar{r} \cdot (1 {+} O(\bar{r})), so

(9)   \begin{equation*}\eta(\bar{r}) \;=\; 1 \,+\, \frac{1}{2} \cdot \bigg\{\frac{w''(0)}{w'(0)}\bigg\} \times \bar{r} \,+\, O(\bar{r}^2)\end{equation*}

If w(x) = \big(\frac{1}{\gamma \cdot \sigma^2}\big) \cdot x, then \frac{\mathrm{d}w}{\mathrm{d}x} = \big(\frac{1}{\gamma \cdot \sigma^2}\big) and \frac{\mathrm{d}^nw}{\mathrm{d}x^n} = 0 for all n \geq 2. Hence, we have \eta(\bar{r}) = 1 for all \bar{r}. However, any preference specification with w(0)=0 and w''(0)=0 would do the same to leading order.

The pieces now assemble by the chain rule. First, take logs of the demand rule

(10)   \begin{equation*}\log \mathrm{Dmnd}_t \;=\; \log w\big(\mathbb{F}_t[\mathrm{Ret}_{t+1}]\big) \,+\, \log \mathrm{Wealth}_t \,-\, \log \mathrm{Price}_t\end{equation*}

Next, differentiate each term with respect to \log \mathrm{Price}_t. The wealth term contributes \theta and the price term contributes -1. Together they are the rebalancing channel, (1{-}\theta).

The first \log w\big(\mathbb{F}_t[\mathrm{Ret}_{t+1}]\big) term is the belief channel. The forecasted return falls by \mu per unit of \log \mathrm{Price}_t, and log dollars move by \big\{ \frac{w'(\bar{r})}{w(\bar{r})} \big\} = \big( \tfrac{1}{\bar{r}} \big) \times \eta per unit of forecasted return. Flipping the sign delivers the headline formula

(11)   \begin{equation*}\nu \;=\; -\frac{\partial \log \mathrm{Dmnd}_t}{\partial \log \mathrm{Price}_t} \;=\; (1 {-} \theta) \;+\; \bigg( \frac{\mu}{\bar{r}} \bigg) \times \eta\end{equation*}

A 1\% price increase lowers next year’s return forecast by \mu percentage points. Dividing by \bar{r} converts this drop into an elasticity of returns, and multiplying by \eta translates it into a demand elasticity.

We can now cleanly state the inelastic-markets result of Gabaix-Koijen. Start with the textbook prediction. Consider a mean-variance investor in a frictionless model where the dividend yield is \overline{\mathrm{DY}} = 3.7\% and the long-run average return is \bar{r}=4.4\%. The predicted belief drag is \mu = (1{+}\overline{\mathrm{DY}}) \times (1-0) = 1.037. This change in next year’s return forecast represents roughly \frac{1.037\%\mathrm{pt}}{4.4\%} \approx 23.6\% of the average return forecast. With \eta = 1, this return elasticity translates to a 23.6\% change in demand. For the market as a whole, \theta = 1 and (1{-}\theta)=0, so \nu = 23.6. For an individual stock, \theta=0 and (1{-}\theta)=1, giving a demand elasticity that is one turn higher, \nu = 1 {+} 23.6 = 24.6. Gabaix-Koijen estimate an aggregate demand elasticity of \hat{\nu} = 0.2. The textbook prediction is off by two orders of magnitude, \frac{23.6}{0.2} \approx 118!

Trailing PE Ratio

Sell-side analysts typically describe setting one-year-ahead price targets using a two-step process. First, an analyst forecasts the stock’s EPS over the next year based on non-price information. Then, the analyst capitalizes this short-term earnings forecast into a price target using the company’s current multiple, \mathrm{PE}_t = \big( \frac{\mathrm{Price}_t}{\mathrm{EPS}_t} \big), which is known as the trailing PE ratio

(12)   \begin{equation*}\mathbb{F}_t[\mathrm{Price}_{t+1}] = \mathbb{F}_t[\mathrm{EPS}_{t+1}] \times \mathrm{PE}_t\end{equation*}

This price forecast implies that next year’s return forecast will consist of two components: the firm’s anticipated dividend yield and forecasted earnings growth

(13)   \begin{align*}\mathbb{F}_t[\mathrm{Ret}_{t+1}] \;&=\; \bigg(\frac{\mathbb{F}_t[\mathrm{Div}_{t+1}]}{\mathrm{Price}_t}\bigg) \,+\, \bigg(\frac{\mathbb{F}_t[\mathrm{Price}_{t+1}] - \mathrm{Price}_t}{\mathrm{Price}_t}\bigg) \\ &=\; \bigg(\frac{\mathbb{F}_t[\mathrm{Div}_{t+1}]}{\mathrm{Price}_t}\bigg) \,+\, \bigg(\frac{\mathbb{F}_t[\mathrm{EPS}_{t+1}] {\times} \mathrm{PE}_t - \mathrm{EPS}_t {\times} \mathrm{PE}_t}{\mathrm{EPS}_t {\times} \mathrm{PE}_t}\bigg) \\ &=\; \bigg(\frac{\mathbb{F}_t[\mathrm{Div}_{t+1}]}{\mathrm{Price}_t}\bigg) \,+\, \bigg(\frac{\mathbb{F}_t[\mathrm{EPS}_{t+1}] - \mathrm{EPS}_t}{\mathrm{EPS}_t}\bigg)\end{align*}

Since the analyst uses today’s PE ratio to forecast next year’s price, the multiple drops out of the price appreciation term. Regardless of the current level, the analyst anticipates that the firm’s price will grow at the same rate as its earnings.

Notice that only one of the two components of the analyst’s return forecast includes the current price. This is clearly going to have implications for demand elasticities. To see what those are, let’s look at a concrete example. Consider a company that realized earnings of \mathrm{EPS}_t = \mathdollar 5.00/\mathrm{sh} last year. Over the next year, analysts anticipate that the company’s earnings will grow by 0.7\% to \mathbb{F}_t[\mathrm{EPS}_{t+1}] = \mathdollar 5.035/\mathrm{sh}. The stock starts out trading at \mathrm{Price}_t = \mathdollar 100.00/\mathrm{sh}, giving the firm a trailing multiple of \mathrm{PE}_t = \frac{\mathdollar 100.00/\mathrm{sh}}{\mathdollar 5.00/\mathrm{sh}} = 20\times. Given how the market is currently pricing the company’s earnings, analysts expect the firm to be trading at a price of \mathbb{F}_t[\mathrm{Price}_{t+1}] = \mathdollar 5.035/\mathrm{sh} \times 20 = \mathdollar 100.70/\mathrm{sh} next year. The company has committed to paying \mathdollar 3.73/\mathrm{sh} in dividends next year, giving the firm an anticipated dividend yield of 3.7\% and a return forecast of \mathbb{F}_t[\mathrm{Ret}_{t+1}] = 0.7\% + 3.7\% = 4.4\%.

Now, imagine that the company’s current price suddenly increases by 1\% to \mathrm{Price}_t = \mathdollar 101.00/\mathrm{sh}. The firm’s earnings over the last twelve months don’t move, \mathrm{EPS}_t = \mathdollar 5.00/\mathrm{sh}. Nothing about the company’s fundamentals change, either. Analysts still have the same next-twelve-month earnings forecast, \mathbb{F}_t[\mathrm{EPS}_{t+1}] = \mathdollar 5.035/\mathrm{sh}. But the company’s higher current price means that this short-term forecast will get capitalized at a higher multiple, \mathrm{PE}_t = \frac{\mathdollar 101.00/\mathrm{sh}}{\mathdollar 5.00/\mathrm{sh}} = 20.2\times, when setting a price target, \mathbb{F}_t[\mathrm{Price}_{t+1}] = \mathdollar 5.035/\mathrm{sh} \times 20.2 = \mathdollar 101.71/\mathrm{sh}. Yet the higher price target has no impact on analysts’ beliefs about future price growth because the current earnings are also being priced using a multiple that is 0.2\times higher. The higher current price level only affects analysts’ return forecast for next year by diluting the dividend yield. Instead of \frac{\mathdollar 3.73/\mathrm{sh}}{\mathdollar 100/\mathrm{sh}} = 3.7\%, analysts now anticipate a dividend yield of \frac{\mathdollar 3.73/\mathrm{sh}}{\mathdollar 101/\mathrm{sh}} = 3.664\%.

Prior to the price increase, the company’s forecasted payout was the \mathdollar 100.70/\mathrm{sh} target price plus the \mathdollar 3.73/\mathrm{sh} forecasted dividend, which came out to \mathdollar 104.43/\mathrm{sh} total. The resale price contributed \big( \frac{1}{1 + \overline{\mathrm{DY}}} \big) \approx 96.3\% of the total payout. Following the unilateral 1\% price increase, the company’s multiple expanded and its price target also rose by 1\%. But its dividend forecast stood still. Hence, analysts’ forecasted payout didn’t rise by a full 1\%. The drag on analysts’ beliefs is thus

(14)   \begin{equation*}\mu \;=\; (1 {+} \overline{\mathrm{DY}}) \times \bigg( 1 - \frac{1}{1 {+} \overline{\mathrm{DY}}} \bigg) \;=\; \overline{\mathrm{DY}} \;=\; 3.7\%\end{equation*}

This drag gets compared to the same average forecast as before, \bar{r} = 4.4\%. But, given the much smaller starting value, 0.037 vs 1.037, the resulting elasticity of returns is much smaller, \big( \frac{\mu}{\bar{r}} \big) = \frac{3.7\%\mathrm{pt}}{4.4\%} \approx 0.8 rather than 23.6. Assuming \eta = 1, you get a single-stock demand elasticity of \nu = 1 + 0.8 \approx 1.8 and an aggregate demand elasticity of \nu \approx 0.8.

The trailing-PE approach gets you from 23.6 down to 0.8. The remaining distance, from 0.8 down to the estimated 0.2, is likely due to \eta rather than beliefs. Mandates, inertia, and the other demand-side frictions at the center of the inelastic-markets literature all mute the position response, which corresponds to \eta < 1. An \eta \approx 0.25 closes the gap. On this reading, the trailing-PE rule and demand-side frictions are complements, not competitors. Beliefs deliver the first factor of \frac{23.6}{0.8} \approx 30. Frictions deliver the last factor of \frac{0.8}{0.2} \approx 4.

Learning Story

The trailing-PE rule hardcodes the link between this year’s multiple and next year’s multiple. A learning story can deliver something similar without hardcoding anything. Think about Grossman-Stiglitz. When the current price goes up by 1\%, an investor might worry that everyone else knows something she doesn’t. And, as a result, she might raise her forecast of the future payout. This is the story in Bastianello (2026).

Consider an investor who solves the following Gaussian inference problem. The investor has prior beliefs about the stock’s future payout

(15)   \begin{equation*}\log \mathrm{Payout}_{t+1} \;\sim\; \mathrm{Normal}\big( \log \mathrm{Prior}_t, \, 1 \big)\end{equation*}

For clarity, I’ve suppressed a constant term, which reflects discounting and the risk premium. I’ve also normalized the prior variance to 1. The current price level is a noisy signal about the future payout

(16)   \begin{equation*}\log \mathrm{Price}_t \;\sim\; \mathrm{Normal}\big( \log \mathrm{Payout}_{t+1}, \; 1/\tau \big)\end{equation*}

\tau > 0 is the precision of the price signal. A larger value of \tau implies that prices are more informative about the stock’s likely payout next year.

Standard Gaussian-updating rules imply that the investor’s posterior beliefs about the payout will be a weighted average

(17)   \begin{equation*}\mathbb{E}_t[\log \mathrm{Payout}_{t+1}|\log \mathrm{Price}_t] \;=\; (1 {-} \lambda) \cdot \log \mathrm{Prior}_t \,+\, \lambda \cdot \log \mathrm{Price}_t\end{equation*}

Her beliefs are a true conditional expectation, so I write them with \mathbb{E}_t[\cdot] rather than \mathbb{F}_t[\cdot]. The weights reflect the precision of the price signal. The investor leans more heavily on the current price when it is a more precise signal about the stock’s future payout, \lambda = \big(\frac{\tau}{1 + \tau}\big).

From here, it’s straightforward to derive the key inputs to the demand-elasticity formula. Start with the belief drag. Differentiating the log expected payout with respect to \log \mathrm{Price}_t gives

(18)   \begin{equation*}\frac{\partial \log \mathbb{E}_t[\mathrm{Payout}_{t+1}|\log \mathrm{Price}_t]}{\partial \log \mathrm{Price}_t} \;=\; \lambda \qquad \rightsquigarrow \qquad \mu \;=\; (1 {+} \overline{\mathrm{DY}}) \cdot (1 {-} \lambda)\end{equation*}

Learning scales the entire textbook drag down by a factor of (1{-}\lambda). There’s no impact on how this drag gets converted into a return elasticity. For the market as a whole, we get a predicted demand elasticity of

(19)   \begin{equation*}\nu \;=\; (1{-}\theta) \,+\, \bigg( \frac{(1{+}\overline{\mathrm{DY}}) \cdot (1{-}\lambda)}{\bar{r}} \bigg) \times \eta\end{equation*}

If the price signal is entirely uninformative, \lambda = 0, you get back the original formula. If the price signal is perfectly revealing, \lambda = 1, the entire belief channel dies. Only the rebalancing term remains.

Partial Symmetry

I motivated the learning story above by pointing out that, if you squint, it looks a bit like the trailing-PE approach. In both cases, an investor sees the current price level change and assumes that most of the change will propagate into the future payout. Using a trailing PE is kind of like viewing the current price level as a very precise signal about the future payout.

Consistent with this intuition, it’s possible to make the two mechanisms produce identical elasticities. All you have to do is equate the belief drags. The trailing-PE approach says \mu = \overline{\mathrm{DY}}. The learning story says \mu = (1{+}\overline{\mathrm{DY}}) \cdot (1{-}\lambda). For both to produce the same value, you need

(20)   \begin{equation*}\overline{\mathrm{DY}} \;=\; (1{+}\overline{\mathrm{DY}}) \cdot (1{-}\lambda^{\star}) \qquad \rightsquigarrow \qquad \lambda^{\star} \;=\; \frac{1}{1{+}\overline{\mathrm{DY}}}\end{equation*}

Assuming \overline{\mathrm{DY}} = 3.7\% would imply that \lambda^{\star} \approx 0.963. A learner who put 96.3\% weight on the current price signal would have the same demand curve as an analyst who set price targets using a trailing PE, \nu = 1.8 for a single stock and 0.8 for the aggregate at \eta = 1. So there is a precise sense in which using a trailing PE and putting a lot of weight on the current price level are symmetric.

Gabaix-Koijen estimates \nu \approx 0.2. To match that result, a learning story would need to put weight on the current price of \lambda \approx 97\% or higher. This is worth pausing on. The learning model in Bastianello is built from standard Bayesian ingredients. The paper never talks in terms of trailing multiples. Yet at the weight implied by the data, the investor submits the same demand curve as an analyst using a trailing PE ratio. Fit to the data, the learning story does not offer an alternative to the trailing-PE mechanism. It approximates it.

But the symmetry isn’t perfect. And the way that it breaks is interesting. The key thing in the trailing-PE story is not the trailing PE. It is that the investor treats next period’s EPS forecast and the current price as unrelated. There is no present-value calculation connecting the two. Her EPS forecast comes from sales and margins, her level comes from the market, and neither number is evidence about the other.

The difference shows up in how a price shift reaches each investor. For the trailing-PE analyst, a 1\% rise in the current price level moves her price target by the same 1\%. So the capital-gain piece of her forecasted return never budges. The entire impact of the price shock arrives through the stock’s dividend yield, and that is why her drag equals \overline{\mathrm{DY}} exactly.

By contrast, an investor who learns about the stock’s future payout from its current price sees both components of her payout forecast change. The price rise is news about fundamentals, which affects her beliefs about next year’s dividend and next year’s resale price. In this scenario, the investor’s belief drag gets spread across the whole forecast rather than concentrated in the dividend.

The apportionment can be made exact. At \lambda = \lambda^{\star} = 96.3\%, both investors mark up their payout forecast by 0.963\%\mathrm{pt} in response to a 1\% price increase, leaving the same 0.037\%\mathrm{pt} shortfall. But the two stories aren’t equivalent. They each place that 0.037\%\mathrm{pt} gap in different places. The trailing-PE analyst increases her price target by 1\%. Her dividend-yield forecast absorbs the entire shortfall. The learner pins 0.036\%\mathrm{pt} on the capital gain and 0.001\%\mathrm{pt} on the dividend yield. In the running example, both investors would forecast the same payout next year, \mathdollar 105.44/\mathrm{sh}. The analyst gets there as \mathdollar 101.71 + \mathdollar 3.73. The learner gets there as \mathdollar 101.67 + \mathdollar 3.77. Same number, different tickets, and the dividend line is the tell.

Filed Under: Uncategorized

Inelastic Markets ~ Flat SML

July 21, 2026 by Alex

Stock returns are typically higher than bond returns. The average difference is somewhere in the neighborhood of \mathbb{E}[\mathrm{Mkt}]{-}\mathrm{rf} \approx 4\%. Academic researchers call this the equity risk premium.

The CAPM is the most famous asset-pricing model in the literature. The theory predicts that each stock’s expected excess return will be proportional to the expected excess return on the market portfolio, \mathbb{E}[\mathrm{Ret}_n {-} \mathrm{rf}] \;=\; \beta_n \times \mathbb{E}[\mathrm{Mkt} {-} \mathrm{rf}]. Suppose you plot each stock’s expected excess return, \mathbb{E}[\mathrm{Ret}_n] {-} \mathrm{rf} (y-axis), against its market beta, \beta_n (x-axis). The resulting line is called the “SML” (security market line)

(1)   \begin{equation*}\mathbb{E}[\mathrm{Ret}_n] {-} \mathrm{rf} \;=\; \beta_n \times \lambda\end{equation*}

The CAPM predicts that you ought to get a line with slope \lambda = \mathbb{E}[\mathrm{Mkt}] {-} \mathrm{rf}.

The same simple model predicts that aggregate demand elasticity should be roughly

(2)   \begin{equation*}\nu \;\approx\; \frac{1}{\mathbb{E}[\mathrm{Mkt}] {-} \mathrm{rf}}\end{equation*}

If the level of the stock market goes up by 1\% for non-fundamental reasons, then it suddenly costs more money to buy the same future cash flows. Textbook theory predicts that investors ought to reduce their holdings. Demand elasticity tells you how much. If \nu = 2, then a unilateral {+}1\% increase in the current price of equities will cause a {-}2\% reduction in investors’ stock holdings.

The security market line (SML) is much flatter than theory predicts. Instead of \lambda = 4\%, the slope of the SML is basically zero. The aggregate stock market is much less elastic than theory predicts. Instead of \nu = \frac{1}{4\%} = 25, Gabaix-Koijen puts the number in the neighborhood of 0.2. Investing an extra \mathdollar 1 in the stock market raises its aggregate value by about \mathdollar 5.

This note shows both findings are related. They’re two perspectives on the same underlying problem.

CARA-Normal Model

Start with the simplest possible model: two periods (today and next year); one investor; one risky asset (the stock market); one riskless bond. Buying a share of the stock market costs \mathrm{Price}_t today. If you own a share of the stock market today, then next year you are entitled to receive \mathrm{Payout}_{t+1}. The riskless bond costs \mathdollar 1 today and will pay (1 {+} \mathrm{rf}) next year.

The representative investor has “constant absolute risk aversion” (CARA) preferences, \mathrm{U}(\mathrm{C}) = {-}\tfrac{1}{\gamma} \cdot e^{-\gamma \cdot \mathrm{C}}. The stock market’s payout next year is normally distributed with variance \sigma^2 > 0. The investor starts with wealth \omega > \mathdollar 0. Today, he must choose how much to consume, \mathrm{C}_t, and how many shares of the risky asset to purchase, \mathrm{Q}_t. His goal is to maximize \mathrm{U}(\mathrm{C}_t) + \mathbb{E}\big[ \, e^{-\rho} \cdot \mathrm{U}(\mathrm{C}_{t+1}) \, \big] where \rho > 0 is his rate of time preference. The investor parks any remaining wealth, (\omega - [\mathrm{C}_t {+} \mathrm{Q}_t \cdot \mathrm{Price}_t]), in the riskfree bond. Next year, the investor eats the combined payout from his risky and safe investments

(3)   \begin{equation*}\mathrm{C}_{t+1} \;=\; (1 {+} \mathrm{rf}) \times \big( \, \omega - [\mathrm{C}_t {+} \mathrm{Q}_t \cdot \mathrm{Price}_t] \, \big) \,+\, \mathrm{Q}_t \cdot \mathrm{Payout}_{t+1}\end{equation*}

Let \psi > 0 denote the supply of shares in circulation. The market clears when the investor’s demand for the risky asset equals the number of available shares, \mathrm{Q}_t = \psi. An equilibrium is an allocation, \{\mathrm{C}_t,\,\mathrm{Q}_t,\,\mathrm{C}_{t+1} \}, and a current price level for the risky asset, \{ \mathrm{Price}_t \}, such that (i) the allocation solves the investor’s optimization problem given the price, and (ii) the price clear the market given the investor’s allocation.

The payout to owning each share of the risky asset is positive on average. So, holding an extra share will lead to slightly higher consumption next year. At the optimum, this benefit will be exactly canceled out by the cost of the required reduction in consumption today, with each side weighted by its marginal utility

(4)   \begin{equation*}\mathrm{U}'(\mathrm{C}_t) \times \mathrm{Price}_t \;=\; \mathbb{E}\big[ \, e^{-\rho} \cdot \mathrm{U}'(\mathrm{C}_{t+1}) \times \mathrm{Payout}_{t+1} \, \big]\end{equation*}

This is the Euler equation. An extra \mathdollar 1 that arrives in bad times (consumption is low; marginal utility is high) counts for more than a \mathdollar 1 that arrives in good times (high consumption; low marginal utility).

Here’s how to solve this model. First, note that the riskless asset costs \mathdollar 1 today and is guaranteed to deliver (1{+}\mathrm{rf}) next year, so its Euler equation is

(5)   \begin{equation*}\mathrm{U}'(\mathrm{C}_t) \;=\; \mathbb{E}\big[ \, e^{-\rho} \cdot \mathrm{U}'(\mathrm{C}_{t+1}) \times (1{+}\mathrm{rf}) \, \big]\end{equation*}

If we replace the \mathrm{U}'(\mathrm{C}_t) in Equation (4) with this expression, then the e^{-\rho} cancels out, and the price becomes a marginal-utility-weighted average of the discounted payout. The definition of a covariance plus Stein’s lemma turn that weighted average into \mathbb{E}[\mathrm{Payout}_{t+1}] - \gamma \times \mathbb{C}\mathrm{ov}[\mathrm{C}_{t+1}, \, \mathrm{Payout}_{t+1}]. What’s more, Equation (3) shows that next year’s consumption will be linear in the payout, so \mathbb{C}\mathrm{ov}[\mathrm{C}_{t+1}, \mathrm{Payout}_{t+1}] = \mathrm{Q}_t \cdot \sigma^2. Given market clearing, \mathrm{Q}_t = \psi, this leads to the following pricing rule

(6)   \begin{equation*}\mathrm{Price}_t = \frac{\mathbb{E}[\mathrm{Payout}_{t+1}] - \gamma \cdot \sigma^2 \cdot \psi}{1 + \mathrm{rf}}\end{equation*}

Each extra share makes next year’s consumption covary more strongly with the payout, so the marginal buyer demands a larger discount. The numerator is the expected payout minus an adjustment for risk. The denominator adjusts for the time cost of money.

Security Market Line

Textbook asset-pricing theory puts every asset on a single line. Expected excess returns ought to be proportional to betas, and the constant of proportionality ought to be the equity risk premium. To see where this prediction comes from, define the stochastic discount factor as discounted marginal utility growth, \mathrm{SDF}_{t+1} = e^{-\rho} \cdot \tfrac{\mathrm{U}'(\mathrm{C}_{t+1})}{\mathrm{U}'(\mathrm{C}_{t})}. Let n = 1, \ldots, N index the cross-section of risky assets… i.e., each stock in the stock market. The same SDF should price every one of them. If you use the SDF to write stock n‘s Euler equation and divide by its current price, then you get a statement about its expected return

(7)   \begin{equation*}1 \;=\; \mathbb{E}\bigg[ \, \mathrm{SDF}_{t+1} \times \underbrace{\bigg(\frac{\mathrm{Payout}_{n,t+1}}{\mathrm{Price}_{n,t}}\bigg)}_{1+\mathrm{Ret}_{n,t+1}} \, \bigg]\end{equation*}

Going forward, I’ll suppress time subscripts where it causes no confusion.

If you subtract the Euler equation for the riskless bond, 1 = \mathbb{E}[ \, \mathrm{SDF} \times (1{+}\mathrm{rf}) \, ], then you get

(8)   \begin{equation*}0 \;=\; \mathbb{E}\big[ \, \mathrm{SDF} \times (\mathrm{Ret}_n{-}\mathrm{rf}) \, \big]\end{equation*}

The difference being priced, (\mathrm{Ret}_n{-}\mathrm{rf}), is stock n‘s excess return. It is the payout from a long/short portfolio that sells riskfree bonds and uses the proceeds to buy shares of the risky asset.

Now consider applying the definition of a covariance, \mathbb{C}\mathrm{ov}[X, \, Y] = \mathbb{E}[X \cdot Y] - \mathbb{E}[X] \cdot \mathbb{E}[Y], to this excess-return SDF formula

(9)   \begin{align*}0 &= \mathbb{E}[ \, \mathrm{SDF} \times (\mathrm{Ret}_{n}{-}\mathrm{rf}) \, ] \\ &= \mathbb{E}[\,\mathrm{SDF}\,] \times (\mathbb{E}[\mathrm{Ret}_{n}]{-}\mathrm{rf}) + \mathbb{C}\mathrm{ov}[ \, \mathrm{SDF}, \, \mathrm{Ret}_{n} \, ]\end{align*}

By rearranging terms, we can arrive at the following expression

(10)   \begin{align*}\mathbb{E}[\mathrm{Ret}_{n}] {-} \mathrm{rf} &= \bigg( \frac{\mathbb{C}\mathrm{ov}[ -\mathrm{SDF}, \, \mathrm{Ret}_{n} ]}{\mathbb{E}[\mathrm{SDF}]} \bigg) \\ &= \underbrace{\bigg( \frac{\mathbb{C}\mathrm{ov}[ -\mathrm{SDF}, \, \mathrm{Ret}_{n} ]}{\mathbb{V}\mathrm{ar}[\mathrm{SDF}]} \bigg)}_{\beta_n} \times \underbrace{\bigg( \frac{\mathbb{V}\mathrm{ar}[ \mathrm{SDF}]}{\mathbb{E}[\mathrm{SDF}]} \bigg)}_{\lambda} \end{align*}

The first \beta_n term tells you how much asset n‘s return tends to comove with the SDF. The SDF captures growth in marginal utility. It is high when the economy enters into bad times. That’s when it becomes more valuable to have an extra dollar. Thus, stocks that tend to do well during booms and poorly during crashes will have large values of \beta_n. The second \lambda term is constant across stocks. It answers the following question: If a stock’s \beta_n goes up by one unit, how much higher will its excess returns be on average?

In the CARA-normal model, the SDF is approximately linear in the change in aggregate consumption

(11)   \begin{equation*}\mathrm{SDF} \;=\; e^{-\rho} \cdot e^{-\gamma \cdot \Delta \mathrm{C}} \;\approx\; a - b \cdot \Delta \mathrm{C} \end{equation*}

And what’s the main driver of the change in aggregate consumption in this model? The payout on the risky asset next year. Equation (3) shows that \mathrm{C}_{t+1} is linear in \mathrm{Payout}_{t+1}, and \mathrm{Payout}_{t+1} = (1 {+} \mathrm{Mkt}_{t+1}) \cdot \mathrm{Price}_t by definition. So the SDF is approximately linear in the market’s return.

Under these assumptions, you can estimate stock n‘s \beta_n by running a time-series regression of realized returns on the market return. Then, if you plot each stock’s average excess return against its estimated \beta_n, the slope of the best-fit line will give you \lambda. The stock market as a whole has \beta_{\mathrm{Mkt}} = 1 and an average excess return of \mathbb{E}[\mathrm{Mkt}]{-}\mathrm{rf} \approx 4\%. The riskfree bond has \beta_{\mathrm{rf}} = 0 and an average excess return of \mathrm{rf}{-}\mathrm{rf} \approx 0\%. Two points define the slope of a straight line. So textbook theory predicts that \lambda = \mathbb{E}[\mathrm{Mkt}]{-}\mathrm{rf} \approx 4\%.

The estimated slope is far below the equity risk premium. Way back in 1972, Black-Jensen-Scholes ran the test on every NYSE stock from 1926 to 1966, grouped into 10 beta-sorted portfolios. Average excess returns do line up with betas. But the fitted line is too flat, \hat{\lambda} \ll 4\%. Low-beta portfolios earn more than the model predicts, and high-beta portfolios earn less. The problem has only gotten worse. In 1992, Fama-French found basically no relation between average returns and betas from 1963 to 1990. Frazzini-Pedersen (2014) document the same issue in both US and global equities.

Demand Elasticity

The demand-system approach to asset pricing takes the same CARA-normal model but solves each investor’s problem before imposing market clearing. Let i = 1, \ldots, I index individual investors, each with his own risk-aversion coefficient, \gamma_i. Repeat the steps that led to the pricing rule in Equation (6), but stop short of market clearing. Isolating investor i‘s demand on the left-hand side, you get the formula below

(12)   \begin{equation*}\mathrm{Q}_i \;=\; \frac{\mathbb{E}[\mathrm{Payout}] - (1 {+} \mathrm{rf}) \cdot \mathrm{Price}}{\gamma_i \cdot \sigma^2}\end{equation*}

If you hold investor i‘s curve fixed and move the price, then the investor’s demand elasticity is given by

(13)   \begin{equation*}\nu_i = - \frac{\partial \log \mathrm{Q}_i}{\partial \log \mathrm{Price}} = \frac{(1 + \mathrm{rf}) \cdot \mathrm{Price}}{\gamma_i \cdot \sigma^2 \cdot \mathrm{Q}_i}\end{equation*}

Define the aggregate risk-aversion parameter, \gamma, as the harmonic average of the individual coefficients, \tfrac{1}{\gamma} = \sum_i \tfrac{1}{\gamma_i}. In equilibrium, each investor’s position in the CARA-normal model will be inversely proportional to his risk aversion, \gamma_i \cdot \mathrm{Q}_i = \gamma \cdot \psi. Thus, the denominator in the elasticity formula is the same for everyone, \nu_i = \nu. A single elasticity describes every investor. More risk-tolerant investors will hold bigger positions, but their percentage response is identical. This is analogous to the common \lambda across assets.

Notice that the denominator in the elasticity formula is just the risk discount in the CARA-normal model, \gamma \cdot \sigma^2 \cdot \psi = \mathbb{E}[\mathrm{Payout}] - (1 {+} \mathrm{rf}) \cdot \mathrm{Price}. Replace the denominator in Equation (15) with this expression and divide through by the current price. If the riskfree rate isn’t too large, then you get

(14)   \begin{equation*}\nu \;=\; \frac{(1 + \mathrm{rf}) \cdot \mathrm{Price}}{\mathbb{E}[\mathrm{Payout}] - (1 + \mathrm{rf}) \cdot \mathrm{Price}} \;\approx\; \frac{1}{\mathbb{E}[\mathrm{Ret}] {-} \mathrm{rf}} \end{equation*}

The reward for bearing a unit of stock-market risk is the equity risk premium, \mathbb{E}[\mathrm{Mkt}]{-}\mathrm{rf} \approx 4\%. A 1\% rise in the price will increase the cost of financing a share by 1\%, consuming roughly a quarter of the 4\% margin and causing demand to fall by \frac{1\%}{4\%} = 25\%. In other words, theory predicts that \nu = 25.

Deeper Connection

The slope of the SML, \lambda, is the exchange rate between risk and expected returns. How much higher must a stock’s expected excess return be in order to compensate investors for holding one more unit of exposure to market risk? One number common to every asset. The demand elasticity, \nu, is the exchange rate between flows and prices. How much do investors have to adjust their holding in response to a 1\% change in the price? One number common to every investor. Every assumption about preferences and beliefs reaches returns data only through \lambda, and reaches price-impact data only through \nu.

In one sense, these two parameters are two sides of the same coin. Neither is consistent with the observed equity risk premium, \mathbb{E}[\mathrm{Mkt}]{-}\mathrm{rf} \approx 4\%. This point estimate is enormous compared to the observed slope of the SML, which is basically zero. However, the same 4\% number implies a demand elasticity of 25, far above the value near 0.2 in the data. What’s more, the two predictions pull in opposite directions. Any effort that pushes \mathbb{E}[\mathrm{Mkt}]{-}\mathrm{rf} down to fit the slope of the SML makes the elasticity error worse and vice versa. The too-flat SML and the too-steep demand curve are both manifestations of the same underlying problem.

But there’s also a deeper connection. A flat SML is a trading opportunity. Buy levered positions in low-beta stocks and short high-beta stocks. Frazzini-Pedersen calls this trade “betting against beta”, and it has been profitable for decades. That sort of thing shouldn’t survive. Investors ought to pour capital into the trade, bidding up the prices of low-beta stocks and pushing down the prices of high-beta stocks until the SML steepened back to 4\%. That correction is a demand response to price, which is exactly what \nu measures. When demand barely responds to price, mispricings do not get traded away. They just sit there. So the too-steep demand curve is not merely a second manifestation of the same problem. It offers a reason why the first one never went away. The betting-against-beta alpha is what inelasticity looks like in returns data.

Filed Under: Uncategorized

Gordon Prime

June 23, 2026 by Alex

At the moment, all of academic finance revolves around present-value logic. Every model starts this way. No one feels compelled to justify the choice. Researchers view the positive-NPV rule as the gold standard of financial decision-making. What would it be like to live on a world where everyone else felt the same way?

Life on Gordon Prime

On Gordon Prime, the whole civilization is organized around discounting and present-value logic. The positive-NPV rule is not an abstract idea that CEOs learn about in business school. It is what everyone does every day, as a matter of course. M&A press releases lead with the NPV surplus the deal creates. Gordian CEOs can quote every major project’s risk-adjusted discount rate and explain how they arrived at the number. NPV calculations are front and center in conference calls and Investor Day presentations.

The citizens of Gordon Prime are not clairvoyant. They cannot sum infinite series in their heads. They need computers to estimate multi-factor models. To cope, Gordian CEOs have developed many helpful shortcuts for picking reasonable discount rates. Like the rule of 72, but for choosing the right \mathrm{r}. Moreover, on Gordon Prime, CEOs routinely make long-term cash-flow forecasts. A CEO who discounts at \mathrm{r} \approx 5\% can predict cash flows \big(\frac{1}{5\%}\big) = 20 years in the future. For projects with lower discount rates, \mathrm{r} \approx 2\%, it is not uncommon to see horizons of \big( \frac{1}{2\%}\big) = 50 years or more.

Corporations on Gordon Prime announce quarterly cash flows. Income statements exist, but CEOs and shareholders treat them as a mere accounting convention. Earnings get reconciled to cash flows, never the other way around. Because everyone instinctively knows to set “price equal to expected discounted payoff”, Gordian children learn the two Modigliani-Miller irrelevance theorems in kindergarten. Many find the results obvious. Investors on Gordon Prime expect a firm’s leverage and dividend policy to be dictated by frictions. Absent such complications, nobody would think to ask a Gordian CEO about either.

Asset pricing on Gordon Prime is a solved science. It has to be. Everyone needs this input to perform the right NPV calculation. John Cochrane has a counterpart on this distant planet. There, he gave an AFA presidential address titled “Discount Rates: We Know How To Calculate Them.” A Gordian CEO expects her decisions to move her firm’s multiple right away. To choose a corporate policy she has to understand how the market will price the resulting change in her firm’s future payout stream.

Research on Gordon Prime

Now put yourself in the shoes of a Gordian corporate-finance researcher. What would your day look like? You certainly would never survey CEOs about whether they use the positive-NPV rule. Of course they do. CEOs on Gordon Prime won’t shut up about it. They talk about the positive-NPV rule incessantly. On Gordon Prime, the question would be as strange as asking whether CEOs perform arithmetic. Instead, you might run surveys asking CEOs how they pick the right discount rate for specific kinds of unusual projects.

Discount rates are a property of a project’s future cash flows, not a characteristic of the firm. So, on Gordon Prime, it does not matter which company is evaluating a project. A CEO who greenlights the investment at one company would make an identical call while running another. Thus, much of corporate-finance research on Gordon Prime would be done at the project level.

Gordian CEOs make decisions by discounting cash flows that are expected to arrive decades in the future. Current interest rates and market conditions play a minor role. The Gordon Prime version of IBES contains 20- and 30-year cash-flow forecasts for most firms because this is the horizon that matters when thinking in present-value terms. The positive-NPV rule compares the present value of a project’s cash flows against the upfront cost, so the funding source is treated as a minor detail. What matters is a project’s total cost, not where the company gets the money from.

Prestige on Gordon Prime would go to researchers who study how corporate investment moves over time in subtle ways. Gordian researchers treat capital structure and dividend policy as second-tier topics, which only matter due to frictions. When these frictions bite, the dose-response curve is the same for every firm. The documented effects are smooth and continuous. Constrained CEOs invest a little less than otherwise-similar unconstrained ones.

A strange little model

One slow afternoon, you (dear Gordian researcher) write down a strange little model where the CEO does not discount anything. Instead, she makes decisions aimed at increasing her firm’s short-term EPS forecast—i.e., expected earnings over the next twelve months divided by current shares outstanding.

What would follow?

Quite a lot! The EPS-maximizing CEO in your model does not use the positive-NPV rule. Instead of converting the project’s future cash-flow stream into an upfront valuation, she translates the upfront cost into an expense flow. Your EPS maximizer only invests in projects that generate enough income next year to cover their own added financing expense using the firm’s cheapest available source of capital. In other words, the CEO funds projects with income yields, \mathrm{IY} = \frac{\mathbb{E}[\Delta \mathrm{NOI}_1]}{\mathrm{Cost}}, that exceed her firm’s cheapest financing yield, \mathrm{FY} = \min\{\mathrm{EY}, \mathrm{i}, \mathrm{rf}\}. This is the accretive investment rule.

One thing that immediately jumps out at you is the nature of this hurdle rate. It is a property of the firm, not the project. Equity would be cheapest for a firm whose earnings yield is below the riskfree rate, \mathrm{EY} < \mathrm{rf}. Such a firm would fund investments by issuing stock even while sitting on cash. By contrast, a firm with \mathrm{EY} > \mathrm{rf} would lever up, making cash by far the cheapest funding source if/when it ever appears. The line between the two types would fall at \mathrm{EY} = \mathrm{rf}, and it would move around as the riskfree rate moved. The EPS-maximizing CEO in your model takes her firm’s current pricing as given, so her decisions track with market conditions rather than her beliefs about cash flows decades in the future.

It is a strange little model. No discounting. No present value. But sharp predictions, so you keep going.

Life on planet Accreton

To better understand your max EPS theory, you run a thought experiment. You imagine a crazy planet where accretive decision-making is the law of the land. You call it, “Accreton”. It must be a wild place, indeed. What would life on this hypothetical planet look like?

On Accreton, press releases would lead with EPS accretion. The deal’s NPV surplus would often go unmentioned, or even (**GASP**) uncalculated. EPS-maximizing CEOs would use rules of thumb, such as IRRs and payback periods, that avoid choosing a project-specific discount rate altogether. Accretonese companies would obsess over short-term earnings, not cash flows. Since the star of the show is next year’s EPS, no one on planet Accreton would bother making 20-year cash-flow forecasts. Maybe companies would talk about cash flows a bit… but only to reconcile why they differ from net income.

Unlike on Gordon Prime, leverage and payout policy would be first-order concerns on planet Accreton. CEOs would labor over these choices. Analysts would ask hard questions to make sure the values were correct. An Accretonese CEO would not need to know the correct asset-pricing model. It would not matter to her where the company’s PE ratio and marginal interest rate came from. All she would need to know is what the current values are. On planet Accreton, a CEO could take these numbers as given and evaluate corporate policies from there.

On planet Accreton, firms on either side of \mathrm{EY} = \mathrm{rf} would pursue two different constellations of corporate policies. A growth stock (\mathrm{EY} < \mathrm{rf}) would see equity as cheap compared to riskfree debt. A value stock (\mathrm{EY} > \mathrm{rf}) would see equity as more expensive. The two kinds of firms could disagree about whether to fund the same project. If the Federal Reserve on planet Accreton cut rates, then the value stock might flip its verdict. Accretonese growth stocks fund investments by issuing equity even when cash is available. By contrast, value stocks would find it accretive to lever up until cash becomes their cheapest available source.

Accretonese research

This thought experiment has been fun so far. You decide to keep pushing. What would life be like for a corporate-finance researcher on planet Accreton? Golly gee willikers, you think to yourself as a proud Gordian citizen, the literature on that other planet sure must be different.

For one thing, you figure that no researcher on planet Accreton would bother running surveys asking whether CEOs use the positive-NPV rule. What would be the point? Every corporate statement on planet Accreton leads with EPS growth. The relevant metric there is obviously not NPV.

You imagine that Accretonese researchers would focus on yield spreads, not discount rates. When a CEO on that planet decides whether to fund a project, what matters is the income-vs-financing yield spread, \mathrm{IY} {-} \mathrm{FY}. The discovery of the growth-versus-value divide at \mathrm{EY} = \mathrm{rf} would be one of the foundational discoveries in the literature. On planet Accreton, capital structure and payout would not be sleepy backwaters. They would have a seat at the big boys’ table, right next to real investment. All three would be seen as ways for a CEO to generate value for her shareholders by increasing the firm’s EPS next year.

A firm’s PE ratio is just another way of writing its earnings yield, \mathrm{PE} = \big( \frac{1}{\mathrm{EY}} \big). You figure that Accretonese researchers would view IRRs and payback periods in a similar light. An IRR is a multi-period generalization of a project’s income yield, \mathrm{IY}. A payback period is the same quantity expressed as a multiple, \big( \frac{1}{\mathrm{IY}}\big). If there’s a database like IBES on planet Accreton, then you have to imagine that it only contains 1- and 2-year EPS forecasts. Why would anyone there bother to forecast cash flows two decades into the future?

Stranger than fiction

Here is the crazy thing. That planet you dreamed up… (Q: You mean the one where accretive decision-making is the law of the land and most people never discount a cent?) Yes, that one. That planet is a pixel-perfect description of Earth. Every line of it. You did not invent Accreton. You described home.

Yet, corporate-finance researchers act as though they are living on Gordon Prime. They spend their days puzzled by the data that keeps arriving. Discount rates that barely move when the cost of capital does. Investment that lurches with cash flow it should not care about. CEOs quote payback periods even though textbooks call the method “stupid”. Researchers cannot tell CEOs which cost of capital to use because they cannot agree themselves. The literature contains a zoo of different factor models. None of it should be puzzling. It is Accretonese data read by researchers who insist they are somewhere else.

A sensible Gordian researcher would never run a survey asking whether CEOs use the positive-NPV rule or an IRR hurdle. On Gordon Prime the answer is plain, they discount, and everyone can see it. A sensible Accretonese researcher would never run that survey either. Here the answer is just as plain, they go by accretion, and every press release says so. The question only occurs to someone who cannot tell which planet she is standing on. It is the question of a lost interstellar traveler.

What regressions show

You can’t convince researchers that they live on Accreton by pitting the accretive rule against the positive-NPV rule in a horse race to see which one fits the data better. That contest is silly. The first Modigliani-Miller paper was published in 1958. Researchers have spent nearly 70 years adding ingredients to the same present-value framework: financing constraints, agency costs, behavioral biases, adjustment costs, etc. Collectively, all that machinery can be fine-tuned to fit almost any pattern. If we live on Accreton, then any explanatory power associated with the positive-NPV rule must come from either overfitting or these after-market add-ons. A comparison of predictive accuracy cannot separate those two stories.

That is the value of this thought experiment. It shows where the diagnostic evidence lives. It is not in the regression R^2. It is in what CEOs say when they announce an M&A deal, in which numbers lead the press release, in how rarely CEOs mention discount rates. Only 1% of conference calls quote a discount rate. The fact that we have to ask CEOs whether they use the positive-NPV rule doesn’t prove we live on planet Accreton. But it sure is hard to square with the claim that Earth is Gordon Prime.

Filed Under: Uncategorized

Why max EPS Persists: Ba (AER 2026) Interpretation

March 29, 2026 by Alex

The question

Researchers currently take it for granted that CEOs should maximize PV[Shareholder Payouts]. Suppose we agree. Under this premise, max EPS is the wrong objective. It’s a misspecified model for how to run a firm. In work with Itzhak Ben-David, I show that EPS maximization solves the 3 core problems in corporate finance: capital structure, real investment, and payout policy.

How does max EPS survive? CEOs have access to textbooks, MBA programs, consultants, and analysts—all of which teach present-value logic. The data generated by decades of corporate decisions are available for everyone to examine. If max EPS is misspecified, why hasn’t it been abandoned?

Ba (2026) provides a formal theory of exactly this phenomenon: when and why misspecified models persist, even when decision-makers are open to switching and have access to infinite data. This note spells out the connection to EPS maximization.

Ba’s framework

An agent uses a subjective model \theta to guide decisions. Each period t, she chooses an action a_t from a finite set \mathcal{A} and observes an outcome y_t drawn from the true (unknown) data-generating process Q^*(\cdot | a_t). Her model \theta is a parametric family of predicted DGPs, \{Q^\theta(\cdot | a, \omega)\}_{a \in \mathcal{A}, \omega \in \Omega^\theta}, where \omega indexes the parameter space \Omega^\theta. The model is correctly specified if some \omega recovers Q^*; it is misspecified otherwise.

The agent holds a prior \pi_0^\theta over \Omega^\theta and updates beliefs via Bayes’ rule within the model. Crucially, she also considers a competing model \theta' with its own parameter space \Omega^{\theta'} and prior \pi_0^{\theta'}. She compares models using the Bayes factor

(1)   \begin{equation*} \lambda_t = \frac{\ell_t(\theta')}{\ell_t(\theta)} \end{equation*}

where \ell_t(\theta) = \sum_{\omega \in \Omega^\theta} \pi_0^\theta(\omega) \ell_t(\theta, \omega) is the marginal likelihood of the data under model \theta, and \ell_t(\theta, \omega) = \prod_{\tau=0}^{t} q^\theta(y_\tau | a_\tau, \omega) is the likelihood conditional on parameter \omega.

As new data rolls in, the agent updates her Bayes factor recursively

(2)   \begin{equation*}  \lambda_t = \lambda_{t-1} \cdot \frac{\sum_{\omega' \in \Omega^{\theta'}} \pi_t^{\theta'}(\omega') \, q^{\theta'}(y_t | a_t, \omega')}{\sum_{\omega \in \Omega^\theta} \pi_t^\theta(\omega) \, q^\theta(y_t | a_t, \omega)} \end{equation*}

If \lambda_t > \alpha where \alpha \geq 1 is a switching threshold, the agent switches to \theta'. If \lambda_t < 1/\alpha, she switches back. The threshold \alpha controls switching stickiness. A larger \alpha requires stronger evidence to switch.

Ba (2026) notation EPS vs. PV application
Initial model \theta Max EPS
Competing model \theta' Max PV[Shareholder Payouts]
Action set \mathcal{A} Corporate decisions: leverage choice, project selection, payout policy
Outcome y_t Observable corporate outcomes: EPS level, EPS growth, stock-price reaction, analyst response
True DGP Q^*(\cdot \mid a) The actual mapping from corporate decisions to outcomes (determined by the full economic environment)
Parameter space \Omega^\theta Parameters of the EPS model (earnings yield, interest rates)
Parameter space \Omega^{\theta'} Parameters of the PV model (discount rates, growth rates, terminal values, risk premia, payout schedules)
Switching threshold \alpha Institutional friction: retraining costs, compensation redesign, regulatory reporting norms, board inertia
Bayes factor \lambda_t Cumulative evidence that PV logic fits the data better than EPS logic

Result 1: Endogenous data lets misspecified models survive forever

The theorem

Theorem 1 (Ba 2026, p. 16). Suppose \alpha > 1. The following are equivalent:

  1. Model \theta is globally robust for at least one full-support prior.
  2. Model \theta is locally robust for at least one full-support prior.
  3. There exists a p-absorbing self-confirming equilibrium (SCE) under model \theta.

A self-confirming equilibrium under \theta is a strategy \sigma supported by a belief \pi^\theta such that (i) \sigma is myopically optimal given \pi^\theta, and (ii) the model’s prediction matches the true DGP on the equilibrium path

(3)   \begin{equation*} q^\theta(\cdot \mid a, \omega) \equiv q^*(\cdot \mid a) \qquad \forall \, a \in \text{supp}(\sigma), \; \forall \, \omega \in \text{supp}(\pi^\theta) \end{equation*}

The strategy is p-absorbing if a dogmatic \theta-modeler eventually plays only actions in \text{supp}(\sigma).

The key insight: a misspecified model need not be globally correct. It only needs to be correct on the equilibrium path—for the actions it actually induces. Off-path misspecification is never revealed because the agent’s own actions determine which data are generated. This is why endogenous data is essential: Ba (2026, p. 16, fn. 19) notes that “in an exogenous-data environment, Theorem 1 implies that the sufficient and necessary condition for both local robustness and global robustness is that the model is correctly specified.”

P-absorbingness adds a dynamic requirement on top of the static SCE condition. It is not enough for an SCE to exist; the agent’s belief dynamics must actually converge to it. Ba’s Section 5.2 (Proposition 2, p. 24) shows that this convergence property depends on the direction of belief reinforcement. When beliefs and actions are complements (so that the bias feeds on itself), the dynamics are positively reinforcing and convergence to the SCE is guaranteed. Hence, the SCE is p-absorbing. When beliefs and actions are substitutes, the bias is self-correcting. Dynamics may oscillate and fail to converge. A SCE exists but is not p-absorbing, and the misspecified model is not robust.

Application to EPS

When a CEO maximizes EPS, her decisions shape the observable corporate outcomes. The EPS model’s predictions are tested only against data generated by EPS-driven actions. Predictions about actions a CEO never takes are never tested.

Leverage. An EPS maximizer borrows when \mathrm{EY} > \mathrm{i} (earnings yield exceeds the interest rate) and uses the proceeds to retire shares. This raises EPS mechanically. The outcome the CEO observes is: EPS went up, the stock price did not collapse, analysts applauded the “accretive” transaction. The EPS model’s prediction that the transaction would be good because it is accretive is confirmed by the data the decision itself generated. The CEO does not observe the counterfactual: what would’ve happened under the PV-optimal leverage choice given frictions.

Investment. An EPS maximizer uses \mathrm{HR} = \min\{\mathrm{EY}, \, \mathrm{i}, \, \mathrm{rf}\} as the hurdle rate, not the WACC. She rejects projects with positive NPV but negative first-year EPS impact (dilutive projects) and accepts projects with negative NPV but positive first-year EPS impact (accretive projects). The observed outcome: EPS did not fall, the project looks like it “worked.” The NPV of rejected projects is never observed.

Payout. An EPS maximizer buys back stock whenever buybacks offer a higher yield than investing cash (\mathrm{EY} > \mathrm{CY}) rather than evaluating the NPV of the buyback. The observed outcome: EPS went up, the market reacted positively to the announcement. The PV counterfactual (could the cash have been better deployed elsewhere?) is off-path.

Let \sigma^{EPS} be the strategy induced by max EPS, and let \hat{\omega} be a parameter value in the EPS model under which the predicted outcome distribution matches Q^*(\cdot | a) for all a \in \text{supp}(\sigma^{\mathrm{EPS}}). Then \sigma^{\mathrm{EPS}} is an SCE under the EPS model. This is plausible because the EPS model does not mispredict the direction of EPS changes from leverage, buybacks, or accretive acquisitions. It correctly predicts that borrowing at \mathrm{i} < \mathrm{EY} raises EPS, that using cash for buybacks at \mathrm{EY} > \mathrm{CY} raise EPS, and so on. What it gets wrong is the welfare interpretation: whether these EPS changes correspond to value creation. But welfare is not directly observed in y_t; what is observed are EPS changes, stock-price reactions, and analyst ratings, all of which are consistent with the EPS model’s on-path predictions.

Moreover, the EPS model’s feedback dynamics are positively reinforcing in the sense of Ba’s Proposition 2: EPS-driven decisions raise EPS, which validates the model, which strengthens conviction, which leads to more EPS-driven decisions. This positive feedback ensures that the SCE is p-absorbing. Contrast this with the dynamics facing a CEO who switches to PV logic: she accepts a dilutive acquisition, EPS falls in the short run, analysts downgrade, the stock price drops, and the PV model appears to have failed, creating pressure to revert. The transition to the correct model generates short-run data that seem to disconfirm it. By Theorem 1, the existence of a p-absorbing SCE under the EPS model is sufficient for it to be globally robust. Hence, EPS maximization can persist against any competitor, including PV logic, with infinite data.

Result 2: Concise models can be more robust than correct ones

The theorem

Theorem 2 (Ba 2026, p. 19). Suppose \alpha > 1 and model \theta has no traps. Then:

  1. Model \theta is globally robust at prior \pi_0^\theta if and only if \pi_0^\theta(C^\theta) \geq 1/\alpha.
  2. Model \theta is locally robust at all full-support priors if and only if C^\theta \neq \emptyset.

C^\theta is the set of consistent parameters: those \omega for which the pure belief \delta_\omega supports a p-absorbing SCE. The model’s prediction under \omega matches the true DGP at every action in the equilibrium strategy’s support.

The condition \pi_0^\theta(C^\theta) \geq 1/\alpha links three forces: the model’s structure (which determines C^\theta), the agent’s prior (which determines how much mass falls on C^\theta), and the switching threshold (which sets the bar). Prior tightness and switching stickiness are substitutes: a higher \alpha lowers the bar for prior concentration, and a tighter prior lowers the bar for stickiness. Any asymptotically accurate model can be globally robust at a given prior, provided switching is sufficiently sticky.

The critical implication: correctly specified models are not necessarily more robust than misspecified ones. A misspecified model with a smaller parameter space |\Omega^\theta| can satisfy the tightness condition more easily. Under an ignorance prior (uniform over \Omega^\theta), each parameter receives weight 1/|\Omega^\theta|. For a model where every parameter is consistent (C^\theta = \Omega^\theta), the tightness condition is automatically satisfied at any \alpha > 1, regardless of the prior—the model is unconditionally globally robust. But a correctly specified model with a large parameter space needs a correspondingly tight prior to be globally robust, and under a uniform prior it may fail. Ba (2026, p. 4): “simple misspecified models equipped with entrenched priors can be more robust than complex correctly specified models.”

In the media-bias application (Section 5.1, Proposition 1), Ba makes this concrete: a two-state misspecified model \hat{\theta} is globally robust at all priors and all \alpha \geq 1, while the correctly specified three-state model \theta is globally robust only if \pi_0^\theta(\omega^M) \geq 1/\alpha. The misspecified model permanently replaces the correct one with arbitrarily high probability as the prior on the extreme states increases.

Application to EPS

The EPS model has a small parameter space. For any given decision (borrow or not, invest or not, buy back or not), it requires the CEO to know essentially two things: the earnings yield \mathrm{EY} = \frac{\mathbb{E}[\mathrm{EPS}]}{\mathrm{Price}} and the relevant financing cost (interest rate \mathrm{i} or risk-free rate \mathrm{rf}). The decision rule is a direct comparison: act if and only if \mathrm{EY} > \mathrm{HR}, where \mathrm{HR} = \min\{\mathrm{EY}, \, \mathrm{i}, \, \mathrm{rf}\}.

The PV model requires knowledge of a much larger parameter space \Omega^{\theta'}: the risk-free rate, market risk premium, firm beta (or multi-factor betas), the project-specific risk adjustment, the terminal growth rate, the expected path of future cash flows, and the probability distribution over states of the world.

Under Theorem 2, the prior tightness condition for global robustness is \pi_0^\theta(C^\theta) \geq 1/\alpha. For the EPS model, if C^\theta encompasses most or all of \Omega^\theta (because the model is consistent for the small set of parameters it uses), then \pi_0^\theta(C^\theta) is close to 1 and the condition is satisfied for any \alpha > 1. The EPS model may be unconditionally globally robust.

For the PV model, even though it is correctly specified (C^{\theta'} \neq \emptyset), the prior mass is spread across a large parameter space. Under a diffuse prior, \pi_0^{\theta'}(C^{\theta'}) may be small. The PV model is locally robust at all priors (by Theorem 2(ii)), but it is globally robust only if \pi_0^{\theta'}(C^{\theta'}) \geq 1/\alpha. With large |\Omega^{\theta'}| and diffuse prior, this can fail.

Moreover, switching stickiness in corporate settings is very large. Compensation contracts are tied to EPS targets. Analyst coverage is organized around EPS estimates, consensus forecasts, and PE multiples. Regulatory reporting (GAAP earnings) makes EPS the most salient and auditable metric, while PV calculations involve subjective inputs (discount rates, growth assumptions) that are harder to audit and verify. Board education is required to shift from a direct comparison (“is this accretive?”) to a multi-parameter model (“what is the NPV at the appropriate risk-adjusted discount rate?”). All of this amounts to a very high \alpha, which further lowers the bar for the prior tightness that the EPS model must satisfy.

Result 3: Even slight switching friction is enough

The theorem

Theorem 3 (Ba 2026, p. 21). Suppose model \theta has no traps and \alpha = 1. Then model \theta is locally or globally robust at any full-support prior \pi_0^\theta if and only if C^\theta = \Omega^\theta.

When switching is non-sticky (\alpha = 1), local and global robustness coincide, robustness at some prior is equivalent to robustness at all priors, and both hold only when every parameter in the model is consistent. This is an extreme demand: the model must be correct for every DGP it entertains, not just on the equilibrium path. Only a model with full prior tightness (C^\theta = \Omega^\theta) can survive frictionless comparison.

The set of robust models shrinks discontinuously at \alpha = 1. For any \alpha > 1, models with C^\theta \neq \Omega^\theta can be robust (provided the prior tightness condition is met). At \alpha = 1, they cannot. Ba (2026, p. 21): “the set of locally robust models and supporting priors shrinks discontinuously at \alpha = 1, which highlights how stickiness helps more misspecified models persist.”

The mechanism: at \alpha = 1, there always exists a nearby competing model that fits the data slightly better than the initial model on some dimension. Because there is no switching friction, this marginal improvement is sufficient to trigger a switch. The proof constructs such a competing model by preserving most DGPs in \theta while slightly improving the accuracy of one DGP associated with a parameter in \Omega^\theta \setminus C^\theta.

Application to EPS

Theorem 3 clarifies that the persistence of max EPS depends on switching friction being positive—but the required friction can be arbitrarily small. For any \alpha > 1 (even \alpha = 1.01), the EPS model can be globally robust provided the prior tightness condition is met. The discontinuity at \alpha = 1 means that even minimal institutional friction (a small cost of retraining, a slight reluctance to abandon a familiar framework) is qualitatively different from zero friction.

This matters because it addresses a potential objection: “surely CEOs could switch to PV logic if they wanted to; there’s no real barrier.” Ba’s result says that even a negligible barrier is enough, as long as it is positive. The EPS model does not need an enormous moat to survive. It needs (i) a p-absorbing SCE (Result 1), (ii) sufficient prior concentration on consistent parameters (Result 2), and (iii) any positive switching friction at all (Result 3). The first two conditions are structural properties of the EPS model. The third is almost trivially satisfied in any real institution.

Conversely, Theorem 3 identifies the knife-edge case where EPS would be displaced: a world with literally zero switching costs (\alpha = 1) and a PV model that is a slight local improvement. In practice, this would correspond to an environment where CEOs face no career risk from short-term EPS misses, no analyst pressure around quarterly earnings, and no cognitive cost of estimating multi-parameter discount rates. These conditions do not describe any real-world setting.

Summary

Ba (2026) provides a formal framework for understanding when misspecified models persist despite competition from correctly specified alternatives. Applied to the EPS-vs.-PV question, the theory identifies three reinforcing mechanisms, one per main result.

Result Mechanism Application to EPS
Theorem 1 Misspecified model is robust iff it admits p-absorbing SCE; endogenous data insulates on-path predictions from off-path errors Accretive actions generate data that confirm max EPS model; positive feedback dynamics ensure convergence to SCE
Theorem 2 Global robustness requires \pi_0^\theta(C^\theta) \!\geq\! 1/\alpha; concise models concentrate priors; stickiness and prior tightness substitutes EPS model’s small parameter space makes tightness condition easy to satisfy; minor frictions and reporting norms matter
Theorem 3 Set of robust models shrinks discontinuously at \alpha = 1; any positive friction qualitatively expands what can persist Even minimal switching costs suffice for EPS to survive; knife-edge \alpha = 1 case (zero friction) does not describe real world

The punchline: even if we take as given that maximizing PV[Shareholder Payouts] is the correct objective, the Ba (2026) framework gives formal reasons—grounded in Bayesian learning theory—for why maximizing EPS can persist indefinitely. The misspecified model is not merely sticky due to inertia or ignorance. It is robust in a precise sense: it admits a self-confirming equilibrium, its directness concentrates prior beliefs, the data it generates through the CEO’s own actions provide continuous apparent validation, and even minimal institutional friction is enough to protect it. These forces can be strong enough that the correct model is permanently abandoned.

Caveat: Is max PV[Shareholder Payouts] actually the correct model?

Everything above assumes that maximizing PV[Shareholder Payouts] is the correctly specified model. It’s the true DGP against which max EPS is judged misspecified. Ba’s framework requires us to designate one model as correct and ask whether the other persists. We chose PV as the correct model because that is what finance theory prescribes. But this assumption deserves scrutiny.

Shareholders do not get to spend corporate earnings. Earnings are an accounting construct; they accrue to the firm, not to the shareholder’s bank account. A dollar of EPS that is retained and reinvested never reaches the shareholder at all. In that sense, maximizing EPS is maximizing a fiction—a number that does not correspond to any cash flow the shareholder actually receives.

But PV[Shareholder Payouts] has a parallel problem. Shareholders do not get to spend the present discounted value of a dollar they expect to receive in 20 years. You cannot eat risk-adjusted returns. The “present value” of a distant payout is a mathematical object, not cash in hand. And yet these distant, heavily discounted payouts are the primary drivers of valuation in the PV framework.

The Gordon growth model makes this concrete. Under the standard formulation, an asset’s price equals expected cash flow next year times a forward-looking multiple

(4)   \begin{equation*} \mathrm{Price} = \mathbb{E}[\mathrm{CF}] \times \bigg( \frac{1}{\mathrm{r} - \mathrm{g}} \bigg) \end{equation*}

For typical parameter values (\mathrm{r} \approx 10\%, \mathrm{g} \approx 5\%), the multiple is roughly \big(\frac{1}{10\%-5\%}\big) = 20\times. But \big(\frac{1}{\mathrm{r}-\mathrm{g}}\big) is also the Macaulay duration of the cash flow stream in years. So the “typical” dollar of present value corresponds to a payout roughly two decades in the future. The PV framework asks the CEO to make decisions today based on the risk-adjusted value of money that shareholders will not receive for 20 years. This is money that never appears on a financial statement, whose value depends on estimates of discount rates and growth rates that are themselves deeply uncertain.

This does not mean PV logic is wrong. It means that both models involve abstractions, and the question of which abstraction is “correct” is less obvious than textbook finance suggests. EPS is a fiction because earnings are not payouts. PV is a fiction because present values are not cash. The Ba (2026) framework shows that even if we grant the PV model the status of correct specification, the EPS model can persist indefinitely. But if we take seriously the possibility that neither model is unambiguously correctly specified, then the persistence of max EPS becomes even less surprising. In Ba’s terms, we may not be in a world where a correctly specified competitor exists at all, in which case the question is not whether EPS will be abandoned but which misspecified model proves more robust. Regardless, history has shown that EPS wins.

Filed Under: Uncategorized

Next Page »

Pages

  • Publications
  • Working Papers
  • Curriculum Vitae
  • Notebook
  • Courses

Copyright © 2026 · eleven40 Pro Theme on Genesis Framework · WordPress · Log in