Research Notebook

How Analysts Do DCF

September 19, 2026 by Alex

Public companies are not obligated to return all available cash to shareholders at the end of the year. When you buy a share of stock, you receive the portion of the firm’s free cash flow that it decides to distribute as dividends. Textbook finance theory assumes that shareholders focus on these future payoffs when pricing a stock. Asset-pricing theory all stems from one simple concept: price equals expected discounted payoff. This is a foundational principle behind every model.

The “cash flows” in the classic discounted cash-flow (DCF) model refer to the money paid to investors, not the money received by the company. The Gordon model says to value a stock by scaling up next year’s per-share dividend forecast to capture the value of all subsequent dividend payments

(1)   \begin{equation*}\text{Gordon Valuation} = \sum_{h=1}^{\infty} \frac{\mathbb{F}[\mathrm{DPS}_h]}{(1{+}r)^h} = \mathbb{F}[\mathrm{DPS}_1] \times \bigg( \frac{1}{r - g} \bigg) \end{equation*}

\mathbb{F}[\mathrm{DPS}_h] is the forecasted dividend per share in h years, and r is the discount rate applied to these future payoffs, which are predicted to grow the same amount every year in perpetuity, \mathbb{F}[\mathrm{DPS}_h] = (1{+}g)^h \times \mathrm{DPS}.

This may be how every academic model works, but it is not how sell-side analysts typically value a company. Most reports set a price target by capitalizing the firm’s predicted earnings at a reasonable recent multiple. The default valuation method is next year’s EPS (earnings per share) forecast times a trailing P/E. Instead of asking, “How much would the firm’s entire future dividend stream be worth in today’s dollars?”, analysts ask, “If the company were to announce next year’s predicted EPS today, how would it be valued given current market pricing?” Only around 30\% of sell-side reports pay lip service to DCF analysis.

Roughly 1 in 10 reports includes a full-blown DCF model. This post examines how analysts actually implement DCF analysis in these reports. How do they arrive at a per-share equity valuation for a particular company? It turns out that the procedure looks very different from what most researchers have in mind. Analysts are not just plugging 3 key numbers into Equation (1). The standard DCF playbook is connected to the Gordon model, but it isn’t approximately Gordon in a way that’s captured by time-varying parameters. The protocol deviates from textbook theory in fundamental ways. My point is not that analysts are wrong. My point is that they’re doing something much more interesting than most researchers currently appreciate.

Standard Playbook

Analysts do not use DCF to compute a company’s share price directly. Instead, the approach gets used to estimate the enterprise value (EV) of the entire firm as a whole. Analysts forecast the company’s free cash flow (FCFF) for the next few years and discount these anticipated cash flows at a weighted average cost of capital (WACC). Then they tack on a terminal value (TV), reflecting the present discounted value of all subsequent free cash flows. The combined expression looks something like this

(2)   \begin{equation*}\mathrm{EV} = \sum_{h=1}^{H} \frac{\mathbb{F}[\mathrm{FCFF}_h]}{(1 {+} r)^h} + \frac{\mathrm{TV}}{(1 {+} r)^H} \end{equation*}

\mathbb{F}[\mathrm{FCFF}_h] denotes the firm’s forecasted FCFF h years from now. H \geq 0 denotes the length of the firm’s transitionary period. During this time, the company’s FCFF is assumed to grow at a faster rate, settling down at a lower sustainable growth rate from year (H{+}1) onward.

After arriving at an enterprise value, the analyst then uses a bridge to convert the number into a per-share equity valuation. First, they deduct the book value of net debt. Then, they divide the result by the company’s current share count. The fair value is thus

(3)   \begin{equation*}\text{DCF Valuation} = \frac{\mathrm{EV} - \text{Net Debt}}{\mathrm{Shares}} \end{equation*}

\text{Net Debt} = \mathrm{Debt} {-} \mathrm{Cash} is the difference between a company’s outstanding debt and its cash holdings.

Running Example

To fix ideas, I’ll be using a running example involving a company in the midst of a transformation. Over the next 4 years, the company’s FCFF is expected to grow quickly: \mathbb{F}[\mathrm{FCFF}_1] = \mathdollar 8\mathrm{M}, \mathbb{F}[\mathrm{FCFF}_2] = \mathdollar 12\mathrm{M}, \mathbb{F}[\mathrm{FCFF}_3] = \mathdollar 16\mathrm{M}, and \mathbb{F}[\mathrm{FCFF}_4] = \mathdollar 20\mathrm{M}. These forecasts give the company a 4-year compound annual growth rate (CAGR) of 26\%. From year 5 onward, the firm’s FCFF will grow at a more modest rate, g = 3\%. The firm has 10\mathrm{M} shares outstanding, and each share is priced at \mathdollar 20/\mathrm{sh}, giving the firm a market cap of \mathdollar 200\mathrm{M}. Assume a cost of equity of 10\%.

Yesterday, the company issued a 3-year coupon bond with a face value of \mathdollar 100\mathrm{M} and a 4\% annual coupon rate. The bond was issued at par. Unfortunately for the firm, the yield on this bond jumped to r=5\% this morning, causing the price to drop by roughly (5\%{-}4\%) \cdot \mathdollar 100\mathrm{M} \cdot 3\text{ years} = \mathdollar 3\mathrm{M} to \mathdollar 97\mathrm{M}. Since the bond was issued at par, the firm’s balance sheet carries it at \mathdollar 100\mathrm{M} because GAAP does not mark debt to market. The firm holds \mathdollar 20\mathrm{M} in cash. Net debt at book is \mathdollar 80\mathrm{M}. Net debt at market would be \mathdollar 97\mathrm{M}{-}\mathdollar 20\mathrm{M}=\mathdollar 77\mathrm{M}. The firm faces a 20\% effective tax rate.

The Discount Rate

Analysts usually pick a discount rate equal to the firm’s weighted average cost of capital (WACC). It is the average tax-adjusted return on the firm’s capital

(4)   \begin{equation*}r = \bigg(\frac{\mathrm{Equity}}{\mathrm{Equity} + \mathrm{Debt}}\bigg) \times r_E + \underbrace{\bigg(\frac{\mathrm{Debt}}{\mathrm{Equity}+\mathrm{Debt}}\bigg)}_{\text{leverage ratio}} \times (1 {-} \tau) \cdot r_D\end{equation*}

Note that the \mathrm{Equity} and \mathrm{Debt} numbers in the WACC formula denote current market pricing, not book value at issuance. An unlevered firm will have a WACC equal to its cost of equity, r_E. A highly leveraged firm will get discounted at a rate closer to its tax-adjusted cost of debt, (1{-}\tau)\cdot r_D, which will be lower.

Plugging in numbers, the firm in our running example would face an annual discount rate of

(5)   \begin{equation*}\underbrace{\bigg(\frac{\mathdollar 200\mathrm{M}}{\mathdollar 200\mathrm{M} + \mathdollar 97\mathrm{M}}\bigg)}_{{\sim}2/3} \times 10\% + \underbrace{\bigg(\frac{\mathdollar 97\mathrm{M}}{\mathdollar 200\mathrm{M}+\mathdollar 97\mathrm{M}}\bigg)}_{{\sim}1/3} \times \underbrace{\!\!\!\!\phantom{\bigg(}(1 {-} 0.2) \cdot 5\%\phantom{\bigg)}\!\!\!\!}_{4\%} \approx 8\%\end{equation*}

The company’s {\sim}1/3 leverage ratio closes 2\%\mathrm{pt} of the gap between r_E=10\% and (1 {-} \tau) \cdot r_D=4\%. Debt gets discounted at 4\% rather than r_D=5\% because interest payments don’t face a 20\% tax.

The Terminal Value

During the next couple of years, the firm is growing quickly. The terminal-value term in Equation (2), \frac{\mathrm{TV}}{(1{+}r)^H}, reflects the present value of the company’s subsequent FCFF stream after this transitionary period has ended. It is usually computed with either a perpetuity or with a trailing/comps multiple

(6)   \begin{equation*}\mathrm{TV} = \mathbb{F}[\mathrm{FCFF}_{H+1}] \times \begin{cases} \big( \tfrac{1}{r - g}\big) &\text{Gordon} \\ \mathrm{EV/FCFF} &\text{Comps} \end{cases}\end{equation*}

The first option assumes that the firm’s FCFF will grow at an annual rate of g from year (H{+}1) onward. The second option uses the firm’s current EV/FCFF multiple or the average multiple for a set of comparable firms to capitalize the company’s anticipated FCF in year (H{+}1).

In our running example, the firm has a 4-year transitionary period in which its FCFF grows at 26\% per year. The analyst believes that its FCFF will grow from \mathbb{F}[\mathrm{FCFF}_1]=\mathdollar 8\mathrm{M} next year to \mathbb{F}[\mathrm{FCFF}_4]=\mathdollar 20\mathrm{M} in year H=4. After that, the company’s FCFF stream will grow at g=3\% per year in perpetuity. This lower more-sustainable growth rate puts the firm’s forecast for year (H{+}1)=5 at

(7)   \begin{equation*}\mathbb{F}[\mathrm{FCFF}_5] = (1{+}3\%)\times \mathdollar 20\mathrm{M} = \mathdollar 20.6\mathrm{M}\end{equation*}

The firm’s 8\% WACC and 3\% growth rate give it a cap rate of (8\%{-}3\%)=5\%. At a terminal multiple of \big( \frac{1}{8\%-3\%} \big) = 20{\times}, the company’s terminal value would be

(8)   \begin{equation*}\mathdollar 20.6\mathrm{M} \times \underbrace{\bigg( \frac{1}{8\%{-}3\%} \bigg)}_{1/5\%=20} = \mathdollar 412\mathrm{M}\end{equation*}

Discounting this terminal value back 4 years at r=8\% gives \frac{\mathdollar 412\mathrm{M}}{(1{+}8\%)^4} \approx \mathdollar 303\mathrm{M}.

All Together Now

Here’s how all the pieces come together in the running example. The table below shows the company’s predicted FCFF stream over the next 4 years as well as its terminal value

    \begin{equation*}\begin{array}{r|ccccc} \text{Horizon, } h & 1 & 2 & 3 & 4 & \mathrm{TV} \\ \hline \mathbb{F}[\mathrm{FCFF}_h] & \mathdollar 8.0\mathrm{M} & \mathdollar 12.0\mathrm{M} & \mathdollar 16.0\mathrm{M} & \mathdollar 20.0\mathrm{M} & \mathdollar 412.0\mathrm{M} \\ \frac{\mathbb{F}[\mathrm{FCFF}_h]}{(1+8\%)^h} & \mathdollar 7.4\mathrm{M} & \mathdollar 10.3\mathrm{M} & \mathdollar 12.7\mathrm{M} & \mathdollar 14.7\mathrm{M} & \mathdollar 302.8\mathrm{M} \end{array}\end{equation*}

Combined, the firm’s first 4 years of FCFF are worth \mathdollar 7.4\mathrm{M} {+} \mathdollar 10.3\mathrm{M} {+} \mathdollar 12.7\mathrm{M} {+} \mathdollar 14.7\mathrm{M} \approx \mathdollar 45\mathrm{M} in today’s dollars. After discounting, its terminal value is worth \mathdollar 303\mathrm{M}. Together, these two components produce a total enterprise value of \mathdollar 45\mathrm{M} {+} \mathdollar 303\mathrm{M} \approx \mathdollar 348\mathrm{M}, with roughly \frac{\mathdollar 303\mathrm{M}}{\mathdollar 348\mathrm{M}} \approx 87\% coming at the end.

An analyst would now need to convert this \mathdollar 348\mathrm{M} enterprise value into a per-share equity valuation. They do this using a bridge. The first step involves deducting the book value of the firm’s net debt, \mathdollar 100\mathrm{M}{-}\mathdollar 20\mathrm{M} = \mathdollar 80\mathrm{M}. The resulting difference is meant to capture the combined value of shareholders’ equity stake. To get a per-share number, the analyst then divides by the company’s current 10\mathrm{M} share count

(9)   \begin{equation*}\frac{\mathdollar 348\mathrm{M} - \mathdollar 80\mathrm{M}}{10\mathrm{M}} = \mathdollar 26.80/\mathrm{sh}\end{equation*}

The company in our running example is trading at \mathdollar 20/\mathrm{sh}, so this implies \frac{\mathdollar 26.80/\mathrm{sh}-\mathdollar 20/\mathrm{sh}}{\mathdollar 20/\mathrm{sh}} \approx 34\% upside.

FCFF, Not Payoffs

Now that we understand what analysts do, let’s talk about the ways that it differs from what researchers assume. First and foremost, the standard DCF implementation capitalizes the cash the firm could pay rather than the cash the firm will pay. The resulting share price doesn’t reflect the discounted payoff stream that investors expect to receive. Analysts use a DCF model to capitalize money received by the firm, not money paid to investors. Every textbook model takes it for granted that investor payoffs and corporate proceeds are the same thing. If that’s the case, then the way that analysts implement DCF analysis isn’t inconsistent with standard theory.

The decision to focus on FCFF is particularly noteworthy. The line item represents the maximum dividend that an unlevered firm could distribute to its shareholders, not the dividend payment shareholders actually receive. The firm in our running example will produce \mathdollar 8\mathrm{M} of FCFF next year. The company has promised \mathdollar 4\mathrm{M} to bondholders. After adjusting for taxes, this leaves \mathdollar 8\mathrm{M}{-}(1{-}0.2)\cdot\mathdollar 4\mathrm{M}=\mathdollar 4.8\mathrm{M}. The firm can use this money to distribute a dividend, buy back shares, pay down its existing debt, acquire another firm, or add to its cash balance. In principle, shareholders might not receive any of these funds. For Equations (1) and (3) to be equivalent, a firm would have to invest any money not paid to *current stakeholders* at exactly 8\%.

Every report that builds a DCF model could have calculated the total payout to all stakeholders. The enterprise-level analog to dividend payments would be

(10)   \begin{align*}\overbrace{\mathrm{FCFF}  - (1 {-} \tau) \cdot \mathrm{Interest} - \Delta \mathrm{Cash}}^{\text{money available to distribute}} = \overbrace{\mathrm{Dividends} + \mathrm{Buybacks}}^{\text{to current shareholders}} + \overbrace{({-}\Delta \text{Debt})}^{\text{to creditors}} + \overbrace{\mathrm{Acquisitions}}^{\substack{\text{to outside} \\ \text{shareholders}}}\end{align*}

All you need to do is take FCFF and subtract off the tax adjusted interest payment and any increase in cash. The resulting amount of money must have been paid to current shareholders (dividends or buybacks), the firm’s creditors (debt paydown), or outside shareholders (acquisitions). Nobody does this.

Transition Period

Roughly 10\% of sell-side reports include a full-fledged DCF model. Almost all of these reports include a transition period with higher FCFF growth. When researchers think about a DCF model, they usually picture some version of the Gordon pricing rule in Equation (1). But the vast majority of DCF models that analysts write down do not assume sustainable long-term growth right away. 9 out of 10 reports that include a detailed DCF analysis set H > 0. The analyst assumes the existence of an initial transition period in which the firm will grow at a much faster rate than g.

In the running example, the transition lasts H=4 years. Think about what would happen if an analyst tried to remove this part of the model. First, suppose the analyst tried to set H=0 without changing any other inputs. In this case, the resulting enterprise value would be far too low. The company is only predicted to generate \mathdollar 8\mathrm{M} of FCFF next year. So, with r=8\% and g=3\%, the Gordon-implied valuation would be

(11)   \begin{equation*}\mathdollar 8\mathrm{M} \times \bigg( \frac{1}{8\% - 3\%} \bigg) = \mathdollar 160\mathrm{M}\end{equation*}

After subtracting off net debt of \mathdollar 80\mathrm{M} and dividing by 10\mathrm{M} shares, the analyst would arrive at a share price of \mathdollar 8.00/\mathrm{sh}. The analyst can’t publish such a low number. The firm is currently trading at \mathdollar 20/\mathrm{sh}.

Maybe the analyst can fix the problem by discounting at a lower discount rate? Nope. This would require discounting FCFF at the firm’s cost of debt, r_D=5\%. To get an enterprise value of \mathdollar 348\mathrm{M}, you’d need

(12)   \begin{equation*}\underbrace{\bigg(\frac{\mathdollar 8\mathrm{M}}{\mathdollar 348\mathrm{M}}\bigg)}_{2.3\%} + \;3\% \approx 5.3\%\end{equation*}

The implied multiple would also be absurdly high, \big( \frac{1}{5.3\% - 3\%} \big) \approx 43{\times}. Those are laughable numbers.

Thus, the requirement that H=0 puts analysts between a rock and a hard place. Allowing for H > 0 provides a way out. A short-term runup makes it possible to quote a valuation of \mathdollar 26.80/\mathrm{sh} without violating professional norms. The analyst gets to use a reasonable sounding 8\% WACC. They can point to a sensible \big( \frac{1}{8\% - 3\%} \big) = 20{\times} multiple. All they have to do is argue that, before settling onto a more stable trajectory, the firm will initially grow at a faster 26\%/\mathrm{yr} clip for the next H=4 years.

Identification Problem

These transition periods create a serious identification problem for researchers looking to back out an analyst’s r discount rate from IBES data. Those numbers contain an analyst’s forecast for next year (ignore the FCFF-vs-EPS distinction), their predicted long-term growth rate, and a price target. The choice of r that justifies this price target depends on the analyst’s choice of transition dynamics. How long will the firm’s transition period last, H? How fast will the firm grow during this time? Neither variable shows up in IBES.

To illustrate the severity of the problem, think about what the firm from the running example would look like in IBES. A researcher would see the analyst’s forecast for next year, \mathbb{F}[\mathrm{FCFF}_1] = \mathdollar 8\mathrm{M}, and the analyst’s long-term growth rate, g = 3\%. IBES might also contain a price target of \mathdollar 26.80/\mathrm{sh}. What IBES doesn’t tell you is the length of the analyst’s assumed transition period, H=4, or the assumed CAGR during this period, 26\%. The table below shows how the implied r changes as you adjust H, holding fixed all observable quantities and setting \mathrm{CAGR}=26\%

    \begin{equation*}\begin{array}{r|rrrrr} H & 0 & 2 & 4 & 6 & 8 \\ \hline \mathbb{F}[\mathrm{FCFF}_{H+1}] & \mathdollar 8.0\mathrm{M} & \mathdollar 12.7\mathrm{M} & \mathdollar 20.6\mathrm{M} & \mathdollar 32.7\mathrm{M} & \mathdollar 51.9\mathrm{M} \\ \text{Implied }r & 5.3\% & 6.4\% & 8.0\% & 9.9\% & 11.9\% \end{array}\end{equation*}

Without knowing how long the analyst assumed that the firm would grow 26\%, a researcher cannot tell whether r=8.0\%, 5.3\%, or 11.9\%. All these numbers are consistent with a \mathdollar 26.80/\mathrm{sh} valuation.

Required Precision

DCF analysis comes with a natural measuring stick for the size of the errors: the present value of the firm’s FCFF forecast for next year. In the running example, this is \frac{\mathbb{F}[\mathrm{FCFF}_1]}{1+r} = \frac{\mathdollar 8\mathrm{M}}{1+8\%} = \mathdollar 7.4\mathrm{M}. If all you’re doing is adding up expected discounted cash flows, then it’d be a problem if you skipped or double-counted one. An error of \mathdollar 7.4\mathrm{M} in enterprise value is the same as forgetting to count next year’s FCFF or including the number twice. Either way, a decision that alters enterprise value by \frac{\mathdollar 7.4\mathrm{M}}{\mathdollar 348\mathrm{M}} \approx 2\% is a big deal.

If the firm’s enterprise value fell to \mathdollar 348\mathrm{M} {-} \mathdollar 7.4\mathrm{M} = \mathdollar 340.6\mathrm{M}, the company’s share price would fall by \mathdollar 0.75/\mathrm{sh} to

(13)   \begin{equation*}\frac{\mathdollar 340.6\mathrm{M} {-} \mathdollar 80\mathrm{M}}{10\mathrm{M}} = \mathdollar 26.05/\mathrm{sh}\end{equation*}

That is a change of roughly \frac{\mathdollar 0.75/\mathrm{sh}}{\mathdollar 26.80/\mathrm{sh}} \approx 3\%. So, anything that moves the company’s share price by more than {\pm}3\% is tantamount to forgetting to count a year of FCFF or double counting that year.

The terminal growth rate, g, describes how a firm’s FCFF will grow in steady state. Forever is a long time. And, with so many years to worry about, small differences in an analyst’s choice of g can compound into huge changes in the firm’s terminal value. The effect of a {\pm}1\%\mathrm{pt} change in the long-run growth rate will scale with the firm’s Gordon multiple

(14)   \begin{equation*}\frac{\Delta \mathrm{TV}}{\mathrm{TV}} \approx \bigg(\frac{1}{r - g}\bigg) \times \Delta g\end{equation*}

The precision with which analysts quote g tells you that they aren’t thinking along these lines. Most quote nice round numbers near 2\% or 3\%.

The firm in our running example has a cap rate of r{-}g = 8\%{-}3\%= 5\%\mathrm{pt}, which implies that a {+}1\%\mathrm{pt} increase in the terminal growth rate will cause the firm’s terminal value to rise by

(15)   \begin{equation*}\bigg( \frac{1}{8\%-3\%} \bigg) \times 1\%\mathrm{pt} = {+}20\%\end{equation*}

Apply that to the \mathdollar 302.8\mathrm{M} present value of the terminal value in the running example. \Delta g = 0.12\%\mathrm{pt} would move this number by roughly \mathdollar 7.4\mathrm{M}, an amount equal to the present value of next year’s anticipated FCFF. Yet, most reports quote g in steps of 0.25\%\mathrm{pt}, an increment half as precise as required by theory.

Market vs Book Value

The firm’s debt shows up in two different places in Equation (3): a) the WACC formula for the discount rate used to compute the firm’s EV; and b) the bridge used to translate this number into a share price. In the first case, the debt is taken at market value. In the second case, the book value of debt is used. For the firm in the running example, the bridge subtracts \mathdollar 100\mathrm{M} for a bond the firm could buy back for \mathdollar 97\mathrm{M}. That \mathdollar 3\mathrm{M} gap understates the firm’s equity value by \mathdollar 3\mathrm{M}, resulting in a share price that is \mathdollar 0.30/\mathrm{sh} too low. The WACC is right. The bridge is wrong. Suppose that the firm’s credit rating deteriorates further so that its \mathdollar 100\mathrm{M} bond trades at \mathdollar 92.5\mathrm{M}, not \mathdollar 97\mathrm{M}. Market net debt would now be \mathdollar 72.5\mathrm{M}, but the bridge would still sit at \mathdollar 80\mathrm{M}. The firm’s equity value would now be understated by \mathdollar 7.5\mathrm{M}, which amounts to \mathdollar 0.75/\mathrm{sh}.

The analyst above calculated his WACC assuming a leverage of \frac{\mathdollar 97\mathrm{M}}{\mathdollar 200\mathrm{M}+\mathdollar 97\mathrm{M}} \approx 33\%. The \mathdollar 200\mathrm{M} equity value in the denominator comes from the fact that the company’s shares currently trade at \mathdollar 20/\mathrm{sh} and there are 10\mathrm{M} shares in circulation. However, this choice of WACC puts the value of each share at \mathdollar 26.80/\mathrm{sh}. At this price point, the firm’s equity value would be \mathdollar 268\mathrm{M}, not \mathdollar 200\mathrm{M}, and the firm’s leverage would be 27\%, not 33\%. The internally consistent calculation would produce a higher WACC, 8.3\%, because there is less debt and less shield. Iterating to this fixed point gives an enterprise value of \mathdollar 327\mathrm{M} and per-share valuation of \mathdollar 24.70/\mathrm{sh}. Nobody does that. Literally. No one. Analysts just ignore this fixed-point problem. The result is a rate that is \Delta r = {-}0.3\%\mathrm{pt} too low and a share price that is overvalued by \frac{\mathdollar 24.70/\mathrm{sh}-\mathdollar 26.80/\mathrm{sh}}{\mathdollar 26.80/\mathrm{sh}} = {-}8\%.

Why does a 0.3\%\mathrm{pt} move in the rate produce an 8\% move in the price? Start with a company that has no runup, just a perpetuity growing at g. Its price is \mathrm{FCFF} times \big(\frac{1}{r - g}\big), so a change in the rate moves prices by

(16)   \begin{equation*}-\bigg(\frac{1}{r - g}\bigg) \times \Delta r = -20 \times 0.30\% = -6\%\end{equation*}

The 20{\times} company in the running example sees its enterprise value fall by 6\% in response to a 0.30\%\mathrm{pt} increase in the rate. The firm’s enterprise value falls by \mathdollar 21\mathrm{M}, from \mathdollar 348\mathrm{M} to \mathdollar 327\mathrm{M}.

The extra 2\%\mathrm{pt} of decline comes from the bridge. Net debt is a fixed \mathdollar 80\mathrm{M}, so the entire {-}\mathdollar 21\mathrm{M} drop in enterprise value is borne by shareholders

(17)   \begin{equation*}{-}\bigg(\frac{\mathdollar 21\mathrm{M}}{\mathdollar 348\mathrm{M} - \mathdollar 80\mathrm{M}}\bigg) = {-}\frac{\mathdollar 21\mathrm{M}}{\mathdollar 268\mathrm{M}} \approx -8\%\end{equation*}

Since equity only constitutes \frac{\mathdollar 268\mathrm{M}}{\mathdollar 348\mathrm{M}} \approx 77\% of enterprise value, the bridge turns the {-}6\% decline in enterprise value into a \big( \frac{1}{77\%} \big) \times ({-}6\%) \approx {-}8\% drop in the share price.

Tax shield of interest

The firm in our running example has an 8\% WACC when including the tax shield of interest. Without the factor of (1{-}\tau), the company’s discount rate would rise to

(18)   \begin{equation*}\bigg(\frac{\mathdollar 200\mathrm{M}}{\mathdollar 200\mathrm{M} + \mathdollar 97\mathrm{M}}\bigg) \times 10\% + \bigg(\frac{\mathdollar 97\mathrm{M}}{\mathdollar 200\mathrm{M}+\mathdollar 97\mathrm{M}}\bigg) \times 5\% \approx 8.4\%\end{equation*}

At this higher rate, the firm’s anticipated FCFF path would be worth just \mathdollar 323\mathrm{M}. In other words, the firm’s enterprise value contains a tax shield worth around \mathdollar 348\mathrm{M}{-}\mathdollar 323\mathrm{M}=\mathdollar 25\mathrm{M}.

The full \mathdollar 25\mathrm{M} does not appear on the firm’s financial statements. Its balance sheet shows a single \mathdollar 100\mathrm{M} bond that will mature in year 3. The shield on that bond is \tau \times \mathrm{Coupon} = 0.2 \times \{4\% \cdot \mathdollar 100\mathrm{M}\} = \mathdollar 0.8\mathrm{M} a year for three years. This \mathdollar 25\mathrm{M}{-}\mathdollar 2.4\mathrm{M} \approx \mathdollar 22.6\mathrm{M} gap is roughly 3{\times} larger than the precision cutoff

(19)   \begin{equation*}\frac{\mathdollar 22.6\mathrm{M}}{\mathdollar 348\mathrm{M}} \approx 6.5\%\end{equation*}

The difference will be even larger for newly distressed firms with high costs of debt. Discounting FCFF at a WACC is only exact when the firm continually rebalances its debt to a fixed share of value.

The Point Of DCF

Each section above describes a discrepancy. Each one vanishes under certain special conditions. Together, these conditions describe a particular kind of firm. For such a company, none of the issues outlined above would matter. Here’s what that company would look like:

  1. The firm must be investment grade. The company’s debt should trade at par. Its bonds ought to be issued recently or carry a floating rate. The firm’s credit rating should be AAA and never budge. Under these conditions, it won’t matter whether you use net debt at book or market value. Both are \mathdollar 100\mathrm{M}.
  2. The firm’s tax shield must be correctly valued. Its debt-to-EV ratio must remain constant. The company must refinance every bond at maturity. And it must have steady taxable income, so that every dollar of interest gets deducted. Under those conditions, the firm will collect the full \mathdollar 25\mathrm{M} shield implied by discounting FCFF at WACC.
  3. The analyst’s valuation must be close to the current market price. Otherwise the leverage ratio used to compute the WACC will not match the leverage ratio implied by the analyst’s own valuation, and the discount rate will be inconsistent with the number it produced. In practice, this means DCF confirms the market rather than disputes it.
  4. The firm must pay out nearly all available free cash flow each year or reinvest the money at the analyst’s WACC. If retained cash and acquisitions are either \mathdollar 0 or NPV neutral, then discounting expected future FCFF is the same thing as discounting expected future payouts. In practice, only large dividend payers meet this condition. Growth-stock initiations never do.
  5. The firm must have a clearly defined runup period, which everyone can agree upon. A massive production facility will come online in 3 years. A new drug will get approved in 5 years. It will take 2 years for a merger to close. The company’s favorable current lease agreement will expire in 10 years. The interim path is a schedule rather than a story, and the terminal year is the first normal year. Very few reports tie the runup to a dated event. The ones that do are hotel chains with an opening schedule that halves at stated dates, a device maker with an approval year, and Tesla with 2013 Model S volumes. Everywhere else the runup is a fade toward g with nothing scheduled.
  6. The firm’s terminal growth rate must be a credible forecast. The firm’s FCFF stream should grow with the economy after completing the runup phase. g at nominal GDP growth must be a credible belief.

These statements describe a mature, investment-grade, fully-distributing firm going through one well-defined transition with a known end date. Think about a regulated utility completing a rate-base build. An oil-pipeline company with a contracted expansion coming online in a year or two. A stable industrial company that is midway through integrating a big recent acquisition.

Now notice what the DCF adds for such a firm. Once the transition ends, the firm trades like its peers. So the terminal multiple is a number the analyst could have looked up. \big( \frac{1}{8\% - 3\%} \big) = 20{\times} and a peer EV/FCFF of 20{\times} are the same number doing the same job. Only one of them can be checked against a screen. That is why analysts so often close with a comp multiple rather than a Gordon multiple. The comp close is the analyst admitting that r and g were never the point.

The point of writing down a DCF model has nothing to do with the core logic behind the Gordon pricing rule in Equation (1). The only thing DCF analysis adds for this kind of firm is the runup. It answers one question: How long until this firm will look like its peers again? For a firm with a real transition, that answer is worth having. For a firm with no transition, there is no point to performing the DCF exercise. You might as well capitalize next year’s FCFF forecast at a reasonable recent multiple and be done with it.

Filed Under: Uncategorized

Excessively Volatile? Or Inexplicably Precise?

August 11, 2026 by Alex

The dividend discount model (DDM) says that a stock’s current price ought to reflect the discounted value of its expected future dividend stream

(1)   \begin{equation*}\mathrm{Price} = \sum_{t=1}^{\infty} \frac{\mathbb{E}[\mathrm{Div}_{t}]}{(1{+}r)^t}\end{equation*}

\mathbb{E}[\mathrm{Div}_t] is the company’s expected dividend in t years, and r > 0\% is the firm’s discount rate.

The Gordon model is a special case in which dividends are assumed to grow at a constant rate, \mathbb{E}[\mathrm{Div}_t] = (1{+}g)^t \cdot \mathrm{Div}_0 for all t \geq 1. Under this assumption, the DDM’s infinite sum reduces to

(2)   \begin{equation*}\mathrm{Price} = \mathbb{E}[\mathrm{Div}_1] \times \bigg( \frac{1}{r-g} \bigg)\end{equation*}

A stock’s price is higher when its next-twelve-month (NTM) dividend forecast is higher (large \mathbb{E}[\mathrm{Div}_1]), when investors don’t discount its future dividend stream very heavily (small r), and when the firm’s expected dividend growth offsets more of the deleterious effects of discounting (large g).

When researchers write down these sorts of models, they typically assume that the relevant parameters are known to all. Shareholders have a good dividend forecast in mind, \mathbb{E}[\mathrm{Div}_1]. They use the right discount rate, r, and hold accurate beliefs about the firm’s long-run growth rate, g.

However, in practice, someone who wanted to use the Gordon model to price a stock would have to estimate all three quantities. This post walks through a simple exercise. Imagine that the price of the the SPDR S&P 500 ETF Trust (SPY) reflects Gordon logic, and investors are able to estimate its cap rate with the same precision that bond traders are able to predict Treasury rates. This is a heroically optimistic assumption. Yet, I show that it would still only pin down SPY’s price to within {\pm}50\%. The excess volatility puzzle should be viewed as an excess precision puzzle. SPY’s return fluctuates by {\pm}20\% from year to year. If you think the price reflects Gordon logic, then how are equity investors keeping things so stable?

Where to spend your energy

SPY is currently trading at \mathrm{Price} = \mathdollar 770/\mathrm{sh}. Investors expect SPY to pay a dividend of \mathbb{E}[\mathrm{Div}_1] = \mathdollar 7.70/\mathrm{sh} over the next year, giving it a dividend yield of \mathrm{DY} = \frac{\mathdollar 7.70/\mathrm{sh}}{\mathdollar 770/\mathrm{sh}} = 1\%. The Gordon model says the index’s dividend yield reflects the difference between its annual discount rate and its expected dividend growth rate, \mathrm{DY} = (r{-}g) = 1\%. SPY’s price level comes from capitalizing its \mathdollar 7.70/\mathrm{sh} dividend forecast at a price-to-dividend multiple of \mathrm{PD} = \big( \frac{1}{1\%} \big) = 100{\times}

(3)   \begin{equation*}\mathrm{Price} = \mathbb{E}[\mathrm{Div}_1] \times \bigg( \frac{1}{r-g} \bigg) = \mathdollar 7.70/\mathrm{sh} \times 100 = \mathdollar 770/\mathrm{sh}\end{equation*}

Gordon offers two ways to shift SPY’s current price level: change its NTM dividend forecast, or change the index’s cap rate. The second channel is way more impactful. To see why, imagine that news comes out that raises SPY’s short-term dividend forecast by {\sim}1\%, from \mathbb{E}[\mathrm{Div}_1] = \mathdollar 7.70/\mathrm{sh} to \mathdollar 7.78/\mathrm{sh}. If SPY’s multiple remains the same, the Gordon model predicts that its price will rise by 1\% as well

(4)   \begin{equation*}\frac{\mathrm{d}\mathrm{Price}}{\mathrm{Price}} = \frac{\mathrm{d}\mathbb{E}[\mathrm{Div}_1]}{\mathbb{E}[\mathrm{Div}_1]}\end{equation*}

A {+}\mathdollar 0.08/\mathrm{sh} increase in SPY’s dividend forecast will lead to a 100 \times \mathdollar 0.08/\mathrm{sh} \approx {+}\mathdollar 8.00/\mathrm{sh} price pop. This is nothing to sneeze at, but SPY’s dividend is fairly stable. Dividend-growth volatility is in the low single digits.

By contrast, when using the Gordon model to value SPY, it is absolutely critical to plug in the right cap rate. The model says that errors in (r{-}g) get magnified by a factor of \mathrm{PD} = 100{\times}

(5)   \begin{equation*}\frac{\mathrm{d}\mathrm{Price}}{\mathrm{Price}} =  -\,\mathrm{PD} \cdot \mathrm{d}(r{-}g)\end{equation*}

Suppose you thought the appropriate cap rate for SPY was 1.1\% rather than 1.0\%. This \mathrm{d}(r{-}g) = {+}10\mathrm{bp} error would cause you to undervalue the index by 100 \times 0.1\% = 10\%. The Gordon-implied price would go from \mathdollar 770/\mathrm{sh} to \mathdollar 7.70/\mathrm{sh} \times \big( \frac{1}{1.1\%} \big) = \mathdollar 700/\mathrm{sh}.

A 1\% increase in SPY’s one-year-ahead dividend forecast would cause its share price to rise by \mathdollar 8.00/\mathrm{sh}. A 10\mathrm{bp} increase in SPY’s cap rate would cause its share price to plummet by \mathdollar 70/\mathrm{sh}. These two channels differ in strength by two orders of magnitude. This is not a coincidence. SPY trades at 100\times its forward dividend. If you want to get SPY’s price level correct using the Gordon model, then you should put almost all your effort into estimating its right cap rate. The question is: how precisely can investors estimate this quantity? To an accuracy of {\pm}100\mathrm{bp}? To within {\pm}10\mathrm{bp}? What’s the tightest plausible error bound?

Treasury forward prices

To answer this question, let’s pivot from talking about SPY to talking about Treasuries. This is the market where rates get estimated most precisely. There are two things about this market which make it especially convenient to estimate a bond’s discount rate. First, there exists an active forward-contract market. A bond’s forward price can be computed from today’s bond price using a no-arbitrage argument. If a dealer quotes a forward price that deviates from this no-arbitrage value, there is a riskless way to make money from the gap.

Consider a 2-year bond that costs \mathrm{Price} = \mathdollar 96 today. This bond promises to pay a coupon of \mathrm{C}=\mathdollar 3 in each of the next two years and then return its face value of \mathrm{FV} = \mathdollar 100 at maturity. The bond trades at a discount to its face value. The going one-year interest rate is r = 5\%. A forward price is the price you agree to today for buying this bond next year after its first coupon has been paid. No one has to guess this price. It can be manufactured. Borrow \mathdollar 96 today and buy the bond. This portfolio would cost you nothing since \mathrm{Price} = \mathdollar 96. One year from now, you would then owe \mathrm{Price} \times (1{+}r) = \mathdollar 96 \times (1{+}5\%) = \mathdollar 100.80 on the short position. But you would also own the bond and be in possession of an extra \mathdollar 3 after collecting the first coupon. Thus, your break-even sale price would be \mathrm{Price} \times (1{+}r) - \mathrm{C} = \mathdollar 100.80 - \mathdollar 3 = \mathdollar 97.80. This is the bond’s one-year-ahead forward price, \mathrm{Fwd}.

The resulting forward price is the break-even resale price. Suppose that the bond actually winds up trading at \mathrm{Fwd} = \mathdollar 97.80 next year. In that case, someone who paid \mathrm{Price} = \mathdollar 96 for the bond today would earn a return equal to the going one-year interest rate

(6)   \begin{align*}\frac{(\mathrm{C} {+} \mathrm{Fwd}) - \mathrm{Price}}{\mathrm{Price}} &= \frac{\mathrm{C}}{\mathrm{Price}} + \frac{\mathrm{Fwd}{-} \mathrm{Price}}{\mathrm{Price}} \\ &= \,\,\frac{\mathdollar 3}{\mathdollar 96}\,\; + \frac{\mathdollar 97.80 {-} \mathdollar 96}{\mathdollar 96} = \, 5\%\end{align*}

If the forward price turns out to match the realized future price on the nose, then the total payout from owning the bond next year would be \mathdollar 100.80. Collect the \mathdollar 3 coupon and sell the bond for \mathdollar 97.80. The \mathdollar 3 coupon amounts to a 3.13\% yield. The \mathdollar 1.80 price increase contributes an extra 1.87\%. The two components sum to deliver r = 3.13\% + 1.87\% = 5\%.

The one-year-ahead forward price is not someone’s idle musings. It is a price forecast that bond traders arrive at with money at stake. A trader who’s convinced that the bond will sell for more than \mathrm{Fwd} = \mathdollar 97.80 next year can buy the forward and wait. A trader who believes that the future price will be lower can do the opposite. Every disagreement is an order, and orders move the quote. There’s no counterpart for SPY. Sure, analysts regularly set one-year-ahead price targets. But different analysts publish different numbers, and there’s no way to arbitrage the discrepancy in their views. You can’t buy or sell an SPY price target.

The bond market’s accuracy

A bond’s forward price is the market’s working prediction of the bond’s price a year in the future, \mathrm{Fwd}_t \approx \mathbb{E}_t[\mathrm{Price}_{t+1}]. Nothing riskless holds the realized price to this prediction. A trader convinced that \mathrm{Price}_{t+1} will come in above \mathrm{Fwd}_t can buy the forward and wait, but waiting entails risk. Rates can move against him before the year is out. So the gap between the forward price and the realized price is a bet, not an arbitrage. The market cannot squeeze it to zero. It can only keep it small. How small? The prediction could be amazingly accurate, or it could be incredibly noisy. To find out, we just need to compare the traded forward price at time t to the bond’s price level a year later. A simple first pass might look at

(7)   \begin{equation*}\sqrt{\frac{1}{T} \cdot \sum_{t=1}^T \, \bigg(\;\frac{\mathrm{Price}_{t+1} - \mathrm{Fwd}_t}{\mathrm{Fwd}_t}\,\bigg)^{\!\!2}}\end{equation*}

If the output is 1\%, then it’d indicate that the realized bond price a year from now is usually {\pm}1\% away from bond traders’ best guess today.

Here’s where the second feature of Treasury markets comes in. If we had run this exercise in equity markets, then there could be two reasons why next year’s price might have shifted: change in next year’s dividend forecast or change in the cap rate. As discussed above, the second channel is more impactful. But the first channel still exists when looking at equities. By contrast, it is completely absent in the Treasury market where a bond’s cash flows are known at the time of purchase. One year from now, our bond will be a 1-year bond with a single remaining payment of \mathrm{C} + \mathrm{FV} = \mathdollar 103. Its price will be that \mathdollar 103 discounted at whatever the 1-year rate turns out to be. Every ingredient of \mathrm{Price}_{t+1} except the rate is fixed in advance. So when the realized price misses the forward, there is exactly one suspect: the discount rate moved. And the conversion is one-for-one at this horizon, since the remaining claim has a duration of one year: a \mathdollar 0.50 price miss is a {\sim}50\mathrm{bps} rate miss. Treasury forward-price errors are estimates of the noise in r, in exactly the units we need, with nothing else mixed in.

When you look at real-world data, how noisy is the bond market’s best guess? Fed-fund futures are forwards written on the overnight rate a few months out. Betting against them has earned {\sim}50\mathrm{bp} per year on average from 1988 to 2003 (Piazzesi and Swanson, 2008). At longer maturities, the errors arrive in price units, so it’s necessary to adjust the formula above to account for duration. A bond’s price miss is roughly its duration times its rate miss. Two-year notes miss by about {\sim}2\% a year, which implies the rate is off by {\pm}100\mathrm{bp}. Ten-year notes miss by {\sim}6\%. With a duration of 8 years, this pricing error translates to rate mistakes of {\pm}75\mathrm{bp}. The long bond misses by {\sim}10\%, which translates to rate noise of {\pm}65\mathrm{bp} when assuming a duration of 18 years. Every point on the curve tells the same story. The bond market’s best forecast of a rate is accurate to within 50\mathrm{bp} to 100\mathrm{bp}.

Equity traders would be thrilled if they could pin down SPY’s cap rate to the same level of precision. Treasuries generate the most accurate estimates for an asset’s discount rate. It is a number the assembled market backed by real money, and anyone holding a better estimate could have traded it into the quote. The errors are pure, because known cash flows leave the rate as the only moving part. Whatever noise survives under these conditions is noise that no investor, anywhere, could have forecasted away.

What it means for SPY

SPY’s cap rate is the same kind of object as the yield on a Treasury bond. It is the rate that turns a stream of future payments into today’s price. Noise in (r{-}g) means the index’s price is wandering off its present-value path, exactly as a bond price wanders off its forward path when the rate moves. The bond market’s best rate forecasts miss by somewhere between 50\mathrm{bp} and 100\mathrm{bp} in a typical year. Take this result seriously and it becomes a precision limit for equity pricing.

The Gordon model says to capitalize SPY’s \mathdollar 7.70/\mathrm{sh} NTM dividend forecast using a 100{\times} multiple. Above, we saw that the same factor of 100{\times} also magnifies cap-rate errors when calculating the percentage price impact. If equity investors cannot hope to estimate SPY’s cap rate to an accuracy better than {\pm}50\mathrm{bp}, then the Gordon price can only be accurate to within

(8)   \begin{equation*}100 \times 0.5\% = {\pm}50\%\end{equation*}

When using (r{-}g)=1\%, the Gordon model says SPY should trade at \mathdollar 7.70/\mathrm{sh} \times \big( \frac{1}{1\%} \big) = \mathdollar 770/\mathrm{sh}. If the correct cap rate could be as high as 1.5\% or as low as 0.5\%, then the true valuation could be anywhere from \mathdollar 513/\mathrm{sh} to \mathdollar 1{,}540/\mathrm{sh}, a range of more than \mathdollar 1{,}000.

This is Fisher Black’s quip about how prices are “correct” to “within a factor of 2”. When using the canonical present-value model, the best achievable estimate of SPY’s cap rate leaves its price level undetermined to within 100 \times 0.5\% \approx 50\%. This is not a prediction of stock-market volatility. It is the size of the price fluctuations that could be explained by the unavoidable noise in the market’s best estimate of (r{-}g). No one knows SPY’s cap rate to an accuracy of {\pm}50\mathrm{bp}. To really drive this point home, note that while SPY currently has a dividend yield of 1\%, its long-run average dividend yield is closer to 2\%. Many academic papers rely on this higher value for calibrations. This is a disagreement of {+}100\mathrm{bp}.

Excess volatility puzzle

The above calculations put an entirely different spin on Shiller’s classic result. In his 1981 paper, he calculated the price implied by the S&P 500’s realized future dividend stream using a constant annual discount rate. He then compared this DDM-implied valuation to the index’s actual price level at the time. Figure 1 in the paper shows that the two time series are wildly different. The S&P 500’s realized price is 10{\times} more volatile than the implied price. The index’s dividend is extremely stable while its returns fluctuate by {\pm}20\% from year to year.

Shiller (1981) studies a model in which investors (a) set price equal to expected discounted payoff, (b) have perfect foresight about those future payoffs, and (c) use a constant discount rate. This model clearly does not fit the level of the S&P 500. How did researchers respond to this finding? Well, they didn’t abandon assumption (a). Instead, they tried to generate additional return volatility by allowing for biased beliefs and letting the discount rate vary over time. e.g., in the late 1980s, Campbell and Shiller produced a dynamic extension of the Gordon model, which allowed r and g to vary over time. It is widely believed that this log-linear approximation to Gordon holds under arbitrary subjective beliefs.

But if you’re unwilling to abandon present-value logic, then this gets the story exactly backwards. Shiller calculated a DDM-implied price for the S&P 500 in an extremely simple way

(9)   \begin{equation*}\text{Implied Price}_t = \sum_{h=1}^{T-t} \frac{\mathrm{Div}_{t+h}}{(1{+}r)^h} \; + \; \frac{\overline{\text{Price}}}{(1{+}r)^{T-t}}\end{equation*}

T=1979 is the last year in the sample. \overline{\mathrm{Price}} is the S&P 500’s average detrended real price level during the sample period. This is not the theoretically correct thing to do. It parks every dividend payment beyond the sample in a single assumed terminal value. It uses the S&P 500’s realized future dividend payments rather than investors’ expectations of these payoffs. The correct discount rate need not be constant.

While not exactly pristine, suppose you think that Shiller’s implied price is roughly correct. Morally speaking, the formula above is clearly a present-value calculation. If you think the outcome of this formula is in the right ballpark, then the question is not: Why is the market price so volatile? The question is: How are market participants keeping the price level so damn close? Suppose the implied price were constant. In that case, an annual return volatility of 20\% would require knowing the S&P 500’s cap rate to within {\pm}20\mathrm{bp}. This level of precision is far below anything observed even in bond markets.

If the bond market cannot pin down next year’s rate to better than {\pm}50\mathrm{bp}, then how on earth are equity investors pricing the S&P 500 in a way that requires knowledge of (r{-}g) to within {\pm}20\mathrm{bp}? If equity investors know SPY’s cap rate to within {\pm}50\mathrm{bp}, then the correct valuation could be anywhere from \mathdollar 7.70/\mathrm{sh} \times \big( \frac{1}{0.5\%} \big) \approx \mathdollar 1{,}540/\mathrm{sh} to \mathdollar 7.70/\mathrm{sh} \times \big( \frac{1}{1.5\%} \big) \approx \mathdollar 513/\mathrm{sh}, a span of over \mathdollar 1{,}000. With a precision of {\pm}20\mathrm{bp}, the range shrinks to \mathdollar 320: \mathdollar 7.70/\mathrm{sh} \times \big( \frac{1}{0.8\%} \big) \approx \mathdollar 962/\mathrm{sh} to \mathdollar 7.70/\mathrm{sh} \times \big( \frac{1}{1.2\%} \big) \approx \mathdollar 642/\mathrm{sh}. This massive improvement in accuracy is hard to fathom on present-value grounds.

Your Honor! I object…

You might balk at me calling the forward-price gap “noise.” Academics usually call it a time-varying risk premium. Fine. Call it whatever you want. Changing the name won’t supply the missing precision. If government bonds carry a risk premium that moves by {\pm}50\mathrm{bp} from year to year, then the S&P 500 should carry a risk premium that moves by at least as much. Stocks are riskier than Treasuries. The return-predictability literature exists because expected equity returns are supposed to swing by percentage points across cycles, not basis points. A cap rate move of {\pm}50\mathrm{bp} implies a price swing of {\pm}50\%. We observe {\pm}20\%. If you want to fly the “risk premium” banner, then you’d have to explain why the discount rate on the riskiest major asset class moves less than half as much as the discount rate on its safest one?

The remaining escape is to argue that r and g move together in a way that cancels out of the difference (r{-}g). But think about what this would require. The wandering in \mathrm{d}r is observed: the risk-free leg alone moves by 50\mathrm{bp} or more in a typical year. So for the difference to stay pinned to within {\pm}20\mathrm{bp}, you would need \mathrm{d}g to shadow \mathrm{d}r nearly move for move. Run the variance arithmetic: the growth forecast needs a spread between roughly 30\mathrm{bp} and 70\mathrm{bp} with a correlation to \mathrm{d}r above 0.9, year after year, decade after decade. Measured long-run growth expectations show nothing like that co-movement with rates. That is not an assumption. It is a century of coincidences stacked one on top of the other.

Filed Under: Uncategorized

Deriving the Gordon Model

July 26, 2026 by Alex

The Gordon model is a mainstay of MBA classes and motivating examples. The model says that a stock’s current share price will equal its expected dividend next year, \mathbb{E}_t[\text{Div}_{t+1}], times a forward multiple, \big( \frac{1}{\mathrm{r} - \mathrm{g}}\big),

(1)   \begin{equation*}\text{Price}_t = \mathbb{E}_t[\text{Div}_{t+1}] \times \bigg( \frac{1}{\mathrm{r} {-} \mathrm{g}}\bigg)\end{equation*}

\mathrm{r} is the stock’s annual risk-adjusted discount rate. A dollar paid out 4 years from now is worth \frac{\mathdollar 1}{(1+\mathrm{r})^4} today. \mathrm{g} = \big(\frac{\mathbb{E}_t[\text{Div}_{t+h}]}{\text{Div}_t}\big)^{1/h}{-}1 is the company’s anticipated dividend-growth rate at every horizon h \geq 1.

Myron Gordon’s idea was to scale up next period’s dividend forecast by a factor of \big( \frac{1}{\mathrm{r}-\mathrm{g}} \big) to capture the present value of the firm’s dividend stream from year (t{+}2) onward. e.g., suppose a firm has promised to pay \mathdollar 5.00/\mathrm{sh} in dividends next year. If investors apply a 10\% discount rate and anticipate 5\% annual dividend growth, then the Gordon model would price the stock at \mathdollar 5.00/\mathrm{sh} \times \big( \frac{1}{10\%-5\%} \big) = \mathdollar 100/\mathrm{sh}.

There’s nothing remotely complicated about this calculation. Researchers all learn the Gordon pricing formula the first week of their PhD program. We’re all very comfortable reasoning in these terms. So it’s easy to forget how much work goes into producing the result. There’s nothing simple or straightforward about it. The key step in the derivation of the Gordon model isn’t about assuming constant parameters. It’s getting rid of the unknown future resale price, which is only a problem when applying present-value logic to stocks.

This post walks through what it takes to derive the Gordon model. I use the Lean proof assistant to do a proper accounting of all the assumptions and steps involved.

Perpetuities

The Gordon model prices stocks by pretending they are bonds. To see the logic, it’s important to understand why it’s easier to apply present-value logic to fixed-income assets. Let’s start with the simplest one: a perpetuity. This is an asset that will pay the same annual coupon starting next year and continuing on until Kingdom come. The present value of this perpetual stream of coupon payments is

(2)   \begin{equation*}\text{Price}_t \;=\; \sum_{h=1}^{\infty} \frac{\text{Coupon}}{(1 {+} \mathrm{r})^h}\end{equation*}

\mathrm{r} > 0\% denotes the annual discount rate.

This geometric series can be simplified as follows

(3)   \begin{equation*}\text{Price}_t \;=\; \text{Coupon} \times \bigg( \frac{1}{\mathrm{r}} \bigg)\end{equation*}

Doubling the annual coupon payment doubles the price of the perpetuity. Lowering the discount rate makes each dollar that a perpetuity delivers in the future more valuable today, thereby increasing the overall price.

Coupon Bonds

An H-year coupon bond works like a perpetuity for the first (H{-}1) years. Both pay the same coupon each year. However, in the final year H, the coupon bond delivers its coupon payment as well as its face value, \mathrm{FV}. The price of a coupon bond reflects the present value of its payout stream

(4)   \begin{equation*}\text{Price}_t \;=\; \sum_{h=1}^{H} \frac{\text{Coupon}}{(1 {+} \mathrm{r})^h} \;+\; \frac{\mathrm{FV}}{(1{+}\mathrm{r})^H}\end{equation*}

The present-value logic is the same. Only the payout stream has changed.

The face value determines the scale of the bond. The size of the coupon is typically reported as a fraction of this number, \mathrm{Coupon} = \mathrm{c} \cdot \mathrm{FV}. Thus, we can write the price as

(5)   \begin{align*}\text{Price}_t \;&=\; \sum_{h=1}^{H} \frac{\text{Coupon}}{(1 {+} \mathrm{r})^h} \;+\; \frac{\mathrm{FV}}{(1{+}\mathrm{r})^H} \\ &=\; \sum_{h=1}^{H} \frac{\mathrm{c} \cdot \text{FV}}{(1 {+} \mathrm{r})^h} \;+\; \frac{\mathrm{FV}}{(1{+}\mathrm{r})^H} \\ &=\; \mathrm{FV} \times \Bigg\{ \sum_{h=1}^{H} \frac{\mathrm{c}\phantom{i}}{(1 {+} \mathrm{r})^h} \;+\; \frac{1\phantom{n}}{(1{+}\mathrm{r})^H} \Bigg\}\end{align*}

The first term in the curly braces, \sum_{h=1}^{H} \frac{\mathrm{c}\,}{(1 {+} \mathrm{r})^h}, is the present value of the coupons spun off by each dollar of face value. The second term, \frac{1\phantom{n}}{(1{+}\mathrm{r})^H}, is the present value of receiving that dollar when the bond matures.

Par Value

The face value is the relevant reference point for pricing bonds. If \text{Price}_t = \mathrm{FV}, then we say that a bond is “priced at par”. Each dollar of face value spins off a \mathrm{c} \cdot \mathdollar 1 coupon once a year for the next H years. When priced at par, the present value of these coupons exactly offsets the loss from having to wait H years to receive the dollar back

(6)   \begin{equation*}\text{@ par:} \qquad \underbrace{\phantom{\Bigg(}\!\!\!\!\mathdollar 1 - \frac{\mathdollar 1\phantom{m}}{(1{+}\mathrm{r})^H}}_{\substack{\text{Loss\phantom{j}from} \\ \text{waiting}}} \;=\; \underbrace{\sum_{h=1}^{H} \frac{\mathrm{c} \cdot \mathdollar 1}{(1 {+} \mathrm{r})^h}}_{\substack{\text{Gain\phantom{j}from} \\ \text{\phantom{t}coupons\phantom{t}}}}\end{equation*}

These two forces offset when the coupon rate equals the discount rate, which is why par bonds have \mathrm{c} = \mathrm{r}.

Bond traders use par pricing as a reference point when performing back-of-the-envelope calculations. When \mathrm{Price}_t = \mathrm{FV}, it doesn’t matter whether you get paid the face value at time (t{+}H) or continue to collect an infinite stream of coupons from year ([t{+}H]{+}1) onward

(7)   \begin{align*}\text{@ par:} \qquad \text{Price}_t \;&=\; \text{FV} \\ &=\; \text{FV} \times \underbrace{\bigg( \frac{\mathrm{c}}{\mathrm{r}} \bigg)}_{=1} \;=\; \underbrace{\text{Coupon} \times \bigg( \frac{1}{\mathrm{r}} \bigg)}_{\text{Perpetuity formula}}\end{align*}

If \mathrm{c} < \mathrm{r}, then the bond’s priced at a discount (below par). If \mathrm{c} > \mathrm{r}, then it’s priced at a premium (above par).

Core Problem

At first glance, it seems like it should be possible to apply the same present-value logic to pricing stocks. The one-year-ahead pricing rule for stocks looks similar to the pricing formula for a one-year coupon bond

(8)   \begin{align*}\text{bond:} \qquad \text{Price}_t \;&=\; \frac{\;\;\!\mathrm{Coupon}\;\;\!}{1{+}\mathrm{r}} + \frac{\;\;\;\;\;\;\;\!\mathrm{FV}\;\;\;\;\;\;\;\!}{1 {+} \mathrm{r}} \\ \text{stock:} \qquad \text{Price}_t \;&=\; \frac{\mathbb{F}_t[\text{Div}_{t+1}]}{1{+}\mathrm{r}} + \frac{\mathbb{F}_t[\text{Price}_{t+1}]}{1 {+} \mathrm{r}}\end{align*}

\mathbb{F}_t[\cdot] denotes investors’ forecast given time-t information. A forecast is just a number in investors’ heads. Nothing guarantees it obeys the laws of probability, so I reserve \mathbb{E}_t[\cdot] for forecasts that do. This distinction will play a big role later on. Chekhov’s gun applies to both screenplays and academic research.

The stock’s forecasted dividend payment next year is kind of like the bond’s coupon. The stock’s anticipated resale price a year from now is sort of like the bond’s face value. However, there’s a key difference. For the bond, \mathrm{Coupon} and \mathrm{FV} are both known at the time of purchase. In ye olde times, when you bought a bond, you received a big piece of paper with a bunch of tabs on the bottom. Each year, you tore off a tab and mailed it in to receive your coupon. When the bond matured, you sent in the last tab and the big sheet of paper to get paid the face value. You couldn’t do this for a stock. Nobody knows \mathrm{Div}_{t+1} or \mathrm{Price}_{t+1} with certainty when you buy a share at time t.

The core problem with using present-value logic to price equities is the forecasted resale price on the right-hand side, \mathbb{F}_t[\mathrm{Price}_{t+1}]. If you’re trying to figure out the functional form of \mathrm{Price}_t, then how are you supposed to know the right value to plug in for next year’s resale price? It is always possible to write an equation in which a stock’s current price equals the discounted payoff to owning a share next year. But for this equation to mean something, you need to remove the dependency of next year’s payoff on the future resale price. Otherwise, the relationship is circular.

Prior to Myron Gordon, people knew how to price bonds using present-value logic. But they didn’t know how to apply similar logic to assets like stocks where fluctuations in the future resale price represent a significant portion of the future payout. Gordon’s 1959 paper showed how to get around this problem by treating stocks like coupon bonds priced at par. Notice that the pricing rule is just a modified perpetuity formula, which includes an adjustment for a growing coupon. This is a bold claim about how stocks get priced. At the very least, it ain’t how people talk about pricing shares of Nvidia or Tesla.

Full Derivation

When researchers describe the Gordon model, they tend to focus on the fact that both \mathrm{r} and \mathrm{g} are constant. This is the least interesting part of the derivation. Let’s walk through what’s required to get from the one-period-ahead present-value formula to Myron Gordon’s result

(9)   \begin{equation*}\text{Price}_t \;=\; \frac{\mathbb{F}_t[\text{Div}_{t+1}] + \mathbb{F}_t[\text{Price}_{t+1}]}{1 + \mathrm{r}_t} \qquad \rightsquigarrow \qquad \text{Price}_t \;=\; \mathbb{E}_t[\mathrm{Div}_{t+1}] \times \bigg( \frac{1}{\mathrm{r}{-}\mathrm{g}} \bigg)\end{equation*}

The discount rate now carries a time subscript. Nothing in one-period-ahead present-value logic requires investors to apply the same discount rate every year, so from here on I let \mathrm{r}_t vary over time. There are 5 steps. The first 4 are where all the real heavy lifting takes place. \mathrm{r} and \mathrm{g} only lose their time subscripts in step #5 after the main formula has been derived. This is just cosmetic tidying-up.

Step #1: Assume Consistent Pricing

To iterate forward, investors must believe the one-period-ahead pricing rule holds at every future date (t{+}h). The same formula that governs today’s price must also govern the price targets in investors’ heads

(10)   \begin{equation*}\text{Price}_{t+h} = \frac{\mathbb{F}_{t+h}[\text{Div}_{(t+h)+1}] + \mathbb{F}_{t+h}[\text{Price}_{(t+h)+1}]}{1 + \mathrm{r}_{t+h}} \qquad \text{for all } h \geq 0\end{equation*}

\mathrm{r}_{t+h} is the one-period discount rate applied to payouts received at time ([t{+}h]{+}1). i.e., each dollar paid the following year is worth \frac{\mathdollar 1}{(1+\mathrm{r}_{t+h})} at time (t{+}h). One more assumption hides in this notation. Discount rates can differ across years, but the entire path \mathrm{r}_t, \mathrm{r}_{t+1}, \mathrm{r}_{t+2}, \ldots is known at time t.

There’s an important economic distinction between applying the formula today, h{=}0, and applying the formula in future years, h \geq 1. Even if most investors don’t think in present-value terms, you could argue that the invisible hand of the market somehow forces the current price to obey the one-period-ahead present-value rule at time t. But you can’t make the same argument for h \geq 1. The pricing formula for these future dates can only exist in investors’ heads. If they don’t think in present-value terms, then there’s no reason for the formula to hold. The claim is a substantive assumption about how investors think.

Step #2: Assume The Tower Property

The tower property says that today’s forecast of next year’s forecast is just today’s forecast

(11)   \begin{equation*}\mathbb{F}_t\big[\mathbb{F}_{t+1}[\,\cdot\,]\big] = \mathbb{F}_t[\,\cdot\,]\end{equation*}

The same is true if we replace next year, h{=}1, with any other longer horizon. Arbitrary forecasts do not have this feature. The tower property is only satisfied by conditional expectations that stem from a well-posed probability space. It is often referred to as the “law of iterated expectations”. I call it the “tower property” to emphasize the distinction between arbitrary forecasts and coherent expectations. From here on out, I write \mathbb{E}_t[\cdot] rather than \mathbb{F}_t[\cdot].

One last thing. Coherent doesn’t mean correct. The law of iterated expectations can be applied to expectations that aren’t objectively correct. It’s a property of the belief structure, not whether these beliefs match the true data-generating process. Biased subjective expectations are a subset of all possible forms of incorrect beliefs. It’s possible to make incorrect forecasts that violate the laws of probability.

Step #3: Iterate Forward Finite Times

The next step is finite induction. Take the one-period-ahead pricing rule and replace the resale price on the right-hand side with its functional form for the following year. If you do this (H{-}1) times, then you get the following expression

(12)   \begin{equation*}\text{Price}_t \;=\; \underbrace{\sum_{h=1}^{H} \frac{\mathbb{E}_t[\text{Div}_{t+h}]}{\prod_{k=0}^{h-1} (1 {+} \mathrm{r}_{t+k})}}_{\text{PV first H dividends}} \;+\; \underbrace{\frac{\mathbb{E}_t[\text{Price}_{t+H}]}{\prod_{k=0}^{H-1} (1 {+} \mathrm{r}_{t+k})}}_{\text{PV resale price}}\end{equation*}

The company’s current share price reflects its expected discounted dividend payments over the next H years plus the present value of the expected resale price H years from now.

Notice that this step doesn’t purge the future resale price from the right-hand side. The date of reckoning has just been pushed farther into the future. In a sense, this makes the original problem worse. If next year’s resale price was hard to fathom, then why would investors have any idea about the price each share might sell for 20 or 100 years in the future? At this point, it’s not obvious progress has been made.

Step #4: Assume Limit Is Well-Behaved

The payoff to iterating forward only occurs when you take the infinite limit, H \to \infty. We’re looking to remove the dependency of the current price on the expected future resale value. For this to happen, we need two things to be true:

  1. Transversality. The expected discounted resale price must go to zero

    (13)   \begin{equation*}\lim_{H \to \infty} \, \frac{\mathbb{E}_t[\text{Price}_{t+H}]}{\prod_{k=0}^{H-1} (1 {+} \mathrm{r}_{t+k})} \;=\; \mathdollar 0\end{equation*}

    The one-period recursion has infinitely many solutions. A rational bubble also satisfies it. Transversality selects the “correct” price, which reflects expected discounted dividends alone.

  2. Convergence. The infinite sum of the stock’s expected discounted dividends must converge to a single finite number, and the answer cannot depend on the order in which the terms get added up. This second requirement is where the bite is. Adding up the discounted dividends in time order and getting a finite limit follows for free from step #3 plus transversality. Absolute convergence does not. Researchers often focus on transversality and take this second condition for granted. But both are strong assumptions. Convergence is a genuine premise of its own, not merely a footnote.

By making both assumptions, it’s possible to eliminate the future resale price entirely. The resulting pricing rule is known as the Dividend Discount Model (DDM). It says that a company’s share price at time t should reflect the discounted value of its expected future dividend stream from time (t{+}1) onward

(14)   \begin{equation*}\text{Price}_t \;=\; \sum_{h=1}^{\infty} \frac{\mathbb{E}_t[\text{Div}_{t+h}]}{\prod_{k=0}^{h-1} (1 {+} \mathrm{r}_{t+k})}\end{equation*}

We’ve now overcome the main challenge in deriving a present-value pricing rule for stocks.

Step #5: Assume Constant Parameters

All the heavy lifting is already done. This last step is about ease-of-use. Most people don’t have clear views about a company’s likely dividend in 2077. They don’t have nuanced views about whether to apply a higher one-year discount rate in 2077 or 2076. So, to make the formula more practical, let’s assume that the stock’s future dividend grows at a constant annual rate

(15)   \begin{equation*}\mathbb{E}_t[\text{Div}_{t+h}] \;=\; (1 + \mathrm{g})^{h-1} \!\cdot \mathbb{E}_t[\text{Div}_{t+1}] \;=\; (1 + \mathrm{g})^h \cdot \text{Div}_t\end{equation*}

Let’s also assume that the same annual discount rate gets applied to every horizon h \geq 0

(16)   \begin{equation*}{\textstyle \prod_{k=0}^{h-1}} (1 {+} \mathrm{r}_{t+k}) \;=\; (1 + \mathrm{r})^h\end{equation*}

The assumption of constant parameters turns the infinite sum with a telescoping product in the denominator into a simple geometric series

(17)   \begin{align*}\text{Price}_t \;&=\; \sum_{h=1}^{\infty} \frac{\mathbb{E}_t[\text{Div}_{t+h}]}{\prod_{k=0}^{h-1} (1 {+} \mathrm{r}_{t+k})} \\ &=\; \sum_{h=1}^{\infty} \frac{(1{+}\mathrm{g})^{h-1} \cdot \mathbb{E}_t[\text{Div}_{t+1}]}{(1 {+} \mathrm{r})^h} \\ &=\; \mathbb{E}_t[\text{Div}_{t+1}] \times \sum_{h=1}^{\infty} \frac{(1{+}\mathrm{g})^{h-1}}{(1 {+} \mathrm{r})^{h\phantom{-1}}} \\ &=\; \mathbb{E}_t[\text{Div}_{t+1}] \times \bigg(\frac{1}{\mathrm{r} {-} \mathrm{g}}\bigg)\end{align*}

Assuming constant \mathrm{r} and \mathrm{g} makes it possible to express the implications of the DDM in a clean way.

With constant parameters, the transversality and convergence assumptions in step #4 boil down to the requirement that \mathrm{r} > \mathrm{g}. If this condition is violated, \mathrm{r} \leq \mathrm{g}, then the present value of the stock’s expected discounted dividend stream will be infinite. e.g., suppose a stock’s future payout stream gets discounted at \mathrm{r}=3\% annually and the company paid a \mathdollar 1.00/\mathrm{sh} dividend last year. If the firm’s dividend-growth rate is \mathrm{g}=4\%, then next year investors expect \mathbb{E}_t[\mathrm{Div}_{t+1}] = \mathdollar 1.04/\mathrm{sh}. Had the firm maintained the same dividend, this cash flow would only be worth \mathdollar 0.97 today. But they expect an extra \mathdollar 0.04 in dividends next year, and this is more than enough to make up for the valuation drag created by discounting.

Assumption Accounting

I use Lean to properly account for all the different assumptions used in the derivation of the Gordon model. The hard part is getting rid of the price forecast on the right-hand side:

  1. Assume that the one-period-ahead pricing rule holds today as well as at every future date. It governs observed prices and the price forecasts in investors’ heads.
  2. Assume that investors’ price forecasts satisfy the tower property. This requires their subjective beliefs to represent conditional expectations that stem from a well-defined subjective probability measure.
  3. Iterate forward a finite number of times, pushing the unknown future resale price far into the future.
  4. Assume that the infinite limit has the properties needed to eliminate the current price’s dependence on the future resale price. These are transversality (a.k.a., no bubbles) and convergence.

The final step is purely cosmetic. It occurs after the troublesome resale price has already been expunged.

  1. Assume constant \mathrm{r} and \mathrm{g}.

The standard telling treats step #5 as the key assumption behind the Gordon model, but the honest ledger shows it is the last and lightest. The core derivation lives in steps #1-4. In addition to maintaining the ledger, the proof in Lean shows that each of the load-bearing assumptions in these steps is necessary: the tower property, transversality, and convergence. There are explicit counterexamples that satisfy everything else and yet break the conclusion.

The point of running the Gordon model through a proof assistant is not the machinery. Every well-trained economist has seen all these ideas before. The issue is that researchers have gotten so familiar with Gordon logic that they often forget all that it requires. The derivation is neither short nor innocent. Lean forces you to reckon with every required step in the proof.

Filed Under: Uncategorized

Trailing PEs Imply Low Elasticities

July 22, 2026 by Alex

A frictionless mean-variance model predicts an aggregate demand elasticity of \nu = 25. Suppose the level of the stock market rises by 1\% on no fundamental news. It’s now 1\% more expensive to buy stocks, but nothing’s changed to make the anticipated payout next year more desirable. Textbook theory says that investors ought to look at this drop in forecasted returns and dump 25\% of their holdings. The data disagrees. There, the aggregate demand elasticity is much much lower. Gabaix-Koijen estimate \nu \approx 0.2.

To get an elasticity that low, investors need to look at the 1\% increase in today’s price and shrug their shoulders. In this note, I show that this is exactly what happens when investors rely on trailing PE ratios when setting price targets. I show that this one simple observation is able to generate a predicted demand elasticity of \nu \approx 0.8. This is well within spitting distance of the estimated 0.2.

The trailing-PE mechanism is kind of like a dogmatic-learning story. Think about a Bayesian investor who treats the current price level as a very precise signal about next year’s payout. Such an investor would face the same demand curve as a trailing-PE user. But the analogy isn’t perfect. The trailing-PE approach doesn’t force next year’s price target to agree with next year’s dividend forecast in present-value terms. When the current price rises, the target rises with it, but the dividend forecast doesn’t budge.

Demand Elasticity

Suppose an asset’s current price changes a tiny bit for non-fundamental reasons. Suppose an investor’s forecasting and allocation rules remain unchanged. How much will her desired position change in response? The answer to this question is called the demand elasticity

(1)   \begin{equation*}\nu \;=\; - \frac{\partial \log \mathrm{Dmnd}}{\partial \log \mathrm{Price}} \;=\; (1 {-} \theta) \;+\; \bigg( \frac{\mu}{\bar{r}} \bigg) \times \eta\end{equation*}

A change in the current price of an asset affects the investor’s demand in two ways. There’s a rebalancing channel, (1{-}\theta), which creates a difference between stock-level and aggregate demand elasticities. There’s also a belief channel. A change in today’s price can impact the investor’s views about next year’s payoff. This is the \big( \tfrac{\mu}{\bar{r}} \big) \times \eta term, and it pins down the overall level. Here’s where this formula comes from.

Let \theta = \mathrm{Dmnd}_t \times \big\{ \frac{\mathrm{Price}_t}{\mathrm{Wealth}_t} \big\} denote an asset’s share of an investor’s wealth. Hold her forecasted return fixed, which switches off the belief channel and keeps a fixed fraction of her wealth in the asset. Under this assumption, a 1\% price rise means the same number of shares now ties up 1\% more of her wealth. Restoring her target weight means trimming shares and parking the proceeds in the rest of the portfolio.

If the investor already holds the asset, then the trim is partly cancelled. The increase in the current price will also revalue her existing position, raising her wealth by \theta percent

(2)   \begin{equation*}\frac{\partial \log \mathrm{Wealth}_t}{\partial \log \mathrm{Price}_t} \;=\; \theta\end{equation*}

This change lifts her target dollar allocation for the asset by the same \theta percent.

The net sale is (1{-}\theta) percent of her shares, which is exactly the share of her portfolio held in other assets. That is the room she has to rebalance into. For a single stock inside a diversified portfolio, we have \theta \approx 0 and (1{-}\theta) \approx 1. For an investor choosing how much to invest in the market as a whole, we have \theta = 1 and (1{-}\theta) = 0. The whole term vanishes.

The belief channel starts with a mapping from the current price to beliefs about future returns. I write the investor’s forecasts as \mathbb{F}_t[\cdot], rather than \mathbb{E}_t[\cdot], because forecasts don’t need to come from a well-posed probability space. A forecast is just a number that the investor writes down.

A share bought today for \mathrm{Price}_t delivers the dollar payout \mathrm{Payout}_{t+1} next year. The realized gross return on this investment will be

(3)   \begin{equation*}1 + \mathrm{Ret}_{t+1} \;=\; \frac{\mathrm{Payout}_{t+1}}{\mathrm{Price}_t} \;=\; e^{\log \mathrm{Payout}_{t+1} - \log \mathrm{Price}_t}\end{equation*}

If you expand the exponential expression around the steady state, then you get the following first-order approximation for the net return

(4)   \begin{equation*}\mathrm{Ret}_{t+1} \;\approx\; (1 {+} \overline{\mathrm{DY}}) \cdot \big( \log \mathrm{Payout}_{t+1} \,-\, \log \mathrm{Price}_t \big) \,+\, \mathrm{constant}\end{equation*}

I use the Gabaix-Koijen calibration values: a long-run dividend yield \overline{\mathrm{DY}} = 3.7\% and an average forecasted return \bar{r} = 4.4\%. A 1\% rise in the payout raises the gross return by (1+\overline{\mathrm{DY}}) \times 1\% \approx 1.037\%, and a 1\% rise in the current price lowers the return by the same amount.

Assumption A1. The investor forecasts the asset’s future payout with a rule \mathbb{F}_t[\mathrm{Payout}_{t+1}] that is differentiable in \log \mathrm{Price}_t, and her return forecast is given by

(5)   \begin{equation*}\mathbb{F}_t[\mathrm{Ret}_{t+1}] \;=\; (1 {+} \overline{\mathrm{DY}}) \cdot \big( \log \mathbb{F}_t[\mathrm{Payout}_{t+1}] \,-\, \log \mathrm{Price}_t \big) \,+\, \mathrm{constant}\end{equation*}

Notice what A1 does not assume: present-value logic. Nothing forces today’s price to equal her payout forecast discounted at a required return, so her payout forecast can move independently of the multiple, and a shock to \mathbb{F}_t[\mathrm{EPS}_{t+1}] need not have any impact on the PE. A1 does not impose Gordon logic, either directly or approximately as in Campbell-Shiller. A1 pins down the investor’s return forecast for next year as a function of her payout forecast and the current price.

Differentiating gives the belief drag, \mu. This parameter represents the price sensitivity of the investor’s forecasted return for the upcoming year

(6)   \begin{equation*}\mu \;=\; {-}\frac{\partial \, \mathbb{F}_t[\mathrm{Ret}_{t+1}]}{\partial \log \mathrm{Price}_t} \;=\; (1 {+} \overline{\mathrm{DY}}) \times \bigg( 1 \,-\, \frac{\partial \log \mathbb{F}_t[\mathrm{Payout}_{t+1}]}{\partial \log \mathrm{Price}_t} \bigg)\end{equation*}

If the current price goes up by 1\% and nothing else changes, how much will the investor’s return forecast fall in response?

In a frictionless mean-variance model, the investor observes the asset’s current price. But this information doesn’t impact how she values the stock. Her payout forecast is built from fundamentals alone, so \frac{\partial \log \mathbb{F}_t[\mathrm{Payout}_{t+1}]}{\partial \log \mathrm{Price}_t} = 0 and \mu = (1 {+} \overline{\mathrm{DY}}) \times (1{-}0) \approx 1.037. She suffers the full drag.

The belief drag has units of percent per year. \mu is a change in the asset’s anticipated return over the next twelve months. The average forecasted return \bar{r} has the same units. So the ratio \big( \tfrac{\mu}{\bar{r}} \big) is dimensionless, which an elasticity term has to be. The two terms combine to form the elasticity of the forecasted return with respect to the current price.

A textbook investor in a frictionless mean-variance model has belief drag \mu = (1 {+} \overline{\mathrm{DY}}) \approx 1.037. A 1\% increase in the current price level will lower her forecasted return for next year by 1.037\%\mathrm{pt}. When we compare this effect to the long-run average return forecast, \bar{r}=4.4\%, we get a relative change of \big( \tfrac{1.037\%\mathrm{pt}}{4.4\%} \big) \approx 23.6\%. The original 1.037\%\mathrm{pt} drag on the asset’s forecasted return may not sound like much, but it’s a big deal compared to the average return forecast.

\big( \tfrac{\mu}{\bar{r}} \big) measures the percent decline in the forecasted return when the price rises 1\%. The parameter \eta converts this elasticity of forecasted returns into an elasticity of demand. You might think this conversion requires a full-fledged asset-pricing model. It turns out any allocation rule with the following form will do.

Assumption A2. The dollar allocation is \mathrm{Wealth}_t \times w\big(\mathbb{F}_t[\mathrm{Ret}_{t+1}]\big) for a smooth increasing rule w(\cdot).

The number of shares that the investor demands can be written as w times the ratio of her initial wealth and the asset’s share price

(7)   \begin{equation*}\mathrm{Dmnd}_t \;=\; w\big(\mathbb{F}_t[\mathrm{Ret}_{t+1}]\big) \times \bigg\{ \frac{\mathrm{Wealth}_t}{\mathrm{Price}_t}\bigg\}\end{equation*}

e.g., mean-variance preferences deliver the special case w(x) = \big(\frac{1}{\gamma \cdot \sigma^2}\big) \cdot x.

\eta(\bar{r}) represents the elasticity of the investor’s dollar position in the asset with respect to her forecasted return next year evaluated at the asset’s long-run average forecast

(8)   \begin{equation*}\eta(\bar{r}) \;=\; \bar{r} \times \bigg\{ \frac{w'(\bar{r})}{w(\bar{r})} \bigg\}\end{equation*}

e.g., if an asset’s forecasted return improves by 1\%, from \bar{r} = 4.4\% to 4.444\%, the investor scales her dollar allocation in the asset up by \eta(\bar{r}) \times 1\%.

Note that \eta(\bar{r}) \approx 1 to leading order. For mean-variance preferences, we have \eta(\bar{r}) = 1 exactly. To see why, take any smooth rule with w(0) = 0. A Taylor expansion gives w(\bar{r}) = w'(0) \cdot \bar{r} \cdot (1 {+} O(\bar{r})), so

(9)   \begin{equation*}\eta(\bar{r}) \;=\; 1 \,+\, \frac{1}{2} \cdot \bigg\{\frac{w''(0)}{w'(0)}\bigg\} \times \bar{r} \,+\, O(\bar{r}^2)\end{equation*}

If w(x) = \big(\frac{1}{\gamma \cdot \sigma^2}\big) \cdot x, then \frac{\mathrm{d}w}{\mathrm{d}x} = \big(\frac{1}{\gamma \cdot \sigma^2}\big) and \frac{\mathrm{d}^nw}{\mathrm{d}x^n} = 0 for all n \geq 2. Hence, we have \eta(\bar{r}) = 1 for all \bar{r}. However, any preference specification with w(0)=0 and w''(0)=0 would do the same to leading order.

The pieces now assemble by the chain rule. First, take logs of the demand rule

(10)   \begin{equation*}\log \mathrm{Dmnd}_t \;=\; \log w\big(\mathbb{F}_t[\mathrm{Ret}_{t+1}]\big) \,+\, \log \mathrm{Wealth}_t \,-\, \log \mathrm{Price}_t\end{equation*}

Next, differentiate each term with respect to \log \mathrm{Price}_t. The wealth term contributes \theta and the price term contributes -1. Together they are the rebalancing channel, (1{-}\theta).

The first \log w\big(\mathbb{F}_t[\mathrm{Ret}_{t+1}]\big) term is the belief channel. The forecasted return falls by \mu per unit of \log \mathrm{Price}_t, and log dollars move by \big\{ \frac{w'(\bar{r})}{w(\bar{r})} \big\} = \big( \tfrac{1}{\bar{r}} \big) \times \eta per unit of forecasted return. Flipping the sign delivers the headline formula

(11)   \begin{equation*}\nu \;=\; -\frac{\partial \log \mathrm{Dmnd}_t}{\partial \log \mathrm{Price}_t} \;=\; (1 {-} \theta) \;+\; \bigg( \frac{\mu}{\bar{r}} \bigg) \times \eta\end{equation*}

A 1\% price increase lowers next year’s return forecast by \mu percentage points. Dividing by \bar{r} converts this drop into an elasticity of returns, and multiplying by \eta translates it into a demand elasticity.

We can now cleanly state the inelastic-markets result of Gabaix-Koijen. Start with the textbook prediction. Consider a mean-variance investor in a frictionless model where the dividend yield is \overline{\mathrm{DY}} = 3.7\% and the long-run average return is \bar{r}=4.4\%. The predicted belief drag is \mu = (1{+}\overline{\mathrm{DY}}) \times (1-0) = 1.037. This change in next year’s return forecast represents roughly \frac{1.037\%\mathrm{pt}}{4.4\%} \approx 23.6\% of the average return forecast. With \eta = 1, this return elasticity translates to a 23.6\% change in demand. For the market as a whole, \theta = 1 and (1{-}\theta)=0, so \nu = 23.6. For an individual stock, \theta=0 and (1{-}\theta)=1, giving a demand elasticity that is one turn higher, \nu = 1 {+} 23.6 = 24.6. Gabaix-Koijen estimate an aggregate demand elasticity of \hat{\nu} = 0.2. The textbook prediction is off by two orders of magnitude, \frac{23.6}{0.2} \approx 118!

Trailing PE Ratio

Sell-side analysts typically describe setting one-year-ahead price targets using a two-step process. First, an analyst forecasts the stock’s EPS over the next year based on non-price information. Then, the analyst capitalizes this short-term earnings forecast into a price target using the company’s current multiple, \mathrm{PE}_t = \big( \frac{\mathrm{Price}_t}{\mathrm{EPS}_t} \big), which is known as the trailing PE ratio

(12)   \begin{equation*}\mathbb{F}_t[\mathrm{Price}_{t+1}] = \mathbb{F}_t[\mathrm{EPS}_{t+1}] \times \mathrm{PE}_t\end{equation*}

This price forecast implies that next year’s return forecast will consist of two components: the firm’s anticipated dividend yield and forecasted earnings growth

(13)   \begin{align*}\mathbb{F}_t[\mathrm{Ret}_{t+1}] \;&=\; \bigg(\frac{\mathbb{F}_t[\mathrm{Div}_{t+1}]}{\mathrm{Price}_t}\bigg) \,+\, \bigg(\frac{\mathbb{F}_t[\mathrm{Price}_{t+1}] - \mathrm{Price}_t}{\mathrm{Price}_t}\bigg) \\ &=\; \bigg(\frac{\mathbb{F}_t[\mathrm{Div}_{t+1}]}{\mathrm{Price}_t}\bigg) \,+\, \bigg(\frac{\mathbb{F}_t[\mathrm{EPS}_{t+1}] {\times} \mathrm{PE}_t - \mathrm{EPS}_t {\times} \mathrm{PE}_t}{\mathrm{EPS}_t {\times} \mathrm{PE}_t}\bigg) \\ &=\; \bigg(\frac{\mathbb{F}_t[\mathrm{Div}_{t+1}]}{\mathrm{Price}_t}\bigg) \,+\, \bigg(\frac{\mathbb{F}_t[\mathrm{EPS}_{t+1}] - \mathrm{EPS}_t}{\mathrm{EPS}_t}\bigg)\end{align*}

Since the analyst uses today’s PE ratio to forecast next year’s price, the multiple drops out of the price appreciation term. Regardless of the current level, the analyst anticipates that the firm’s price will grow at the same rate as its earnings.

Notice that only one of the two components of the analyst’s return forecast includes the current price. This is clearly going to have implications for demand elasticities. To see what those are, let’s look at a concrete example. Consider a company that realized earnings of \mathrm{EPS}_t = \mathdollar 5.00/\mathrm{sh} last year. Over the next year, analysts anticipate that the company’s earnings will grow by 0.7\% to \mathbb{F}_t[\mathrm{EPS}_{t+1}] = \mathdollar 5.035/\mathrm{sh}. The stock starts out trading at \mathrm{Price}_t = \mathdollar 100.00/\mathrm{sh}, giving the firm a trailing multiple of \mathrm{PE}_t = \frac{\mathdollar 100.00/\mathrm{sh}}{\mathdollar 5.00/\mathrm{sh}} = 20\times. Given how the market is currently pricing the company’s earnings, analysts expect the firm to be trading at a price of \mathbb{F}_t[\mathrm{Price}_{t+1}] = \mathdollar 5.035/\mathrm{sh} \times 20 = \mathdollar 100.70/\mathrm{sh} next year. The company has committed to paying \mathdollar 3.73/\mathrm{sh} in dividends next year, giving the firm an anticipated dividend yield of 3.7\% and a return forecast of \mathbb{F}_t[\mathrm{Ret}_{t+1}] = 0.7\% + 3.7\% = 4.4\%.

Now, imagine that the company’s current price suddenly increases by 1\% to \mathrm{Price}_t = \mathdollar 101.00/\mathrm{sh}. The firm’s earnings over the last twelve months don’t move, \mathrm{EPS}_t = \mathdollar 5.00/\mathrm{sh}. Nothing about the company’s fundamentals change, either. Analysts still have the same next-twelve-month earnings forecast, \mathbb{F}_t[\mathrm{EPS}_{t+1}] = \mathdollar 5.035/\mathrm{sh}. But the company’s higher current price means that this short-term forecast will get capitalized at a higher multiple, \mathrm{PE}_t = \frac{\mathdollar 101.00/\mathrm{sh}}{\mathdollar 5.00/\mathrm{sh}} = 20.2\times, when setting a price target, \mathbb{F}_t[\mathrm{Price}_{t+1}] = \mathdollar 5.035/\mathrm{sh} \times 20.2 = \mathdollar 101.71/\mathrm{sh}. Yet the higher price target has no impact on analysts’ beliefs about future price growth because the current earnings are also being priced using a multiple that is 0.2\times higher. The higher current price level only affects analysts’ return forecast for next year by diluting the dividend yield. Instead of \frac{\mathdollar 3.73/\mathrm{sh}}{\mathdollar 100/\mathrm{sh}} = 3.7\%, analysts now anticipate a dividend yield of \frac{\mathdollar 3.73/\mathrm{sh}}{\mathdollar 101/\mathrm{sh}} = 3.664\%.

Prior to the price increase, the company’s forecasted payout was the \mathdollar 100.70/\mathrm{sh} target price plus the \mathdollar 3.73/\mathrm{sh} forecasted dividend, which came out to \mathdollar 104.43/\mathrm{sh} total. The resale price contributed \big( \frac{1}{1 + \overline{\mathrm{DY}}} \big) \approx 96.3\% of the total payout. Following the unilateral 1\% price increase, the company’s multiple expanded and its price target also rose by 1\%. But its dividend forecast stood still. Hence, analysts’ forecasted payout didn’t rise by a full 1\%. The drag on analysts’ beliefs is thus

(14)   \begin{equation*}\mu \;=\; (1 {+} \overline{\mathrm{DY}}) \times \bigg( 1 - \frac{1}{1 {+} \overline{\mathrm{DY}}} \bigg) \;=\; \overline{\mathrm{DY}} \;=\; 3.7\%\end{equation*}

This drag gets compared to the same average forecast as before, \bar{r} = 4.4\%. But, given the much smaller starting value, 0.037 vs 1.037, the resulting elasticity of returns is much smaller, \big( \frac{\mu}{\bar{r}} \big) = \frac{3.7\%\mathrm{pt}}{4.4\%} \approx 0.8 rather than 23.6. Assuming \eta = 1, you get a single-stock demand elasticity of \nu = 1 + 0.8 \approx 1.8 and an aggregate demand elasticity of \nu \approx 0.8.

The trailing-PE approach gets you from 23.6 down to 0.8. The remaining distance, from 0.8 down to the estimated 0.2, is likely due to \eta rather than beliefs. Mandates, inertia, and the other demand-side frictions at the center of the inelastic-markets literature all mute the position response, which corresponds to \eta < 1. An \eta \approx 0.25 closes the gap. On this reading, the trailing-PE rule and demand-side frictions are complements, not competitors. Beliefs deliver the first factor of \frac{23.6}{0.8} \approx 30. Frictions deliver the last factor of \frac{0.8}{0.2} \approx 4.

Learning Story

The trailing-PE rule hardcodes the link between this year’s multiple and next year’s multiple. A learning story can deliver something similar without hardcoding anything. Think about Grossman-Stiglitz. When the current price goes up by 1\%, an investor might worry that everyone else knows something she doesn’t. And, as a result, she might raise her forecast of the future payout. This is the story in Bastianello (2026).

Consider an investor who solves the following Gaussian inference problem. The investor has prior beliefs about the stock’s future payout

(15)   \begin{equation*}\log \mathrm{Payout}_{t+1} \;\sim\; \mathrm{Normal}\big( \log \mathrm{Prior}_t, \, 1 \big)\end{equation*}

For clarity, I’ve suppressed a constant term, which reflects discounting and the risk premium. I’ve also normalized the prior variance to 1. The current price level is a noisy signal about the future payout

(16)   \begin{equation*}\log \mathrm{Price}_t \;\sim\; \mathrm{Normal}\big( \log \mathrm{Payout}_{t+1}, \; 1/\tau \big)\end{equation*}

\tau > 0 is the precision of the price signal. A larger value of \tau implies that prices are more informative about the stock’s likely payout next year.

Standard Gaussian-updating rules imply that the investor’s posterior beliefs about the payout will be a weighted average

(17)   \begin{equation*}\mathbb{E}_t[\log \mathrm{Payout}_{t+1}|\log \mathrm{Price}_t] \;=\; (1 {-} \lambda) \cdot \log \mathrm{Prior}_t \,+\, \lambda \cdot \log \mathrm{Price}_t\end{equation*}

Her beliefs are a true conditional expectation, so I write them with \mathbb{E}_t[\cdot] rather than \mathbb{F}_t[\cdot]. The weights reflect the precision of the price signal. The investor leans more heavily on the current price when it is a more precise signal about the stock’s future payout, \lambda = \big(\frac{\tau}{1 + \tau}\big).

From here, it’s straightforward to derive the key inputs to the demand-elasticity formula. Start with the belief drag. Differentiating the log expected payout with respect to \log \mathrm{Price}_t gives

(18)   \begin{equation*}\frac{\partial \log \mathbb{E}_t[\mathrm{Payout}_{t+1}|\log \mathrm{Price}_t]}{\partial \log \mathrm{Price}_t} \;=\; \lambda \qquad \rightsquigarrow \qquad \mu \;=\; (1 {+} \overline{\mathrm{DY}}) \cdot (1 {-} \lambda)\end{equation*}

Learning scales the entire textbook drag down by a factor of (1{-}\lambda). There’s no impact on how this drag gets converted into a return elasticity. For the market as a whole, we get a predicted demand elasticity of

(19)   \begin{equation*}\nu \;=\; (1{-}\theta) \,+\, \bigg( \frac{(1{+}\overline{\mathrm{DY}}) \cdot (1{-}\lambda)}{\bar{r}} \bigg) \times \eta\end{equation*}

If the price signal is entirely uninformative, \lambda = 0, you get back the original formula. If the price signal is perfectly revealing, \lambda = 1, the entire belief channel dies. Only the rebalancing term remains.

Partial Symmetry

I motivated the learning story above by pointing out that, if you squint, it looks a bit like the trailing-PE approach. In both cases, an investor sees the current price level change and assumes that most of the change will propagate into the future payout. Using a trailing PE is kind of like viewing the current price level as a very precise signal about the future payout.

Consistent with this intuition, it’s possible to make the two mechanisms produce identical elasticities. All you have to do is equate the belief drags. The trailing-PE approach says \mu = \overline{\mathrm{DY}}. The learning story says \mu = (1{+}\overline{\mathrm{DY}}) \cdot (1{-}\lambda). For both to produce the same value, you need

(20)   \begin{equation*}\overline{\mathrm{DY}} \;=\; (1{+}\overline{\mathrm{DY}}) \cdot (1{-}\lambda^{\star}) \qquad \rightsquigarrow \qquad \lambda^{\star} \;=\; \frac{1}{1{+}\overline{\mathrm{DY}}}\end{equation*}

Assuming \overline{\mathrm{DY}} = 3.7\% would imply that \lambda^{\star} \approx 0.963. A learner who put 96.3\% weight on the current price signal would have the same demand curve as an analyst who set price targets using a trailing PE, \nu = 1.8 for a single stock and 0.8 for the aggregate at \eta = 1. So there is a precise sense in which using a trailing PE and putting a lot of weight on the current price level are symmetric.

Gabaix-Koijen estimates \nu \approx 0.2. To match that result, a learning story would need to put weight on the current price of \lambda \approx 97\% or higher. This is worth pausing on. The learning model in Bastianello is built from standard Bayesian ingredients. The paper never talks in terms of trailing multiples. Yet at the weight implied by the data, the investor submits the same demand curve as an analyst using a trailing PE ratio. Fit to the data, the learning story does not offer an alternative to the trailing-PE mechanism. It approximates it.

But the symmetry isn’t perfect. And the way that it breaks is interesting. The key thing in the trailing-PE story is not the trailing PE. It is that the investor treats next period’s EPS forecast and the current price as unrelated. There is no present-value calculation connecting the two. Her EPS forecast comes from sales and margins, her level comes from the market, and neither number is evidence about the other.

The difference shows up in how a price shift reaches each investor. For the trailing-PE analyst, a 1\% rise in the current price level moves her price target by the same 1\%. So the capital-gain piece of her forecasted return never budges. The entire impact of the price shock arrives through the stock’s dividend yield, and that is why her drag equals \overline{\mathrm{DY}} exactly.

By contrast, an investor who learns about the stock’s future payout from its current price sees both components of her payout forecast change. The price rise is news about fundamentals, which affects her beliefs about next year’s dividend and next year’s resale price. In this scenario, the investor’s belief drag gets spread across the whole forecast rather than concentrated in the dividend.

The apportionment can be made exact. At \lambda = \lambda^{\star} = 96.3\%, both investors mark up their payout forecast by 0.963\%\mathrm{pt} in response to a 1\% price increase, leaving the same 0.037\%\mathrm{pt} shortfall. But the two stories aren’t equivalent. They each place that 0.037\%\mathrm{pt} gap in different places. The trailing-PE analyst increases her price target by 1\%. Her dividend-yield forecast absorbs the entire shortfall. The learner pins 0.036\%\mathrm{pt} on the capital gain and 0.001\%\mathrm{pt} on the dividend yield. In the running example, both investors would forecast the same payout next year, \mathdollar 105.44/\mathrm{sh}. The analyst gets there as \mathdollar 101.71 + \mathdollar 3.73. The learner gets there as \mathdollar 101.67 + \mathdollar 3.77. Same number, different tickets, and the dividend line is the tell.

Filed Under: Uncategorized

Inelastic Markets ~ Flat SML

July 21, 2026 by Alex

Stock returns are typically higher than bond returns. The average difference is somewhere in the neighborhood of \mathbb{E}[\mathrm{Mkt}]{-}\mathrm{rf} \approx 4\%. Academic researchers call this the equity risk premium.

The CAPM is the most famous asset-pricing model in the literature. The theory predicts that each stock’s expected excess return will be proportional to the expected excess return on the market portfolio, \mathbb{E}[\mathrm{Ret}_n {-} \mathrm{rf}] \;=\; \beta_n \times \mathbb{E}[\mathrm{Mkt} {-} \mathrm{rf}]. Suppose you plot each stock’s expected excess return, \mathbb{E}[\mathrm{Ret}_n] {-} \mathrm{rf} (y-axis), against its market beta, \beta_n (x-axis). The resulting line is called the “SML” (security market line)

(1)   \begin{equation*}\mathbb{E}[\mathrm{Ret}_n] {-} \mathrm{rf} \;=\; \beta_n \times \lambda\end{equation*}

The CAPM predicts that you ought to get a line with slope \lambda = \mathbb{E}[\mathrm{Mkt}] {-} \mathrm{rf}.

The same simple model predicts that aggregate demand elasticity should be roughly

(2)   \begin{equation*}\nu \;\approx\; \frac{1}{\mathbb{E}[\mathrm{Mkt}] {-} \mathrm{rf}}\end{equation*}

If the level of the stock market goes up by 1\% for non-fundamental reasons, then it suddenly costs more money to buy the same future cash flows. Textbook theory predicts that investors ought to reduce their holdings. Demand elasticity tells you how much. If \nu = 2, then a unilateral {+}1\% increase in the current price of equities will cause a {-}2\% reduction in investors’ stock holdings.

The security market line (SML) is much flatter than theory predicts. Instead of \lambda = 4\%, the slope of the SML is basically zero. The aggregate stock market is much less elastic than theory predicts. Instead of \nu = \frac{1}{4\%} = 25, Gabaix-Koijen puts the number in the neighborhood of 0.2. Investing an extra \mathdollar 1 in the stock market raises its aggregate value by about \mathdollar 5.

This note shows both findings are related. They’re two perspectives on the same underlying problem.

CARA-Normal Model

Start with the simplest possible model: two periods (today and next year); one investor; one risky asset (the stock market); one riskless bond. Buying a share of the stock market costs \mathrm{Price}_t today. If you own a share of the stock market today, then next year you are entitled to receive \mathrm{Payout}_{t+1}. The riskless bond costs \mathdollar 1 today and will pay (1 {+} \mathrm{rf}) next year.

The representative investor has “constant absolute risk aversion” (CARA) preferences, \mathrm{U}(\mathrm{C}) = {-}\tfrac{1}{\gamma} \cdot e^{-\gamma \cdot \mathrm{C}}. The stock market’s payout next year is normally distributed with variance \sigma^2 > 0. The investor starts with wealth \omega > \mathdollar 0. Today, he must choose how much to consume, \mathrm{C}_t, and how many shares of the risky asset to purchase, \mathrm{Q}_t. His goal is to maximize \mathrm{U}(\mathrm{C}_t) + \mathbb{E}\big[ \, e^{-\rho} \cdot \mathrm{U}(\mathrm{C}_{t+1}) \, \big] where \rho > 0 is his rate of time preference. The investor parks any remaining wealth, (\omega - [\mathrm{C}_t {+} \mathrm{Q}_t \cdot \mathrm{Price}_t]), in the riskfree bond. Next year, the investor eats the combined payout from his risky and safe investments

(3)   \begin{equation*}\mathrm{C}_{t+1} \;=\; (1 {+} \mathrm{rf}) \times \big( \, \omega - [\mathrm{C}_t {+} \mathrm{Q}_t \cdot \mathrm{Price}_t] \, \big) \,+\, \mathrm{Q}_t \cdot \mathrm{Payout}_{t+1}\end{equation*}

Let \psi > 0 denote the supply of shares in circulation. The market clears when the investor’s demand for the risky asset equals the number of available shares, \mathrm{Q}_t = \psi. An equilibrium is an allocation, \{\mathrm{C}_t,\,\mathrm{Q}_t,\,\mathrm{C}_{t+1} \}, and a current price level for the risky asset, \{ \mathrm{Price}_t \}, such that (i) the allocation solves the investor’s optimization problem given the price, and (ii) the price clear the market given the investor’s allocation.

The payout to owning each share of the risky asset is positive on average. So, holding an extra share will lead to slightly higher consumption next year. At the optimum, this benefit will be exactly canceled out by the cost of the required reduction in consumption today, with each side weighted by its marginal utility

(4)   \begin{equation*}\mathrm{U}'(\mathrm{C}_t) \times \mathrm{Price}_t \;=\; \mathbb{E}\big[ \, e^{-\rho} \cdot \mathrm{U}'(\mathrm{C}_{t+1}) \times \mathrm{Payout}_{t+1} \, \big]\end{equation*}

This is the Euler equation. An extra \mathdollar 1 that arrives in bad times (consumption is low; marginal utility is high) counts for more than a \mathdollar 1 that arrives in good times (high consumption; low marginal utility).

Here’s how to solve this model. First, note that the riskless asset costs \mathdollar 1 today and is guaranteed to deliver (1{+}\mathrm{rf}) next year, so its Euler equation is

(5)   \begin{equation*}\mathrm{U}'(\mathrm{C}_t) \;=\; \mathbb{E}\big[ \, e^{-\rho} \cdot \mathrm{U}'(\mathrm{C}_{t+1}) \times (1{+}\mathrm{rf}) \, \big]\end{equation*}

If we replace the \mathrm{U}'(\mathrm{C}_t) in Equation (4) with this expression, then the e^{-\rho} cancels out, and the price becomes a marginal-utility-weighted average of the discounted payout. The definition of a covariance plus Stein’s lemma turn that weighted average into \mathbb{E}[\mathrm{Payout}_{t+1}] - \gamma \times \mathbb{C}\mathrm{ov}[\mathrm{C}_{t+1}, \, \mathrm{Payout}_{t+1}]. What’s more, Equation (3) shows that next year’s consumption will be linear in the payout, so \mathbb{C}\mathrm{ov}[\mathrm{C}_{t+1}, \mathrm{Payout}_{t+1}] = \mathrm{Q}_t \cdot \sigma^2. Given market clearing, \mathrm{Q}_t = \psi, this leads to the following pricing rule

(6)   \begin{equation*}\mathrm{Price}_t = \frac{\mathbb{E}[\mathrm{Payout}_{t+1}] - \gamma \cdot \sigma^2 \cdot \psi}{1 + \mathrm{rf}}\end{equation*}

Each extra share makes next year’s consumption covary more strongly with the payout, so the marginal buyer demands a larger discount. The numerator is the expected payout minus an adjustment for risk. The denominator adjusts for the time cost of money.

Security Market Line

Textbook asset-pricing theory puts every asset on a single line. Expected excess returns ought to be proportional to betas, and the constant of proportionality ought to be the equity risk premium. To see where this prediction comes from, define the stochastic discount factor as discounted marginal utility growth, \mathrm{SDF}_{t+1} = e^{-\rho} \cdot \tfrac{\mathrm{U}'(\mathrm{C}_{t+1})}{\mathrm{U}'(\mathrm{C}_{t})}. Let n = 1, \ldots, N index the cross-section of risky assets… i.e., each stock in the stock market. The same SDF should price every one of them. If you use the SDF to write stock n‘s Euler equation and divide by its current price, then you get a statement about its expected return

(7)   \begin{equation*}1 \;=\; \mathbb{E}\bigg[ \, \mathrm{SDF}_{t+1} \times \underbrace{\bigg(\frac{\mathrm{Payout}_{n,t+1}}{\mathrm{Price}_{n,t}}\bigg)}_{1+\mathrm{Ret}_{n,t+1}} \, \bigg]\end{equation*}

Going forward, I’ll suppress time subscripts where it causes no confusion.

If you subtract the Euler equation for the riskless bond, 1 = \mathbb{E}[ \, \mathrm{SDF} \times (1{+}\mathrm{rf}) \, ], then you get

(8)   \begin{equation*}0 \;=\; \mathbb{E}\big[ \, \mathrm{SDF} \times (\mathrm{Ret}_n{-}\mathrm{rf}) \, \big]\end{equation*}

The difference being priced, (\mathrm{Ret}_n{-}\mathrm{rf}), is stock n‘s excess return. It is the payout from a long/short portfolio that sells riskfree bonds and uses the proceeds to buy shares of the risky asset.

Now consider applying the definition of a covariance, \mathbb{C}\mathrm{ov}[X, \, Y] = \mathbb{E}[X \cdot Y] - \mathbb{E}[X] \cdot \mathbb{E}[Y], to this excess-return SDF formula

(9)   \begin{align*}0 &= \mathbb{E}[ \, \mathrm{SDF} \times (\mathrm{Ret}_{n}{-}\mathrm{rf}) \, ] \\ &= \mathbb{E}[\,\mathrm{SDF}\,] \times (\mathbb{E}[\mathrm{Ret}_{n}]{-}\mathrm{rf}) + \mathbb{C}\mathrm{ov}[ \, \mathrm{SDF}, \, \mathrm{Ret}_{n} \, ]\end{align*}

By rearranging terms, we can arrive at the following expression

(10)   \begin{align*}\mathbb{E}[\mathrm{Ret}_{n}] {-} \mathrm{rf} &= \bigg( \frac{\mathbb{C}\mathrm{ov}[ -\mathrm{SDF}, \, \mathrm{Ret}_{n} ]}{\mathbb{E}[\mathrm{SDF}]} \bigg) \\ &= \underbrace{\bigg( \frac{\mathbb{C}\mathrm{ov}[ -\mathrm{SDF}, \, \mathrm{Ret}_{n} ]}{\mathbb{V}\mathrm{ar}[\mathrm{SDF}]} \bigg)}_{\beta_n} \times \underbrace{\bigg( \frac{\mathbb{V}\mathrm{ar}[ \mathrm{SDF}]}{\mathbb{E}[\mathrm{SDF}]} \bigg)}_{\lambda} \end{align*}

The first \beta_n term tells you how much asset n‘s return tends to comove with the SDF. The SDF captures growth in marginal utility. It is high when the economy enters into bad times. That’s when it becomes more valuable to have an extra dollar. Thus, stocks that tend to do well during booms and poorly during crashes will have large values of \beta_n. The second \lambda term is constant across stocks. It answers the following question: If a stock’s \beta_n goes up by one unit, how much higher will its excess returns be on average?

In the CARA-normal model, the SDF is approximately linear in the change in aggregate consumption

(11)   \begin{equation*}\mathrm{SDF} \;=\; e^{-\rho} \cdot e^{-\gamma \cdot \Delta \mathrm{C}} \;\approx\; a - b \cdot \Delta \mathrm{C} \end{equation*}

And what’s the main driver of the change in aggregate consumption in this model? The payout on the risky asset next year. Equation (3) shows that \mathrm{C}_{t+1} is linear in \mathrm{Payout}_{t+1}, and \mathrm{Payout}_{t+1} = (1 {+} \mathrm{Mkt}_{t+1}) \cdot \mathrm{Price}_t by definition. So the SDF is approximately linear in the market’s return.

Under these assumptions, you can estimate stock n‘s \beta_n by running a time-series regression of realized returns on the market return. Then, if you plot each stock’s average excess return against its estimated \beta_n, the slope of the best-fit line will give you \lambda. The stock market as a whole has \beta_{\mathrm{Mkt}} = 1 and an average excess return of \mathbb{E}[\mathrm{Mkt}]{-}\mathrm{rf} \approx 4\%. The riskfree bond has \beta_{\mathrm{rf}} = 0 and an average excess return of \mathrm{rf}{-}\mathrm{rf} \approx 0\%. Two points define the slope of a straight line. So textbook theory predicts that \lambda = \mathbb{E}[\mathrm{Mkt}]{-}\mathrm{rf} \approx 4\%.

The estimated slope is far below the equity risk premium. Way back in 1972, Black-Jensen-Scholes ran the test on every NYSE stock from 1926 to 1966, grouped into 10 beta-sorted portfolios. Average excess returns do line up with betas. But the fitted line is too flat, \hat{\lambda} \ll 4\%. Low-beta portfolios earn more than the model predicts, and high-beta portfolios earn less. The problem has only gotten worse. In 1992, Fama-French found basically no relation between average returns and betas from 1963 to 1990. Frazzini-Pedersen (2014) document the same issue in both US and global equities.

Demand Elasticity

The demand-system approach to asset pricing takes the same CARA-normal model but solves each investor’s problem before imposing market clearing. Let i = 1, \ldots, I index individual investors, each with his own risk-aversion coefficient, \gamma_i. Repeat the steps that led to the pricing rule in Equation (6), but stop short of market clearing. Isolating investor i‘s demand on the left-hand side, you get the formula below

(12)   \begin{equation*}\mathrm{Q}_i \;=\; \frac{\mathbb{E}[\mathrm{Payout}] - (1 {+} \mathrm{rf}) \cdot \mathrm{Price}}{\gamma_i \cdot \sigma^2}\end{equation*}

If you hold investor i‘s curve fixed and move the price, then the investor’s demand elasticity is given by

(13)   \begin{equation*}\nu_i = - \frac{\partial \log \mathrm{Q}_i}{\partial \log \mathrm{Price}} = \frac{(1 + \mathrm{rf}) \cdot \mathrm{Price}}{\gamma_i \cdot \sigma^2 \cdot \mathrm{Q}_i}\end{equation*}

Define the aggregate risk-aversion parameter, \gamma, as the harmonic average of the individual coefficients, \tfrac{1}{\gamma} = \sum_i \tfrac{1}{\gamma_i}. In equilibrium, each investor’s position in the CARA-normal model will be inversely proportional to his risk aversion, \gamma_i \cdot \mathrm{Q}_i = \gamma \cdot \psi. Thus, the denominator in the elasticity formula is the same for everyone, \nu_i = \nu. A single elasticity describes every investor. More risk-tolerant investors will hold bigger positions, but their percentage response is identical. This is analogous to the common \lambda across assets.

Notice that the denominator in the elasticity formula is just the risk discount in the CARA-normal model, \gamma \cdot \sigma^2 \cdot \psi = \mathbb{E}[\mathrm{Payout}] - (1 {+} \mathrm{rf}) \cdot \mathrm{Price}. Replace the denominator in Equation (15) with this expression and divide through by the current price. If the riskfree rate isn’t too large, then you get

(14)   \begin{equation*}\nu \;=\; \frac{(1 + \mathrm{rf}) \cdot \mathrm{Price}}{\mathbb{E}[\mathrm{Payout}] - (1 + \mathrm{rf}) \cdot \mathrm{Price}} \;\approx\; \frac{1}{\mathbb{E}[\mathrm{Ret}] {-} \mathrm{rf}} \end{equation*}

The reward for bearing a unit of stock-market risk is the equity risk premium, \mathbb{E}[\mathrm{Mkt}]{-}\mathrm{rf} \approx 4\%. A 1\% rise in the price will increase the cost of financing a share by 1\%, consuming roughly a quarter of the 4\% margin and causing demand to fall by \frac{1\%}{4\%} = 25\%. In other words, theory predicts that \nu = 25.

Deeper Connection

The slope of the SML, \lambda, is the exchange rate between risk and expected returns. How much higher must a stock’s expected excess return be in order to compensate investors for holding one more unit of exposure to market risk? One number common to every asset. The demand elasticity, \nu, is the exchange rate between flows and prices. How much do investors have to adjust their holding in response to a 1\% change in the price? One number common to every investor. Every assumption about preferences and beliefs reaches returns data only through \lambda, and reaches price-impact data only through \nu.

In one sense, these two parameters are two sides of the same coin. Neither is consistent with the observed equity risk premium, \mathbb{E}[\mathrm{Mkt}]{-}\mathrm{rf} \approx 4\%. This point estimate is enormous compared to the observed slope of the SML, which is basically zero. However, the same 4\% number implies a demand elasticity of 25, far above the value near 0.2 in the data. What’s more, the two predictions pull in opposite directions. Any effort that pushes \mathbb{E}[\mathrm{Mkt}]{-}\mathrm{rf} down to fit the slope of the SML makes the elasticity error worse and vice versa. The too-flat SML and the too-steep demand curve are both manifestations of the same underlying problem.

But there’s also a deeper connection. A flat SML is a trading opportunity. Buy levered positions in low-beta stocks and short high-beta stocks. Frazzini-Pedersen calls this trade “betting against beta”, and it has been profitable for decades. That sort of thing shouldn’t survive. Investors ought to pour capital into the trade, bidding up the prices of low-beta stocks and pushing down the prices of high-beta stocks until the SML steepened back to 4\%. That correction is a demand response to price, which is exactly what \nu measures. When demand barely responds to price, mispricings do not get traded away. They just sit there. So the too-steep demand curve is not merely a second manifestation of the same problem. It offers a reason why the first one never went away. The betting-against-beta alpha is what inelasticity looks like in returns data.

Filed Under: Uncategorized

Next Page »

Pages

  • Publications
  • Working Papers
  • Curriculum Vitae
  • Notebook
  • Courses

Copyright © 2026 · eleven40 Pro Theme on Genesis Framework · WordPress · Log in