Noise-State Game Explorer · risk aversion

Risk Aversion, Entropically

A risk-averse player, with a coefficient θ, minimises θ−1 log E eθC of its cost C: the certain cost it would trade its uncertain one for. It believes exactly what a risk-neutral player would. What changes is where it judges a deviation from its plan: at a point moved by an affine map of its beliefs. Here is why, and what follows.

Judging a deviation

Whether a small change to a plan pays is a derivative: the change in cost, averaged over the outcomes the player thinks possible. For a risk-neutral player that average is taken with its beliefs. For a risk-averse one, differentiating log E eθC gives the same average with every outcome weighted by eθC: the costly outcomes count for more. Two facts make that simple. The weight is the exponential of a quadratic, so weighted Gaussian beliefs are Gaussian again, with a moved centre. And the change in cost from a small deviation is a straight line in the shocks, and the average of a straight line is the line at the average. So the risk-averse judgment is the risk-neutral one, taken at the moved centre.

One shock w. Grey: the player’s belief about it. Orange: the weight eθC on a log scale, where it is the parabola θC, the cost C being quadratic in w. Dashed blue: the belief times the weight, again a bell curve, shifted and widened. It is not a new belief: the player still believes the grey curve, and the blue one only says how much each outcome counts in the average that judges a deviation. Below: the change in cost from a small deviation, a straight line in w. Its average under each curve (computed by summing over w) equals the line at that curve’s centre.
Where the judgment is taken, against what the player believes. Wherever the belief sits, the judgment point moves by the same stretch and the same shift: an affine map, steeper as θ grows. At θ* the weighted curve stops being a bell (it no longer has an average) and the map breaks down.

Here the belief is w ~ N(μ, 0.8²), the cost C = ½w² + ½w, so the weighted centre is (μ + θ·0.8²·½) / (1 − θ·0.8²), and θ* = 1/0.8². With many shocks the stretch is (I − θΣK)−1, Σ the uncertainty and K the cost’s quadratic part, and the shift comes from its linear part: the appendix to Chapter 1.

A past shock and a future one

With more than one shock the moved centre can land where the belief puts nothing. Take a shock that has happened, seen through noise, and one still to come, with the cost the square of their sum. The belief about the future shock is centred at zero; the weighted one is not, because the costly outcomes are the ones where both push the same way. A risk-neutral player can ignore what it has not seen yet. A risk-averse one cannot: certainty equivalence fails exactly here.

Grey: what the player believes, the past shock about 0.6 and the next one anybody’s guess. Dashed blue: the same belief with every outcome weighted by eθC. It leans along the costly diagonal, and its hollow centre, the point where a deviation is judged, moves off zero on the future shock’s axis. The player does not expect the next shock to push; it weighs the outcomes where it does.

In the tracking game

The same for player 1 of Chapter 1’s tracking game, halfway through, having seen the state pushed up. Its noise-state is its belief about every shock, past and future. Its first-order condition is the risk-neutral one evaluated at the risk-adjusted noise-state, Ŵθ = (I − θΣK)−1Ŵ (no shift here: the game has no targets), where Σ is how uncertain player 1 still is about the shocks and K is how they enter its cost. Like an adjoint, it is a point at which a derivative is taken, not a belief.

Grey: what player 1 believes, its estimate of the state so far and its forecast of the rest; the thin dark line is what actually happened. Dashed blue: the same path averaged over outcomes weighted by their cost, where the risk-averse player judges a change in its plan. It is neither an estimate nor a forecast: it leans up, past and future, because the outcomes where the state ran high are the ones that count most.
The judgment itself: how much player 1’s cost changes if, now, it pushes a little harder against the state. With its belief the answer is zero (up to the time cells), as it must be: the plan is optimal for a risk-neutral player. Cost-weighted, the push lowers the cost, more so as θ grows: a risk-averse player would push harder. That is where the equilibrium below goes.
Shock by shock: player 1’s estimate of each shock path (grey) and the same averaged over outcomes weighted by their cost (dashed blue, not an estimate). Every future shock averages zero in the belief, flat after the line; weighted by cost, the future shocks average a drift that pushes the state further, and the past is weighted toward the bigger pushes. Thin dark line: what actually happened.

Because the map is affine, the first-order system stays a linear system in noise-states, with the same filters: CARA fits the noise-state method without changing its machinery. Past θ*, when I − θΣ1/2KΣ1/2 stops being positive, the weighting has no average. These drawings hold the strategies at the risk-neutral equilibrium’s; the equilibrium section solves for risk-averse ones. Computed on 100 time cells; the belief matches player 1’s own filter in noisestate to within the cells.

Weight, then look; or look, then weight

Chapter 1’s appendix and the noisestate solver compute this weighting in different orders. The appendix looks first: each time the player sees something, it takes the player’s updated belief and weights it by eθC, so the weighting is redone at every moment. The solver weights first: it weights every outcome of the whole game once, then conditions on what the player has seen. Could the two disagree? They cannot: by Bayes’ rule the two orders give the same weighted curve, so the derivative that decides whether a plan is best comes out the same either way. Here are two shocks, one the player has seen and one still to come, correlated, with cost C = ½(seen + to come)².

Weight, then lookthe solver: once, for the whole game

prior over both shocks

weight every outcome by eθC

weighted (not a belief)

condition on what was seen (read along the line)

weighted curve (not a belief)

Look, then weightthe appendix: redone at each moment

prior over both shocks

condition on what was seen (read along the line)

the player’s belief

weight by eθC

weighted curve, and the row above’s
In every panel the shock still to come runs across. In the squares the shock already seen runs up, and the dark line is its value; conditioning on it reads the square along that line. The two shocks are correlated, so what was seen moves the player’s belief. The last panel draws both routes’ curves: they are one curve. Its centre (the circle) is the point at which a derivative of the cost is taken, the same either way. In symbols: eθC(x,y)p(x, y) read at x = x0 is eθC(x0,y)p(y | x0), up to a constant. The prior is standard normal with correlation 0.6, and θ* = 1/3.2; near it the weighting spreads along the costly diagonal.

An equilibrium of risk-averse players

Now let both players of the tracking game be risk averse, and let each best-respond to the other. Every player then judges its plan at its risk-adjusted noise-state, and so the kernels, the adjoints and the costs that define each player’s adjustment all come from the risk-averse players themselves: one fixed point, not a correction to the risk-neutral one.

The life of a state shock at t = 0.2: the state, and player 1’s push back. Risk-averse players push back harder, and the shock dies sooner.
Each player’s entropic cost rises with θ here. Its expected cost first falls, a little, to its lowest near θ = 0.9: when both push harder, both gain on average, which suggests the risk-neutral players each leave some of the pushing to the other. Far enough out, they over-insure, and at θ ≈ 3.04 the equilibrium breaks down. That is later than the tracking game’s θ* = 1.93: here the players push harder, damp the shocks, and so can bear more risk aversion.

Solved by noisestate, in its development version, and checked against a brute-force solver.

A single controller

One controller, alone, steering a state that shocks keep pushing: dX = D dt + dW, at a cost X² + 0.1 D² per unit of time. The risk-neutral controller pushes back with gain √10 ≈ 3.2. A risk-averse one pushes back harder, with gain √10 / √(1−0.2θ): it pays more for control to keep the state from wandering into a bad run. At θ* = 5 no gain is enough.

The gain against θ, with the breakdown at θ* = 5.
The state under both controllers, driven by the same shocks, with the band each keeps it in most of the time (two standard deviations): the risk-averse one’s is tighter.

The gain is the risk-sensitive regulator’s closed form; minimising the exact entropic cost rate of the controlled process over the gain gives the same number at every θ.