Noise-State Game Explorer · risk aversion
Risk Aversion, Entropically
A risk-averse player, with a coefficient θ, minimises θ−1 log E eθC of its cost C: the certain cost it would trade its uncertain one for. It believes exactly what a risk-neutral player would. What changes is where it judges a deviation from its plan: at a point moved by an affine map of its beliefs. Here is why, and what follows.
Judging a deviation
Whether a small change to a plan pays is a derivative: the change in cost, averaged over the outcomes the player thinks possible. For a risk-neutral player that average is taken with its beliefs. For a risk-averse one, differentiating log E eθC gives the same average with every outcome weighted by eθC: the costly outcomes count for more. Two facts make that simple. The weight is the exponential of a quadratic, so weighted Gaussian beliefs are Gaussian again, with a moved centre. And the change in cost from a small deviation is a straight line in the shocks, and the average of a straight line is the line at the average. So the risk-averse judgment is the risk-neutral one, taken at the moved centre.
Here the belief is w ~ N(μ, 0.8²), the cost C = ½w² + ½w, so the weighted centre is (μ + θ·0.8²·½) / (1 − θ·0.8²), and θ* = 1/0.8². With many shocks the stretch is (I − θΣK)−1, Σ the uncertainty and K the cost’s quadratic part, and the shift comes from its linear part: the appendix to Chapter 1.
A past shock and a future one
With more than one shock the moved centre can land where the belief puts nothing. Take a shock that has happened, seen through noise, and one still to come, with the cost the square of their sum. The belief about the future shock is centred at zero; the weighted one is not, because the costly outcomes are the ones where both push the same way. A risk-neutral player can ignore what it has not seen yet. A risk-averse one cannot: certainty equivalence fails exactly here.
In the tracking game
The same for player 1 of Chapter 1’s tracking game, halfway through, having seen the state pushed up. Its noise-state is its belief about every shock, past and future. Its first-order condition is the risk-neutral one evaluated at the risk-adjusted noise-state, Ŵθ = (I − θΣK)−1Ŵ (no shift here: the game has no targets), where Σ is how uncertain player 1 still is about the shocks and K is how they enter its cost. Like an adjoint, it is a point at which a derivative is taken, not a belief.
Because the map is affine, the first-order system stays a linear system in noise-states, with the same filters: CARA fits the noise-state method without changing its machinery. Past θ*, when I − θΣ1/2KΣ1/2 stops being positive, the weighting has no average. These drawings hold the strategies at the risk-neutral equilibrium’s; the equilibrium section solves for risk-averse ones. Computed on 100 time cells; the belief matches player 1’s own filter in noisestate to within the cells.
Weight, then look; or look, then weight
Chapter 1’s appendix and the noisestate solver compute this weighting in different orders. The appendix looks first: each time the player sees something, it takes the player’s updated belief and weights it by eθC, so the weighting is redone at every moment. The solver weights first: it weights every outcome of the whole game once, then conditions on what the player has seen. Could the two disagree? They cannot: by Bayes’ rule the two orders give the same weighted curve, so the derivative that decides whether a plan is best comes out the same either way. Here are two shocks, one the player has seen and one still to come, correlated, with cost C = ½(seen + to come)².
Weight, then lookthe solver: once, for the whole game
weight every outcome by eθC
condition on what was seen (read along the line)
Look, then weightthe appendix: redone at each moment
condition on what was seen (read along the line)
weight by eθC
An equilibrium of risk-averse players
Now let both players of the tracking game be risk averse, and let each best-respond to the other. Every player then judges its plan at its risk-adjusted noise-state, and so the kernels, the adjoints and the costs that define each player’s adjustment all come from the risk-averse players themselves: one fixed point, not a correction to the risk-neutral one.
Solved by noisestate, in its development version, and checked against a brute-force solver.
A single controller
One controller, alone, steering a state that shocks keep pushing: dX = D dt + dW, at a cost X² + 0.1 D² per unit of time. The risk-neutral controller pushes back with gain √10 ≈ 3.2. A risk-averse one pushes back harder, with gain √10 / √(1−0.2θ): it pays more for control to keep the state from wandering into a bad run. At θ* = 5 no gain is enough.
The gain is the risk-sensitive regulator’s closed form; minimising the exact entropic cost rate of the controlled process over the gain gives the same number at every θ.