One controller, one hidden state. Estimate the state, then act on the estimate exactly as you would on the state itself. How you learn and how you act are separate problems.
The separation principle. Simple, and its intuition has held back the multi-player version for decades.
Add a second player who watches the same state through a signal of their own. Now every push you give the state also shows up in what they see. Your action does something and says something at once.
They revise their forecast, change what they do, and that moves the state again.
Optimal control prices a move of the state with a shadow price: how much better the future looks if the state is a little higher now. The best action sets the cost of effort against that price.
In a game there is a second thing a push moves: the other player’s noise-state. The information wedge is its price, the value of changing an opponent’s beliefs. It is one more term in the backward equation for the shadow price.
All the strategic complexity sits in deterministic kernels, solved once at equilibrium.
Here it is, computed. Player 1’s first-order condition at mid-horizon, split into its physical part and the wedge, for each shock it responds to.
The physical part is what the condition would be if nobody reacted to player 1’s deviation. The wedge is the rest: player 2 revising their forecast.
Now cut the loop. Blur player 2’s signal until they see nothing. Player 1’s pushes teach them nothing, and the wedge vanishes. Separation comes back.
The wedge breaks separation, and vanishes when the loop is cut.
What the wedge buys. A planner has a fixed budget of signal precision to split between the two players, and player 2 finds effort costlier. It looks like a tradeoff: favor the productive player, or balance the precision to keep both restrained.
It is not. Give it all to player 1 and player 2, starved of information to manipulate with, stops pulling against them. The total cost falls the whole way.