Chapter 1Baseline Linear-Quadratic-Gaussian Games
One player’s action changes what the others learn, their response feeds back into the state, and the first player must price it. The noise-state represents each player’s beliefs about the shocks that generate histories, so higher-order beliefs are generated by composition rather than tracked as separate state variables. In continuous-time LQG, beliefs, prices, and policy kernels are deterministic impulse-response kernels, and equilibrium is a deterministic fixed point in those functions. The adjoint system contains a shadow price of changing opponents’ posteriors, the information wedge; in a two-player benchmark it separates the statistical and strategic value of information and turns signal allocation into an instrument for controlling inefficient effort.
1.1Introduction
Dynamic games with dispersed private information pervade economics. Firms infer competitors’ pricing from market outcomes, traders learn from the order flow their own trades generate, central banks know their announcements reshape private agents’ signals. In each setting, actions feed back into the information environment that agents optimize against, generating an infinite hierarchy of beliefs. Townsend [109] identified that hierarchy, and Sargent [104] computed equilibria in linear-Gaussian environments by positing low-dimensional ARMA laws rather than deriving them.
Beliefs here are beliefs about the underlying sources of randomness rather than hierarchies of beliefs about the endogenous state. A researcher who specifies payoffs, a state equation, and a signal structure gets back equilibrium impulse-response functions; the representation handles the belief hierarchy internally.
Call this representation the noise-state.
In a game of poker, the common approach describes each player’s beliefs about opponents’ hands, their beliefs about opponents’ beliefs, and so on. The noise-state approach instead tracks each player’s belief about the underlying source of randomness (here, the order the cards were shuffled) and how that belief depends on the truth; higher-order beliefs follow by composing these dependencies. The two representations are equivalent, but approximating a composition of deterministic dependencies is natural, whereas truncating a belief hierarchy is not.
A deviation enters both the physical state and opponents’ noise-state coordinates. The Markov state consists of the physical state, the players’ noise-states, and, when needed, the blueprints that translate histories into beliefs.
The construction extends beyond the LQG setting studied here, but the LQG benchmark is the clean case in which it becomes explicitly solvable, since the kernel form relies on conditional Gaussian filtering. It is also the base case for a broader noise-state program in conditional-Gaussian linear economies, including strategic trading, CARA risk aversion, and replicated finite symmetric networks in which local strategic feedback can affect aggregate response coefficients.
Characterizing equilibrium produces a shadow price of manipulation, the information wedge . Every effect of belief manipulation enters the dynamics of the state price (Theorem 1.8) through that single term. In the two-player benchmark, the wedge explains why most gains from pooling information are strategic rather than purely statistical, and why optimal precision allocation can generate information starvation by concentrating signals on the player who moves the state more cheaply, leaving the other player with none. Changing signal precision changes equilibrium policy rules, not only posterior variances, so separation fails. As a diagnostic, Corollary 1.9 shows the wedge vanishes when observation laws are fixed, recovering the standard fixed-information LQG optimality system. Section 1.2 works these economic consequences out in a two-player game, and Section 1.4 builds the general theory. Proposition 1.12 gives a unique unilateral best response for every finite horizon. On short horizons, Theorem A.7 gives a unique equilibrium on a strategy ball, with geometric Picard convergence. For an arbitrary horizon, existence and uniqueness on a strategy ball hold under a weighted monotonicity condition (Theorem A.9) that Corollary A.11 can certify numerically.
1.2Bilateral belief manipulation in a two-player game
A two-player tracking game isolates the local feedback channel (Figure 1.1). Every figure in this section, and the solver behind it, runs in the browser at sbabichenko.com/noisestate with adjustable parameters.
1.2.1Primitives and objectives
Fix a horizon and a filtered probability space supporting a standard three-dimensional Brownian motion , where is the completed, right-continuous natural filtration of . The state is the real-valued, -adapted process satisfying where is the fundamental volatility. Player does not observe directly; instead it observes the private diffusion channel The numbers are signal precisions. Player chooses using only its private observation history. Formally, is the completed, right-continuous augmentation of , and an admissible control is -progressively measurable and square-integrable, . The two controls are adapted to different filtrations, to and to , and in general neither filtration contains the other.
Running costs are symmetric tracking losses with quadratic effort penalty : Player 1 wants the state near ; player 2 wants it near .
Each player minimizes its own cost, taking the opponent’s strategy as given. A pair of admissible strategies is a Nash equilibrium if for every admissible , and symmetrically for player 2. Section 1.3.3 restates the definition for the -player game.
1.2.2Impulse responses and wedges
Equilibrium controls belong to the noise-state linear class (Definition 1.4) and can be written in two equivalent impulse-response coordinates, with the noise-state defined on the middle line: The first line is the strategy as the player implements it: a deterministic mean plus a reading of the noise-state. The third line is the same control composed with the player’s filter and written directly over the primitive shocks. , the policy kernel, weights estimates measurable to the player, while describes how the choice loads on the truth and involves every player’s filter.
Proposition 1.1 (Best responses and equilibrium in the two-player game). For every finite horizon and every fixed linear strategy of the opponent, each player has a unique best response over the full admissible class, and that response is noise-state linear. There is such that for the game admits a noise-state linear equilibrium, unique in the strategy ball of Theorem A.7, Picard iteration converges to it geometrically, and the equilibrium is Nash over the full admissible class. For an arbitrary horizon, any strategy ball satisfying the weighted monotonicity and residual-radius conditions of Theorem A.9 contains a unique fixed point, and that fixed point is a Nash equilibrium over the full admissible class.
1.2.3The life of a spike deviation
A spike by one player propagates through the loop of Figure 1.1. Player 1 adds an impulse to its control. The state shifts by and stays shifted. Player 2’s innovation carries the part of the shift player 2 has not yet learned, , and each coordinate of its noise-state moves through the unresolved response, , the part of the shock’s effect player 2 has not yet resolved, with the coordinate born at equal to . Player 2’s control moves with its coordinates and the state with the control. Player 2’s target is , so the response opposes the shift.
Player 1 pays for the shift and for the response. The shift is priced at the running cost. The response is priced through player 2’s noise-state, one coordinate at a time: of (1.4.8), the belief price, is the marginal cost to player 1 of having moved player 2’s estimate of the shock at . Setting the variation to zero for every -measurable ,
which is (1.4.7) integrated, with the wedge (1.4.9) at the example’s coefficients. Without player 2 only the running-cost term remains; is the change in player 2’s beliefs times their price. Theorem 1.8 is the same calculation for players, a vector state, and a general observation map.
Player 2 sees the state only through its signal, so it cannot tell a unit spike by player 1 from a fundamental shock of size . At , player 2’s curves in Figure 1.2 therefore show how player 2 responds to a unit spike by player 1. The left panel traces a single state shock through the actual state and through each player’s posterior estimate of it, each a linear readout of the game’s Markov state (the physical state together with every player’s noise-state; Section 1.3). The shock moves the state. Each player’s estimate catches up at a speed set by its signal precision. The right panel traces the same shock through the players’ actions. Each player responds to the shock through what it has learned, the better-informed player earlier and more strongly. Both control responses die out at the horizon, when nothing remains to be gained from moving the state. These plots can be redrawn at any parameters in the browser version.
1.2.4Computation
The unknowns of the fixed point are the kernels , , and ; the adjoint equations (1.4.7)–(1.4.8) are solved for these kernels rather than for a value function on , because each player’s response depends on every opponent’s filter. Equilibrium is computed as a fixed point of the simultaneous best-response map on the deterministic causal policy kernels, discretized on a uniform grid of points; no Monte Carlo simulation is used. Starting from zero policy kernels, each outer iteration first solves the forward state–filter system generated by the current profile, then solves both players’ backward physical and belief-price equations, and finally applies the first-order conditions to obtain a new pair of policy kernels. Both new kernels come from the same old profile, so the update is simultaneous rather than alternating.
Write for this discretized best-response map and . Convergence is measured by the normalized Frobenius residual over every time pair on the discrete causal triangle and every primitive-shock coordinate, The first update is a relaxed Picard step with relaxation parameter ; subsequent updates use depth-five Anderson acceleration with mixing parameter . The algorithm stops when the residual falls below . Within each outer iteration, the coupled forward state–filter calculation uses two relaxed updates of the state impulse response and six filter sub-iterations. After the policy kernels converge, a separate relaxed fixed-point iteration, initialized at zero and solved to tolerance , gives the deterministic mean controls and mean state. Parameter sweeps may be warm-started from the equilibrium at a neighboring parameter value.
Unless otherwise stated, , , , , and there is no terminal cost. In an optimized native C++ build with OpenMP, the symmetric benchmark at converges in outer iterations and takes about seconds; at the same benchmark takes about seconds.
The scheme is first order in the grid spacing. On the symmetric benchmark with targets , player 1’s cost is , , and at , , and , an observed order of ; the policy kernels converge at the same rate. A single run is therefore accurate to about 3 percent on costs and mean actions, and considerably worse on quantities small early in the horizon, such as the posterior estimates just after a shock and the information wedges. Unless a caption says spectral solution, figures and reported values below are Richardson extrapolations from two nested grids, , which agree with the extrapolation from and to within 0.3 percent except at the first two grid points of the wedge curves.
The solvers and figures of this section are the work of Claude Fable 5, a large language model built by Anthropic [7], working under the author’s direction. The interactive version at sbabichenko.com/noisestate runs the same solver.
1.2.5Mean actions, separation failure, and the welfare cost of belief manipulation
Since the stochastic integral in the first line of (1.2.4) has zero unconditional mean, isolates the deterministic tug-of-war over the target , while the response term captures impulse responses to posterior shocks.
Separation failure.
Under perfect information, the mean policy is independent of signal precision. Decentralized information breaks this. Changing signal precision changes how much of each shock remains unresolved for the opponent, which changes the belief price, so the policy rule moves independently of the state estimate (Remark 1.13). The mechanism is the information wedge.
Figure 1.3 plots the resulting failure. The equilibrium mean control path shifts with the signal precision , and the partial-information paths lie between two limits independent of . With no information the players cannot see each other and the mean paths are the open-loop Nash equilibrium of the deterministic game, . With perfect information they are its closed-loop Nash equilibrium, in which each player’s feedback on the state is priced by the other. As rises the mean control moves monotonically from the first to the second, and at it is still ten percent above the closed-loop path, and at seven percent. Under a separation principle the colored paths would sit on the closed-loop line at every precision. With , opposite targets, and , the two mean controls cancel and at every precision. The movement in Figure 1.3 is effort spent holding a path that does not change. It vanishes in models where information and incentives can be treated separately. Because the shift is in the mean rather than in the zero-mean response term, it does not cancel when many independent copies of the market are averaged; Chapter 5 works this out.
Information pooling as a corrective intervention.
Figure 1.4 compares private signals with a pooled signal of precision under opposing targets and and a common target . At and pooling is Pareto-improving at every , and the gain is an order of magnitude larger under competition. Since the target change affects only the deterministic mean system, the common-target case isolates the pure estimation benefit. Each player still pays its own effort cost, so this case is a Nash game rather than the team problem of Section 1.4.1. The excess gain under competition is the welfare cost of bilateral belief manipulation.
The private-signal curves also show who wants the precision. Under the common target the better-informed player has the higher cost. At player 2 pays against player 1’s . The better-informed player does more of the work, so each player would rather be the less precise one. Under opposing targets the ordering reverses, against , and precision is an advantage over the opponent. In absolute terms, both players gain when either one is better informed, including player 1, whose own precision is fixed and whose opponent improves. The competitive case does not change that even though it raises every cost by an order of magnitude. At the pooled cost and player 2’s private cost have nearly met, since a player who observes the state almost exactly has little left to learn from a pooled signal. With stronger conflict they cross. The policy kernels do not depend on the targets, and the mean paths are linear in them, so at targets the cost is , with the cost at the common target and the mean part at . Pooling still lowers player 2’s variance cost, but it raises its mean cost, because player 2 gives up the mean-path advantage of being the better informed. Above the second effect wins. Pooling then costs player 2 ( at ) while player 1 gains . Under strong conflict the better-informed player prefers to keep its edge. Under the common target player 1’s private cost meets the pooled cost too ( against ), so an almost perfectly informed teammate is nearly as good as a shared signal. Under opposing targets it stays well above ( against ). An almost perfectly informed opponent is not.
1.2.6Optimal precision allocation
Information pooling is blunt. A planner with a fixed precision budget must still choose how to split it between the two players.
Setup.
Retain the state dynamics (1.2.1) and observation channels (1.2.2) with opposing targets (), but allow asymmetric precisions subject to and asymmetric effort costs with , so player 1 is the more efficient mover of the state. A planner chooses the split to minimize aggregate equilibrium cost; for each candidate allocation the solver of Section 1.2.4 computes the full equilibrium. Parameters: , , with the fundamental volatility in (1.2.1) reduced to . Because the players have opposing targets, their mean controls push in opposite directions rather than jointly stabilizing the state. Total destructive effort measures the intensity of this arms race,
Information starvation.
Under cooperation, a planner should assign precision to the efficient mover. Under competition, the natural guess is instead that a planner should balance precision to limit manipulation. Figure 1.5 shows the opposite. With asymmetric effort costs , concentrating precision on the efficient player shifts the mean state toward that player’s target and sharply reduces destructive effort. On net, both players’ costs fall as destructive effort moves toward the perfect-information benchmark.
Denying the inefficient player signal access weakens the bilateral arms race: commitment in the sense of Schelling [105]. Giving all precision to the inefficient player produces the worst outcome, because the efficient player pushes hard against a well-informed opponent. The planner’s optimum is not balanced precision but an information monopoly for the efficient player. The right panel shows separation failure again. Certainty-equivalent controls are invariant to the split, while equilibrium controls move with the belief prices.
1.3The decentralized LQG game
Informal description of the game.
The games here have a common evolving state, private noisy observation channels, and controls that feed back into the state and so into others’ signals. These features generate the Townsend hierarchy of forecasting the forecasts of others [104, 109]; quadratic objectives and linear-Gaussian primitives make the problem tractable. The section generalizes Section 1.2: the same noise-state coordinates, filters, and prices, now with players, vector states, and general linear dynamics.
1.3.1State dynamics and objectives
Consider players interacting over a finite horizon on a filtered probability space .
State dynamics.
The state evolves as where is deterministic and bounded, is a -dimensional Brownian motion, is deterministic and bounded, and is player ’s control, entering through the deterministic, bounded gain (the identity in the examples).
Objectives.
Let denote the linear state-cost coefficient. In the baseline case it is deterministic and bounded.1 Each player incurs the cost
where the joint running Hessian , with , is symmetric positive semidefinite, and are positive semidefinite, and denotes opponents’ controls. Control coercivity needs the Schur complement rather than , which permits the form . Semidefiniteness gives , so the pseudo-inverse term is exact. The cross term lets the price of the control depend on the state (an input bill paid at a market price, a trading profit ); is the tracking case, where . Each player seeks to minimize Opponents’ controls enter player ’s cost only through , which they move through (1.3.1) and which player observes through (1.3.5).
1.3.2Information structure
Observations.
Player does not observe directly. Instead, its private information comes through the diffusion signal
where is a deterministic observation gain and is the block selector extracting player ’s noise channel (). All Brownian motions are mutually independent. The precision matrix measures the Fisher information about the state per unit time. The block-selector form is without loss, since replacing by a general invertible noise loading substitutes for everywhere. Throughout the main characterization, the gain is fixed; the two-player benchmark studies comparative statics and planner-chosen precision.
Notation for primitive shocks.
Collect all primitive shocks into a single vector . The direct projection is the identity on the -th block and zero elsewhere. The block-selector structure ensures each player’s observation noise is an independent component of . Player observes directly up to a drift term and infers from drift-based learning (Theorem 1.5).
Player ’s information at time is its observation history, . An admissible control is -adapted with , where denotes the Euclidean norm on .
Unresolved response.
For each player , define the component of the state’s impulse response to the -th shock that player has not yet resolved at time . It is large when the shock is recent or the observation channel weak, and it shrinks as observations accumulate. Formally, ; it is the filtering gain in the noise-state update (Theorem 1.5), with its explicit formula in (1.A.4).
1.3.3Nash equilibrium
Each player’s drift depends on opponents’ estimates . The following equilibrium concept fixes opponents’ strategy maps rather than their realized actions; Section 1.4 needs this to show that a best response to noise-state linear opponents is itself noise-state linear.
Definition 1.2 (Nash equilibrium in private-signal strategies). Throughout, a player’s choice is a strategy map: for each , the action is a (progressively measurable) functional of the private observation history . Write , where the bracket notation denotes the map from the observation path to the action at time . Admissible strategy maps must satisfy and must also be causally well posed: against any fixed profile of the opponents’ complete causal maps, the coupled state–observation–strategy equations admit a unique strong causal solution, adapted to the control-free observation filtration. Strategies use no independent randomization; any private randomization, if admitted, is included in the player’s filtration.
A profile is a Nash equilibrium if for every player and every admissible alternative strategy map ,
Holding opponents’ strategy maps fixed is substantive because signals are endogenous. A unilateral deviation changes the drifts of opponents’ observation paths, and with them the likelihoods they assign to histories. Under admissible square-integrable controls this only changes drifts, so the induced observation laws are mutually absolutely continuous on every finite horizon. A perfect-Bayesian formulation is simpler here than in discrete signal games, because Bayes’ rule applies on a common support and the fixed strategy maps are evaluated on whatever observation paths occur.
Noise-state.
On a given profile, the noise-state is the conditional mean path of the aggregated primitive shock: Against fixed opponent maps, player ’s realized control is a known input. Proposition 1.6 shows that its full downstream effect on the state, on opponents’ responses, and so on is a deterministic causal functional of the realized control path. Subtracting that effect gives a control-free Gaussian observation with the same filtration. Off the candidate profile, then denotes the conditional projection of in this equivalent reference coordinate. The deviating player does not choose a filter. The actual conditional state estimate is the reference estimate plus the known state response, while the innovation and its covariance are unchanged. On the candidate profile this agrees with (1.3.6).
Definition 1.3 ((Primitive-shock) linear process). An stochastic process is a (primitive-shock) linear process if there exist deterministic functions such that, for every , When convenient, extend by for .
Definition 1.4 (Noise-state linear control / strategy). An admissible strategy map is noise-state linear if, for the fixed filter described above, there exist deterministic functions such that, for every private history and every , Extend by for and drop the argument on the equilibrium path.
1.4General theory
The primitive-shock space plays the role of a dynamic Harsanyi type space [52]: conditional beliefs are projections of onto private signal histories. Higher-order beliefs are compositions of these projections. Player ’s estimate of player ’s estimate of the state is ; since is a linear functional of with the deterministic state kernel , its projection onto is the same functional applied to , which is (1.4.1) for player evaluated on player ’s noise-state. Every order of belief is therefore a finite composition of the kernels , and no new state variable is needed. Actions reshape the signals from which beliefs are formed, so those maps must be solved as a fixed point (Figure 1.6).
Beliefs as deterministic kernels.
The noise-state of (1.3.6) has a deterministic-kernel representation for its estimated-noise increments on the causal triangle , extended by zero off . Write for increments in the path index (the source time ) and for updates in time. The state has the form , with the impulse-response kernel. Appendix 1.A gives the general linear-filtering result; here .
Theorem 1.5 (Beliefs are deterministic kernels: Filtering Closure). Fix a player . Assume the state admits the primitive-shock linear form and player observes with deterministic and block selector .
Let and , and define the innovation .
(i) The estimated-noise path increments admit the decomposition where the deterministic kernel is given by The pair is the blueprint of player ’s beliefs.
(ii) The mixed -update satisfies Each old coordinate moves with the innovation, and is the gain. The singular term is the coordinate’s birth, its entry into the filter at the current date.
The unresolved response is determined algebraically from by (1.A.4); the filter kernel is the unique solution of the forward evolution system (1.A.5)–(1.A.6) (Definition 1.16).
Proposition 1.6 (Against frozen opponents, the deviating player’s own control is a known input to its filter). Fix player and the opponents’ complete noise-state linear strategy maps, including their filters. Let and be any two admissible controls generated by causally well-posed strategy maps, and write The corresponding differences in the state, opponents’ noise-states, and opponents’ controls satisfy the exact deterministic causal system stated in (1.B.1)–(1.B.3). That system is linear in the realized input process even when the feedback rule generating is nonlinear. No small-deviation approximation is used.
The own-observation difference is and is a deterministic causal functional of . Taking , define the control-free (reference) observation by subtracting this known response from the actual observation. Causal well-posedness gives an invertible causal correspondence between the actual and control-free histories, so they generate the same completed filtration. If and denote the control-free state and observation, then is the same standard Brownian innovation under every admissible strategy of player . The deviating player’s filter is then the conditional filter implied by the frozen environment and the known input.
Lemma 1.7 (Noise-state integrals as innovation integrals). Fix and let be deterministic and square-integrable. In the reference innovation coordinate of Proposition 1.6, The closed linear span of the coordinates contains every -measurable affine Gaussian random variable, so each has a deterministic noise-state representation. Equation (1.4.5) gives the corresponding innovation coefficient.
Proof. Integrating the mixed update in Theorem 1.5 from the birth time to gives, as a measure in , Substitution followed by stochastic Fubini gives (1.4.5). For the span claim, conditional expectation is the orthogonal projection of linear functionals of the primitive shock onto the -measurable ones. The projected primitive increments then generate precisely that subspace. ◻
Best-response verification and information wedges.
The maximum-principle calculation identifies deterministic adjoint kernels and the information wedge . The exact finite-deviation identity below then verifies the resulting candidate against the full admissible class.
Theorem 1.8 (Belief prices and information wedges). Fix a player and suppose each opponent uses a noise-state linear strategy with deterministic impulse-response maps, which pins down a deterministic forward environment (state impulse response , filtering objects , and policy kernels ). There exist linear processes and , , such that the following system holds for , with (1.4.8) holding for old-history coordinates . For each opponent , define The state price and belief prices satisfy with terminal conditions and . All coefficients, , , , and , are determined by opponents’ frozen strategy maps and do not change when player deviates.
Proof. Appendix 1.B.2: the spike variation, the opponent filter variation, the coupled state–filter system, and the backward adjoints. The state price is the marginal cost of a spike that moves the state at , and the marginal cost of one that moves opponent ’s estimate of the shock at . Every channel through which player ’s action moves what opponents believe, and every consequence of those moved beliefs, enters the dynamics of the state price only through the term in (1.4.7). To write explicitly, define the filter contraction of a kernel , , and its transpose action on a row-valued kernel by with . Then . The contraction of (1.4.9) is opponent ’s filter applied to a belief price: the birth term for the shock arriving now, plus the integral of against the innovation kernel for the shocks already in the history. Chapters 2, 4, and 6 reuse it with replaced by each chapter’s observer pair. In (1.4.8), is what player pays, at the state price, for the action opponent takes on a moved coordinate.
For a fixed profile, the continuation value is a functional of the physical state and the players’ noise-states, with adjoints and . All coefficients in (1.4.7)–(1.4.8) are deterministic, so the mean and response-map coefficients decouple.
Corollary 1.9 (No observational externality). Suppose the state components that opponents observe are unaffected by player ’s control. Then a unilateral deviation in leaves every opponent’s observation law, and hence the fixed filter for each , , unchanged; , the belief-price equations (1.4.8) decouple from the state price, and the optimality system reduces to the standard LQG game with fixed information structures.
Conversely, a player can be made to ignore the effect of its own actions on the other players’ beliefs by setting its information wedge to zero, which is how the exogenous-signal benchmarks of the later chapters are constructed.
Lemma 1.10 (Strict convexity on the fixed reference filtration). Fix opponents’ complete causal maps, use the reference filtration of Proposition 1.6, and suppose the joint running Hessian is positive semidefinite and (1.3.3) holds. Then player ’s objective is strictly convex and coercive on the admissible control space. For any two controls the state difference is the exact linear causal response to their control difference, and the quadratic part of the cost difference is half the joint form in , which is nonnegative and dominates after completing the square in .
Proof. Proposition 1.6 makes the state and all induced opponent responses affine functions of the realized own-control input on one fixed adapted Hilbert space. Substituting them, the running integrand is half the joint form in ; completing the square in leaves plus a nonnegative term, so (1.3.3) gives the claim. ◻
A candidate that solves the stationarity system is a best response against every admissible deviation, including those outside the linear class.
Corollary 1.11 (Exact finite-deviation identity and best response over the full admissible class: Best Response Closure). Under the conditions of Theorem 1.8 and the Schur condition (1.3.3), let be a noise-state linear candidate whose deterministic kernels satisfy the stationarity system (1.B.13), and let be the corresponding state price. For every admissible strategy map of any form, set Then the following identity is exact:
The kernel stationarity equations make the first line on the right vanish for every adapted finite deviation. is then the unique best response over the full admissible class.
Proof. Appendix 1.B derives (1.4.10) by applying the physical and belief prices to the exact finite-difference system. Stationarity gives . The remaining terms are the joint quadratic form, nonnegative by Lemma 1.10, and bounded below by after completing the square. ◻
Corollary 1.11 takes a solution of the stationarity equations as given; the next proposition shows the best response exists.
Proposition 1.12 (Global solvability of unilateral best responses). Fix a finite horizon and complete linear causal strategy maps for all opponents, including their fixed filters from observations to controls. Under Assumption A.1, player ’s optimization problem has a unique best response over the full admissible control class. By Corollary 1.11, this best response is noise-state linear.
The proof is in Appendix A.2.
Remark 1.13 (Separation and the information wedge). Standard LQG separation holds the feedback gain fixed as signal precision changes. Here the FOC, , depends on the player’s conditional expectation of the state price. When actions affect opponents’ signals, that state price contains the wedge, which depends on opponents’ filters and belief prices. Changing information then generally changes the policy rule itself, not only the posterior variance.
1.4.1Identical-interest LQG games with control externalities
The baseline decentralized LQG game lets opponents’ controls affect player ’s payoff only through the state . In the identical-interest variant, every player evaluates the same quadratic cost of the full control profile. The state and observation equations (1.3.1) and (1.3.5) are unchanged. Let Fix deterministic weights together with deterministic linear terms and and a deterministic constant . The common cost is Every player has the same objective, A profile is a cooperative Nash equilibrium if, for every player and every admissible private-signal strategy , Any global solution of the decentralized team problem is a cooperative Nash equilibrium.
Identical-interest games with dispersed information are the team problem of Radner [97] and Marschak and Radner [82]; the common-information approach of Nayyar et al. [89] solves the case in which every player also observes a common signal history. Here the information structure remains decentralized: player still computes the best response in the filtration , and a unilateral deviation still propagates through the other players’ private filters. In the identical-interest case the wedge no longer prices an adversarial manipulation of beliefs.
The filtering equations are unchanged, since the observation system is unchanged. Only the source terms in the control first-order system change. Let For player , the state price satisfies
For each opponent , the belief price becomes The information wedge is the same filtering chain-rule object as before, (1.4.6). The new source term is the effort cost of the control opponent takes, through its future posterior, in response to player ’s deviation. Under the common objective, player pays that cost too.
The best-response first-order condition is Equivalently, writing the block decomposition and assuming , the cooperative Nash policy satisfies In the separable-effort case, this reduces to but the opponent-effort terms still enter the belief-price equations through .
The objective remains quadratic and the state, filters, and frozen opponent maps remain linear in the unilateral best-response problem, so the best response is again noise-state linear. A fixed point of the resulting cooperative best-response system is then a Nash equilibrium in private-signal strategies with a common objective.
1.5Conclusion
In the continuous-time LQG benchmark, beliefs, prices, and policy kernels are deterministic impulse-response kernels. Corollary 1.11 verifies that a fixed point in the noise-state linear class is a Nash equilibrium against arbitrary admissible deviations. The two-player examples show that the wedge changes welfare comparisons, information design, and the response to signal precision. Even with identical interests, teammates’ effort enters the belief prices.
The following chapters keep the same objects and change the environment. Chapter 2 adds delayed observation rows, Chapter 3 rewrites the kernel system in stationary lag form, and Chapter 4 applies the adjoint logic to order-flow trading. Chapter 5 uses symmetry to solve finite local strategic-information games in relative-position coordinates, then aggregates their equilibrium outcomes across markets. Chapter 6 makes the observability of a deviation player-dependent.
Appendices to Chapter 1
The two appendices following Chapter 1 contain the canonical derivations used throughout the dissertation. Appendix 1.A derives the filtering closure and the primitive-noise representation of the noise state. Appendix 1.B derives the control first-order conditions, belief adjoints, and information wedge. Later chapters modify these objects rather than rederive them from scratch.
Proof-local notation.
The symbols this appendix introduces for test functions, filter errors, and characteristic-function arguments are local to the proof in which they appear.
1.AFiltering
Theorem 1.5 follows by specializing a generic linear-filtering result to the physical observation model and then converting from innovations to primitive-shock coordinates.
1.A.1Dynamics of the noise-state
Fix a player . Recall the noise-state , , and the observation channel where is deterministic. Since is deterministic, set without loss by replacing with .
Standing notation. Write in bold (to distinguish it from the player index ) and ; take all spaces to be complex; normalize the observation noise so that , so ; and assume for each , so that .
Lemma 1.14 (Test functionals). For let solve Each is bounded, and is total in . Consequently an -adapted process equals if it matches the test integrals
Proof. Each is -measurable, and Itô’s formula with gives ; so is deterministic and is bounded. For a Brownian motion the family is total by [14, Lemma B.39]. The drift of is square-integrable, so for the Brownian motion with the same diffusion coefficient (Girsanov), with density . Totality transfers. If for all , then is orthogonal to every under , so by B.39 and then . ◻
Proposition 1.15 (Dynamics of ). The observation follows with for each . Then there exist deterministic matrix-valued functions and such that, for each fixed , with where is the deterministic conditional error-covariance kernel.
Proof. Step 1 (Test family). Let () be the test functionals of Lemma 1.14; each is bounded, solves , and , justifying the interchanges below.
Step 2 (Characterizing identity). As is -measurable, the tower property gives for all . By Lemma 1.14 it suffices to substitute the ansatz (1.A.1), differentiate in , and match.
Step 3 (Left side). Split ; only the martingale part has quadratic variation, so (using ; the drift affects the -comparison, not the bracket). Itô’s product rule gives
Step 4 (Right side). With (Gaussian conditional law) and the conditional Itô isometry,
Step 5 (Matching). Equate the derivatives for all ; the common terms cancel. Varying the free endpoint (which leaves fixed) splits the identity into its -linear and -independent parts: (i) the -coefficient, deterministic after cancellation, gives ; (ii) the remainder for all , so separation (Lemma 1.14) and Itô isometry force . ◻
1.A.2Specialization to the physical observation model
Specialize to the observation equation of Section 1.3, which corresponds to setting in the general setting of Proposition 1.15.
Innovation process.
Define the innovation where . By the innovations theorem for Gaussian linear filtering (Liptser–Shiryaev [75], Theorem 7.12), is a standard -Brownian motion.
Innovation form of the dynamics.
Substituting into Proposition 1.15 and using , the dynamics (1.A.1) become The noise-state moves only with the innovation, through a covariance gain and an indicator term.
Identifying as the filtering gain.
For , the indicator term in (1.A.2) is locally constant in , so the mixed -update takes the form The gain in (1.A.3) is the conditional cross-covariance density, divided by , of and given . Since , Itô isometry gives .
Algebraic identity for in terms of .
Define the induced estimated-state kernel . Since ,
1.A.3The blueprint: direct projection and filter kernel
Define by its evolution equation, with given by (1.A.4).
Definition 1.16 (Blueprint system). Fix player . A deterministic causal kernel on solves the blueprint system if, with defined by (1.A.4), it satisfies and
Since (1.A.4) expresses as a deterministic linear functional of , the blueprint system (1.A.5)–(1.A.6) is a closed forward ODE for , quadratic through the substitution; its right side is locally Lipschitz in on bounded sets, so the solution is unique. It fixes the map from the primitive shocks to the player’s estimated shocks.
Convert the innovations-based characterization of Subsections 1.A.1–1.A.2 into primitive-shock coordinates. What is needed is a deterministic kernel such that for every and , carries the contribution at , and carries the density across the other source times.
Derivation
Return to the innovation form of the estimated-noise dynamics (equation (1.A.2)): for , where the first term is the direct observation at time and the second is the accumulated drift-based revision over . To pass from innovations to the primitive shock, expand each in terms of .
Expanding the innovation.
Under the physical measure the innovation decomposes as Substituting the linear forms and using : The drift is the state error written against the primitive shocks.
Expanding the direct term.
Substituting (1.A.7) at into : The second term is absolutely continuous in with primitive-shock density for . In kernel coordinates over the primitive shock, the direct term is the singular diagonal part, a Dirac delta at , while collects everything with a density. Carrying the delta as the matrix coefficient keeps a function.
Expanding the integral term.
Substituting (1.A.7) into : Term (a) contributes for . Term (b) has a deterministic integrand, so the stochastic Fubini theorem [113, Thm. 2.2] licenses exchanging the outer time integral with the inner Itô integral ; its hypothesis here reduces to the deterministic square-integrability which holds because the impulse-response kernels and the precision are bounded on the finite horizon . Re-expressing the integration region as then gives Term (b) contributes, at each source time , the inference accumulated after .
Assembling the density.
Collecting the coefficient of from all three contributions, with causality ( for ) zeroing the irrelevant indicators:
Kernel representation.
Fix ; the primitive Brownian motion has dimension . Under the boundedness assumptions on , , , and , the expression in (1.A.10) defines a deterministic kernel Applying stochastic Fubini to the preceding identity gives, for every deterministic , The notation is understood in the integrated sense of (1.A.11).
The kernel is unique up to Lebesgue-null sets. Suppose and both satisfy (1.A.11), and write . Then, for every , The Itô isometry implies The Hilbert–Schmidt operator generated by then vanishes on , so Accordingly, almost everywhere on .
1.BControl stationarity and finite-deviation verification
This appendix derives the stationarity condition (1.B.13), the closed backward system for the adjoint kernels, and the exact finite-deviation identity used in Corollary 1.11. Nothing here assumes that the best response is noise-state linear. Opponents’ complete causal maps stay frozen, while the deviating player may use any causally well-posed square-integrable observation-feedback strategy.
1.B.1Frozen opponents, exact finite differences, and the reference innovation
Fix player and the opponents’ complete noise-state linear maps, including their filters. Let and be any two admissible own controls and define For , the direct primitive-shock part of the two fixed filters cancels, so write The differences satisfy the exact system The opponent-control difference is . Subtracting the two closed-loop systems gives these identities. They are exact and linear in the realized input ; the rule generating that input may be nonlinear in the observation history.
The own-observation difference is Since (1.B.1)–(1.B.3) have deterministic causal coefficients, is known from the realized own-control difference up to time . Taking defines the control-free pair . The actual observation comes from by adding the known response of the realized control, and causal well-posedness gives the reverse reconstruction. The completed filtrations coincide. With the common reference innovation is It is a standard Brownian motion whose law does not depend on player ’s strategy, so the deviating player filters against the frozen environment and this known input.
1.B.2Spike variation and the first-variation system
Fix and a control spike . The spike perturbation induces the normalized first variation with , the induced state impulse.
Opponents apply fixed linear maps, so is linear with deterministic coefficients. The spike-induced state response is known to the deviating player and enters the reference conditional mean, so it creates no new own-innovation increment. The reference innovation (1.B.4) is unchanged by construction.
Opponents’ filter response.
For each , with . Existing coordinates satisfy and the coordinate born at satisfies The opponent control variation is , giving the coupled state system After the spike, the state moves through its own dynamics and through the opponents’ control variations.
1.B.3Transition kernels
The coupled system (1.B.5)–(1.B.7), with birth condition (1.B.6), has bounded deterministic coefficients and admits a unique deterministic two-parameter evolution family.
Definition 1.17 (Block transition kernels). For , define and (, ) by for initial data and the given . In the spike deviation , so . The notation denotes the birth value for the coordinate born at time .
These blocks satisfy with and . The terms involving come from the birth condition (1.B.6).
1.B.4First variation of costs and the adjoint kernels
Fix player . Using and the linear form of , the normalized cost variation is where the running marginal cost of the state now includes the cross term with the player’s own control, , and The upper limit in (1.B.10) may be written as since for .
Stationarity condition (globally necessary and sufficient).
Setting for all -measurable gives with equilibrium coefficients
Equivalently, . The player cannot act on shocks it has not observed, so the control is the conditional expectation of the state price given its own history. This is not a separation principle (Remark 1.13), because the state price itself is an equilibrium object shaped by opponents’ filters and belief prices.
1.B.5Closed backward system for the adjoints
Define the belief-price coefficients, the prices of moving opponent ’s noise-state coordinates, Equivalently, introduce the noise-linear adjoint processes Here and below, denotes the kernel of the process , while denotes a process whose kernel carries the second argument . The argument count differs between the two families because is itself indexed by a coordinate. For each opponent , set The first term carries the coordinate born at in (1.B.6), the second the accumulated drift inference from (1.B.5). Together they are the information wedge, the price player pays for what player learns from player ’s action. In coefficients,
The adjoints satisfy the two backward equations The terminal conditions are
1.B.6Exact finite-deviation verification
Let be the noise-state linear candidate satisfying (1.B.13), and let be any admissible finite deviation. Put and let solve (1.B.1)–(1.B.3). Expanding the quadratic objective gives
Apply integration by parts to the paired forward–backward quantity Set The product rule also picks up a birth contribution at the newly born coordinate . Explicitly, Substitute the exact finite-difference equations (1.B.1)– (1.B.3) and the adjoint equations (1.B.19)–(1.B.20). The physical-drift pairings and the opponent-control pairings cancel, leaving The last two terms in braces are, respectively, the old-history and newly born components of by (1.B.16). The braces then vanish opponent by opponent, and Integrating from to , the initial term vanishes because and the belief-coordinate integral is empty. At the other endpoint, , while . The terminal physical pairing is exactly the terminal linear term in (1.B.22), and the running pairing cancels the cross linear term of (1.B.22). What remains from all linear terms is Since is adapted to the common reference filtration,
Substituting gives the exact identity (1.4.10). The stationarity condition annihilates its linear term for every finite adapted deviation, while the remaining quadratic terms are nonnegative. Under (1.3.3), equality is possible only when in .
1.B.7CARA / risk-sensitive extension: exponential utility FOC
Theorem 1.18 (Risk-sensitive extension). Take , the tracking case. Replace (1.3.4) with the entropic risk measure where is the random cost (1.3.2). A Taylor expansion gives , so governs risk aversion (the risk-neutral case recovers (1.3.4)).
Under a linear profile, the stationarity condition is identical to (1.B.13) with replaced by the risk-adjusted noise-state where is the deterministic conditional covariance kernel of given (deterministic by Gaussianity of the conditional law under the linear profile), is the symmetric quadratic kernel of in the representation (1.B.28), and is the corresponding linear kernel (1.B.27). The formula is valid when , which is precisely finiteness of . The filtering structure is unchanged, but certainty equivalence fails; the system is the multi-agent, endogenous-information analogue of linear-exponential-quadratic-Gaussian control [62] and risk-sensitive certainty equivalence [115].
The spike variation of Section 1.B.2 carries over; differentiating the objective (1.B.26) in the spike amplitude gives with the pathwise first variation whose expectation is (1.B.10).
Quadratic structure of the cost.
Expanding through the linear forms and and applying stochastic Fubini gives the representation where absorbs deterministic and trace terms, and the linear and symmetric quadratic kernels are
Tilted conditional FOC.
Since is quadratic in , the weight tilts a Gaussian measure by a quadratic form, preserving Gaussianity. Setting for all -measurable and applying the tower property gives where .
Risk-adjusted noise-state.
Since is linear in with deterministic kernels, the FOC reduces to computing the tilted conditional mean . Under the deterministic-kernel profile, is Gaussian with mean and deterministic covariance . The quadratic tilt shifts this to provided . Here and act as integral operators on . The kernel of (1.B.28) is supported on and couples past and future coordinates. is the conditional covariance of the full path on , so restricting to would drop the past–future coupling that the tilt introduces. On this domain, is the resolvent (Neumann series) , convergent precisely when . Rearranging gives and substituting into the tilted FOC gives The pathwise cost is convex in the control path, is convex and increasing, and of a log-convex family is convex, so is convex in the control. Stationarity is also sufficient, and (1.B.30) characterizes the best response. The linear kernel (1.B.27) shifts the conditional mean toward lower expected cost; the quadratic kernel (1.B.28) reweights conditional precision toward the directions of along which cost variance is large. As , and (1.B.30) recovers (1.B.13).