Chapter 6Monitored Deviations: Naive Distortions and Privy Responses

6.1Introduction

Chapter 4 retains the classical Kyle–Back division of roles, in which informed traders choose strategic controls while the competitive market maker’s pricing rule is pinned down by conditional-expectation consistency. The price is then not a control the market maker can deviate from unilaterally. A strategic market maker requires a different equilibrium test. Its quote must be an admissible control, and changing the quote unilaterally must not improve its payoff.

Such a deviation is not hidden inside aggregate order flow. In the transparent market, traders observe the information needed to recognize a departure from the market maker’s quote rule. The deviation propagates through the other players’ strategies before returning to the market maker’s objective. Pricing this feedback requires a coordinate for each player who sees the deviation and responds to it, and an adjoint valuing the resulting chain of responses. The finite-horizon response calculus below defines these objects. Applying it to a strategic stationary market maker additionally requires specifying the market maker’s objective, solving the market equilibrium, and translating the gains, one for each pair of a deviating player and a player who responds to it, into the lag coordinates of Chapter 3.

The preceding chapters vary what players know about nature, the projection , while every deviation stays hidden inside others’ filtering. This chapter varies what players know about the origin of an off-equilibrium deviation.

In Chapter 1 the deviation’s origin is known exclusively to the deviating player. The deviating player knows the drift it inserted, so its own innovation is invariant, while every other player attributes the resulting movement to primitive shocks and filters it accordingly. The information wedge prices this unrecognized-deviation channel. It does not record how another player changes its strategy after recognizing who deviated.

Here a deviation may instead be monitored. A player who knows its origin, a privy player, treats it as a labeled, perfectly observed, zero-variance shock and responds through a separate strategy coordinate. A player who does not know its origin, a naive player, continues to filter its effect as primitive shocks. The same seed can generate privy responses for some players and naive filtering for others. The convention of Chapter 1 is the corner in which no opponent monitors the deviation, while a publicly attributable market-maker quote lies in the opposite corner in which every relevant opponent is privy. When every opponent is privy, a deviation moves no opponent’s beliefs, the information wedge vanishes, and equilibrium play is the separation principle applied to the closed-loop strategies of the perfect-information game [16, 25]. In the classical perfect-information game every player sees the shocks and is privy. The monitoring relation (Definition 6.1) moves the game between these corners; Proposition 6.13 recovers the two endpoints exactly. The naive players here differ from the deceived agents of the adversarial-deception literature [67, 119]. Their strategies remain equilibrium best responses, and their inference is Bayes-correct given their information (Remark 6.3); they lack only the origin label. Table 6.1 locates the four regimes. This chapter mixes the two columns: the same deviation can be observed by some players and unobserved by others.

Table 6.1. Four information regimes. The monitoring relation moves the game along the bottom row, and Proposition 6.13 recovers its two corners exactly; the top row is the classical pair.
deviations
2-3 shocks unobserved observed
observed open-loop Nash closed-loop Nash
filtered Chapter 1 all-privy corner

Mathematically, monitoring enlarges the forward system. A labeled deviation gives every player who sees it a response coordinate, in addition to any shifts in ordinary noise-state estimates, so merely removing privy players from the wedge sum would omit their strategic responses. The response is the linear perturbation implied by the equilibrium policy kernels, not a re-solution of the game after the deviation. For each responder–origin pair one backward equation, (6.5.16), gives a gain , and the response to a seed at any time is read off by (6.5.23), so no forward–backward system is solved per seed time.

Index roles are fixed throughout this chapter. The pair always reads “responder , origin .” Symbols carrying or together with identify the ordinary noise-state row or column inside that responder–origin map.

6.2Baseline noise-state setting

The objects of Chapter 1 carry over. Let be the player set and let the primitive uncertainty be a Brownian shock path . The state and player ’s observation evolve as so that player ’s control enters the state through the control-to-state map , and is the state-to-observation map; and load the common shocks into the state and into player ’s observation. Player observes a filtration and carries the noise-state and its equilibrium action is noise-state linear, Here “linear” refers to linearity in the represented shock coordinates, not to linearity of the kernels, themselves determined by nonlinear fixed-point equations.

Each player’s running cost is quadratic in the joint physical-state and control variable, as in (1.3.2), with Hessian Throughout, satisfies the Schur condition (1.3.3) of Chapter 1.

6.3Deviation origins and monitoring

6.3.1The monitoring relation

The monitoring relation is directed. Every player knows the origin of its own deviation, so the relation is reflexive and the set of privy players includes the deviating player itself. The blip-of-irrationality convention of Section 6.4.2, a single forced departure after which the deviating player resumes equilibrium play, fixes what this means for the deviating player’s own continuation.

Definition 6.1 (Monitoring relation). Write if player can resolve deviations whose origin is player . Equivalently, when player creates a deviation seed, the impulse that starts the deviation, player observes its origin perfectly and treats it as a separate deviation shock rather than filtering it as part of the primitive uncertainty. The relation is reflexive, so for every . Throughout, a deviation is a control spike of player , and is the state impulse it produces; the seed carried by the forward system is , not . It is the state impulse alone; Chapter 4’s seed was the filter displacement a spike leaves.

For each origin , define The set contains every player who observes ’s deviation seed, the deviating player among them, while contains the players who do not.

A common source of the relation is information inclusion. If every information source available to player is also available to player (for example , with observation of ’s policy rule and action channel), it is natural to set . Inclusion is sufficient: can compute ’s equilibrium action, and any gap between it and the realized action is ’s deviation. It is not necessary: may observe ’s action directly, a posted quote for instance, without observing ’s signals, as in Section 6.8. The calculus below requires only the relation ; the modeler specifies the institutional or informational mechanism that generates it.

6.3.2Transitivity

Transitivity is assumed: The condition is substantive. Suppose is privy to and so responds to ’s seed. If is privy to but not to , then sees ’s response while lacking the coordinate needed to represent the cause of that response. The response chain is then ill-typed: can identify the proximate actor but not the root seed, a murder mystery.

Contrapositively, if is naive to and is privy to , then must also be naive to ’s response to : Transitivity is the monitoring analogue of partial nestedness in team theory [55]. There, an agent affected by another’s action observes that other’s information; here, a player who sees a response also sees its origin.

Lemma 6.2 (Naive sets). Under transitivity (6.3.3), for every privy player . A player naive to the deviating player is naive to every response by a privy player . Section 6.5 reads this as a containment of index sets: the rows of every gain, indexed by , contain its columns, indexed by .

Proof. If and , then (6.3.4) gives , so . ◻

Remark 6.3 (No smoke without fire). Consider the ill-typed case above, in which observes player ’s response to an -origin seed but not the origin . Seeing only the response move the state, could not treat it as a labeled coordinate. It would have to attribute the movement by Bayes’ rule, forming a posterior over which origin produced the response, supported on every deviation cannot rule out. That posterior is not a free belief but an equilibrium object. Transitivity is exactly the condition that forbids this case. On the privy side the attribution never arises: a privy player’s response kernel is indexed by the originating seed , not by an inferred mixture over possible origins. To rule out this attribution problem, no player may see smoke without seeing the fire that caused it.

With reflexivity and transitivity, is a preorder. If and , then and are informed about each other’s deviation origins.

Assumption 6.4 (Perfect transitive monitoring). The relation is reflexive and transitive, and monitoring is perfect: a player either observes a deviation origin exactly or not at all, with no intermediate noisy observation. Formally, means that player ’s strategies may depend on the coordinate of Remark 6.5.

This assumption is maintained for the rest of the chapter; Remark 6.12 returns to the role of perfect observation and to what fails when it is dropped.

6.4Naive versus privy coordinates

Monitoring breaks a convention every other chapter relies on. The control is noise-state linear, (6.2.3), and everything a player could react to reaches it through , an opponent’s deviation included, since a deviation enters the observation stream and is filtered like any other movement of the state. A monitored seed is not like that. It moves the physical state, and it moves the beliefs of the players who do not recognize it, but for a player who does recognize it there is nothing to filter. The seed has zero variance under the equilibrium law, so it is not a coordinate of any noise-state, and a strategy written as a function of alone cannot respond to it. The response has to live somewhere else. A strategy must therefore split into two parts. The on-path part is the policy kernel of the other chapters. The response part is a kernel on the labeled seed coordinate, zero on the equilibrium path and pinned only by what the player would do off it. Section 6.4.3 makes the decomposed object the strategy of the monitored game.

6.4.1Naive and privy players

A naive player does not recognize the deviation seed as a deviation. Its physical effect enters ’s observation stream, and filters that effect as if it were evidence about the primitive shock path. The ordinary noise-state changes: Throughout, is the noise-state perturbation in density form, ; the singular part does not vary. For a privy player , the ordinary shock estimate does not change when deviates: Instead, receives a separate input coordinate, the deviation seed (Remark 6.5). The seed enlarges player ’s sufficient statistic from to This labeled coordinate needs a response kernel of its own.

Remark 6.5 (Deviations on their own factor). Deviations live on their own factor. Let be player ’s admissible deviations, the control paths of Chapter 1, and . Player ’s strategy is measurable in and in the factors with , its own among them, and constant in the others. Under the equilibrium law the seed has zero variance, so a naive player’s filter carries no coordinate for it and never anticipates it; a privy player reads it as a known input. At fixed equilibrium gains the displaced state, filters, and responses are affine in , so the coefficients below are exact for finite deviations, and two deviations, from one origin or several, displace by the sum. The seed affects conditional means but not conditional covariances or the ordinary filter for the primitive shocks. A monitored deviation can move a privy player’s action but cannot manipulate that player’s ordinary noise-state. That perfect observation keeps the privy response affine in the seed.

As in the Chapter 1 spike variation, denotes the induced state impulse, the state movement a control spike produces. The forward variational system starts from If the underlying control-space spike is , then , and the deviating player’s first-order condition pulls the state price back to control space through . The kernel represents a privy player’s response per unit of observed state impulse; a control-space spike first maps into the state impulse and then propagates.

6.4.2Blip of irrationality

The convention that fixes how privy players continue off equilibrium is a blip of irrationality, in which a deviation is a single forced departure from ’s best response, after which continues through its own response kernel . The other players do not read the blip as evidence about ’s type or as a loss of reputation. Continuation play treats the deviation as bygone, as in Markov perfect equilibrium [83]. The responders price the mechanical effect of the seed and carry on as though is again rational.

The seed spawns no independent root seeds. It induces responses that change the state and feed naive filters and other controls, but each response, the deviating player’s own included, is a deterministic linear function of within the per-origin system.

Chapter 1 needed no such convention. No opponent observes a deviation’s origin, so how the deviating player’s own future departures are grouped into seeds is invisible to every other player, and only the total control path matters. Once some players are privy, the grouping matters. A responder reacts to each labeled seed it observes, so what counts as one seed with its continuation, rather than several separate seeds, determines the responses. The blip convention pins this down. Given it, the deviating player’s own deviation calculus does not depend on how a departure is decomposed.

Lemma 6.6 (Blip and frozen continuations). Fix an origin . For a seed time in the horizon of Chapter 1, let denote the control spike at followed by the frozen continuation, as in Corollary 1.11. Let denote the same spike followed by the blip continuation, in which resumes its equilibrium response. Then where is the bounded strictly causal kernel generated by the deviating player’s response to its own seed, and is invertible on the space of admissible deviations, the control class of Corollary 1.11. Consequently the families and span the same space of admissible deviations, and the deviating player’s first-order condition (6.5.36) is the same under either convention. The deviation quadratic forms of Corollary 1.11, written in the two bases as and , are related by the invertible change of variable , with , so one is positive definite (semidefinite) exactly when the other is.

Proof. The two continuations agree on and differ on by the response path that the blip convention appends, the linear image of the seed under the deviating player’s own response kernels propagated through (6.5.8). This difference is an admissible control path supported on , and every such path is a superposition of frozen spikes at times ; reading off the coefficient of defines . The kernels are bounded on the compact horizon (admissibility), so is a bounded Volterra operator and is invertible. The inverse is given by the series , because is strictly causal, feeding a spike at only into times after . Both families parameterize the same space. A linear functional vanishes on a spanning family exactly when it vanishes on the space, which gives the first-order claim. Substituting into the deviation quadratic form gives , and an invertible change of variable preserves positive definiteness (semidefiniteness). ◻

The exact finite-deviation identity of Corollary 1.11 transfers through , so verifying a blip-convention profile against all admissible deviations reduces to the Chapter 1 identity.

Monitoring creates two belief freedoms that the all-naive setting of Chapter 1 lacks. A player who sees a movement may need to know who caused it, and a player who sees a seed may need to know what the deviating player will do next. Transitivity closes the first, so no attribution belief is ever needed (Remark 6.3), and the blip convention closes the second by fixing the continuation.

6.4.3The monitored game and its equilibrium

The primitives of the monitored game are the baseline game of Section 6.2, the monitoring preorder of Definition 6.1 under Assumption 6.4, and the blip convention. A strategy for player has two parts. On the equilibrium path it reduces to the noise-state linear map (6.2.3). Off the path, player ’s strategy may depend on the coordinates with (Remark 6.5), and it specifies a response kernel pair for each origin with . A player’s full feedback map is the pair of on-path kernels and response kernels.

Definition 6.7 (Monitored-deviation equilibrium). A profile of full feedback maps is a monitored-deviation equilibrium if

  1. for every player , no admissible control deviation improves ’s payoff against the other players’ full feedback maps. The deviation’s consequences are generated by the monitoring structure: each privy player responds through its response kernels and each naive player filters the deviation’s effects with its equilibrium filter;

  2. for every origin , seed, and privy player , player ’s response is optimal for conditional on the observed seed, against the other players’ full feedback maps.

Condition (ii) is sequential rationality, optimality conditional on events that equilibrium play never produces. A seed occurs with probability zero, so (ii) changes no player’s ex-ante payoff. It pins the response kernels. For it pins the deviating player’s own post-blip continuation. At first order, (ii) is the privy player’s first-order condition (6.5.9), with the response kernels (6.5.23), and (i) is the deviating player’s condition (6.5.36); both come out of the rectangular backward system of Proposition 6.9. Lemma 6.6 lets one verify condition (i) beyond first order. The finite-deviation identity of Corollary 1.11, transported to the blip basis, expresses the payoff change of an arbitrary admissible deviation as its first-order term plus a quadratic form in the deviation. Verification requires this quadratic form to be positive semidefinite.

Existence is a separate question. At the corners, Proposition 6.13 reduces the concept to systems with known existence theory, the coupled Riccati equations of the closed-loop game [16, 26] and the finite-horizon results of Chapter 1. Away from the corners the profile and the response kernels form one fixed-point problem, the one the market of Section 6.8 poses; a general existence theorem is open.

6.5The seed-dependent system

Fix a deviating player and a seed time , with privy set and naive set as in (6.3.2). The state dynamics are (6.2.1), with player ’s control entering through . We first derive the forward equations and the privy player’s response to the seed. We then write the backward system in operator form and in coefficient blocks, before treating the deviating player’s response to its own seed.

6.5.1Forward deviation equations

The initial condition for a state impulse by player at time is For , the displaced state is The state equation is For each naive player , define the naive error variation The kernel is the impulse response of the state at to a primitive shock at . It carries no player index, since the dynamics are common; integrating it against the truth gives the state, and against a player’s estimates gives that player’s estimate of the state. Player ’s unresolved-state kernel of Chapter 1 is the part of that response player has not yet resolved, and its innovation kernel is The naive filter variation has an interior revision equation and a birth term Substituting (6.5.4) into (6.5.6), the coefficient of in is This is the belief–belief block of the generator, and it is not symmetric in : the lag being revised enters through the error kernel, the lag of the shock being estimated through the state kernel. That asymmetry makes the generator non-self-adjoint.

6.5.2A privy player’s response

Fix a privy player . To compute player ’s response to the origin- seed, one tests whether would profitably deviate from its proposed response. This introduces a second, hypothetical, -origin spike. The -seed and this test -seed are distinct labels. In the gain defined below, the rows are indexed by , the players naive to , because they come from player ’s own deviation problem; the columns by , the players naive to , because they come from the already-realized -origin displacement.

Let the closed-loop coefficients for an -origin deviation be These coefficients use the response kernels appropriate to the origin label . When they propagate the observed -seed; when they are the row-side coefficients that price a hypothetical deviation by .

The -variation of player ’s closed-loop adjoint has rows and for , dual to the row state of (6.5.11). The next two subsections write it as one rectangular gain, first in operator form and then in coefficient blocks, with source .

Player ’s first-order condition, conditioned on its own information , gives its response. Player ’s strategies are measurable in the factor (Remark 6.5), so the origin- seed is a known input to . Subtracting the condition player satisfies on the equilibrium path from the one it satisfies after the -seed gives Player observes the -seed, and the forward and adjoint systems are linear with deterministic coefficients, so the conditional expectation collapses:

Remark 6.8 (Two labels, no cross term). The forward system is affine in the seeds, so the - and -labeled forward effects add, and the test uses the same gains as the ordinary -origin deviation problem. The -seed shifts the value of player ’s adjoint while leaving that gain system intact.

6.5.3Operator form and source

The operator notation compresses the seed-dependent physical and noise-state equations above. Stacked symbols collect Chapter 1 coordinates, and the rectangular gain collects their adjoint responses into one backward solve.

An -origin seed moves the noise-states of , while a hypothetical deviation by moves those of . There is then no single displaced state, so the origin label indexes the state. For any origin label , define the state attached to that origin by The superscript is the origin label, and the noise-state fields are those of the players naive to origin .

For an -origin deviation, write The deviation dynamics are where is the controlled-state transport part and is the filter birth-and-revision part. Throughout, is the full generator and its physical-state block.

In coordinates, the row is (6.5.3) with the coefficients (6.5.8) at label , and the filter rows are (6.5.6)–(6.5.7) for . The same non-self-adjoint generator template is used for every origin label. On the row side of a Sylvester equation it is ; on the column side, .

Fix an observed seed of origin and a privy player . The -origin state supplies the columns and the rows, because player ’s adjoint lies in the dual of . The rectangular gain is then one operator In block form,

At the operator level, the backward equation is the rectangular Sylvester equation

The source is player ’s ordinary running cost, differentiated once along ’s own deviation direction and once along the -seed direction, the Hessian evaluated on the two directions. Let be the physical-state projection. Let be player ’s response to its own seed, and let be player ’s response to an -origin seed. Then For , the row and column maps coincide and is symmetric. For , is generally rectangular and nonsymmetric.

Write the response kernels as Then the source blocks of (6.5.17) are

The response kernels themselves are Together, (6.5.16), (6.5.17), and (6.5.23) form a rectangular Riccati–Sylvester fixed point.

6.5.4Coefficient blocks

The gain equations are the coefficient form of the seed-dependent system. The four block equations below share one backward template, specialized to a row coefficient and a column coefficient.

Proposition 6.9 (Rectangular gain reduction). Fix an origin and a privy player . Under Assumption 6.4, suppose the variations and , , are linear in with bounded deterministic coefficients. Then they take the rectangular form The coefficient functions solve (6.5.28)–(6.5.32), with source blocks (6.5.19)–(6.5.22). Equation (6.5.23) then recovers player ’s response to any realized origin- seed without solving a new forward-backward system for each seed time.

The filter contraction.

The filter generator enters the adjoint through one operator and its transpose. For a naive player , the filter contraction of Chapter 1, equation (1.4.9), acts here on a noise-state lag field by the filter birth at lag together with the revision at interior lags. Its transpose acts on a field from the right, Contracting the responder’s -row on the left, for , gives the cost to player of having its own hypothetical deviation treated as noise by the players naive to it, the sum over in (6.5.28). Contracting the deviating player’s -column on the right, for , gives the term through which the deviating player’s seed enters the naive filters, the sum over there. Each contraction leaves the other lag as the displayed argument. The observation map multiplies the column contraction on the right in the gain equations below.

Figure 6.1 assembles the four equations as a block. The state and the naive coordinates () and () index the rows and columns, and .

Figure 6.1. Reconstructing the backward system . Row terms come from (left), column terms from (top); the -row revision and the -column residual term carry the opposite sign. Each cell is completed by adding the two terms on its row margin and the two on its column margin; the dot is the free index, filled by the cell’s column on the row margins and by its row on the column margins.
Rows.

Matching the coefficient gives Here and below the terminal condition appears blockwise; all non- terminal blocks are zero, matching the operator terminal of (6.5.16).

For each , matching the coefficient of gives The term inherits its sign from the residual convention For each row , matching the coefficient gives For each row and column , matching the coefficient of gives Together these are the rectangular backward system. It is nonlinear because the response kernels in (6.5.23) and the source blocks (6.5.19)–(6.5.22) depend algebraically on the same coefficients.

6.5.5The deviating player’s own response and its first-order condition

By reflexivity, the deviating player is one of the privy players. Under the blip-of-irrationality convention (Section 6.4.2), the deviating player’s post-blip continuation is its equilibrium continuation law, so its response to its own seed is the diagonal response . The diagonal solves a backward equation of the form , a source plus the generator acting on each side. It is the rectangular Sylvester equation (6.5.16) at , expanded into the blocks (6.5.28)–(6.5.32). It is special in two ways. First, the deviating player’s continuation control moves with the seed, so its own response is charged in the second variation rather than dropped. Second, the row and column maps coincide. The source is ’s running-cost Hessian evaluated on its own response kernel, Here is the physical-state projection of (6.5.17). Contracting the case of (6.5.19)–(6.5.22) against the seed gives, on the row, and, on each belief row,

The deviating player’s first-order condition sets to zero the coefficient of an arbitrary -measurable control spike: or equivalently Read as a response kernel rather than a level, this condition is the case of (6.5.10).

Remark 6.10 (Symmetry of the own-response equation without a self-adjoint generator). The diagonal equation has the form The full generator is not self-adjoint, by the belief–belief asymmetry noted after (6.5.6), so the row and column equations need not look term-by-term symmetric. What is symmetric is the diagonal source and terminal condition, so the solution preserves the adjoint symmetry provided the backward system has a unique solution.

6.6Superposition under perfect monitoring

For a fixed origin , the privy players in all see the same seed and solve one self-consistency problem, coupled only through the shared forward state. The monitoring relation orders who sees whose seeds. Each origin produces its own gain system of the form (6.5.28)–(6.5.32); the systems differ because the naive/privy split depends on the origin.

Write for the per-unit privy response to an origin- seed at time , and for the response to a configuration of seeds.

Proposition 6.11 (Superposition of monitored seeds). Under Assumption 6.4, the forward system is affine in the seeds, so the response to several monitored seeds is the sum of the single-seed responses,

If player blips, then the response is the same regardless of whether some other player blipped sooner or not. Without superposition, verifying would mean thinking about an infinite tree of deviations rather than treating them one at a time. The gains themselves are solved in order rather than independently, since the source (6.5.19) for contains , player ’s response to its own seed, so each solve uses the solve first. The pairs with for one origin are coupled through , which contains every , .

Remark 6.12 (The role of perfect observation). The linear theory relies on binary monitoring. A player either observes a deviation seed perfectly, as a known zero-variance input not jointly estimated with the fundamentals, or not at all, in which case it is folded into the ordinary filter with the equilibrium filter fixed. Keep two relaxations apart. In the first, a player observes the seed through an additive noisy channel. A noisy reading of a point mass at zero carries no information, so this channel is informative only if the seed carries a nondegenerate prior. Deviations then occur with positive probability on the equilibrium path, and the object being solved is no longer a deviation calculus. In the second, transitivity fails and a player sees a response whose origin it cannot resolve. The response discrepancy is identically zero on the equilibrium path, so the observation is a null event, and attribution conditional on it is a mixture over the origins the player cannot rule out. The mixture weights play the role that off-path beliefs play in sequential equilibrium [68]. Given the weights the theory remains linear, but superposition fails, since attribution couples the per-origin blocks into one system.

6.7The two corners

For a fixed origin , solve (6.5.28)–(6.5.32) for each privy player . By Lemma 6.2 the columns are contained in the rows .

The variations and in (6.5.24)–(6.5.25) are the marginal values of the seed to player . The responder’s own first-order condition pins each response kernel; none is assumed. This is the consistency required of conjectured reactions in oligopoly theory [22], and the selection that closed-loop differential games otherwise leave open [16].

The two corners of the monitoring relation recover the two classical regimes.

Proposition 6.13 (Corner reductions). Maintain Assumption 6.4. (i) (All privy.) Fix an origin with . Then the diagonal of (6.5.28)–(6.5.32) retains only its block, a Riccati equation in with row and column coefficient . If for every origin , so every naive set is empty, then every solve retains only its block and its solution does not depend on the origin, for all . Writing for the common response kernel , the pairs solve the coupled Riccati system of the closed-loop Nash equilibrium of the perfect-information game [16, 26], No seed moves any player’s ordinary filter. (ii) (All naive.) Fix an origin with . Then only the diagonal remains, and its rows and columns are both indexed by . The system (6.5.28)–(6.5.32) is the coefficient form of the -variation of the adjoint system of Theorem 1.8 of Chapter 1, written in the blip basis of Lemma 6.6, and the exact verification identity of Corollary 1.11 transfers through the basis map.

Proof. (i) With the diagonal has no rows or columns beyond , so only (6.5.28) remains, its source the case of (6.5.19) and its row and column coefficient . With every naive set empty the rows and columns are absent for every pair. The candidate matches the terminal condition for every pair. Substituting it into (6.5.23) gives independent of the origin, so for every origin label in (6.5.8), and the source (6.5.19) becomes the first four terms of (6.7.1). The right sides of (6.5.28) and (6.7.1) then coincide, so the candidate solves the corner system, and it is the solution whenever the closed-loop Riccati system (6.7.1) has a unique solution. That no seed moves any ordinary filter is (6.4.2) applied to every player.

(ii) With only the diagonal pair remains, with rows and columns both indexed by the deviating player’s naive set. Expanding the -variation of (1.4.7)–(1.4.8) of Chapter 1 in the coefficients of and comparing with (6.5.28)–(6.5.32), the terms of (6.5.19) in and the closed-loop part of the row and column coefficient differ from the frozen-continuation expansion by by the case of (6.5.23) and its transpose, and blockwise the same for the rows and columns, so the two systems differ by the operator . The two conventions are the two continuations of Lemma 6.6, whose change of variable carries this term, and the lemma transfers the verification identity. ◻

Remark 6.14 (Lyapunov and Riccati forms). The frozen continuation holds the deviating player’s control fixed after the spike, so its backward equation is linear in , a Lyapunov equation. The blip continuation lets the deviating player re-optimize, and substituting the gain (6.5.23) makes the equation quadratic in , a Riccati equation. The term is the value of that re-optimization.

6.8A strategic market maker: transparent and opaque markets

In the Kyle–Back markets of Chapter 4 the market maker is not a player. It posts the conditional expectation of the fundamental given order flow, and the only strategic question is how the informed traders trade against that rule. Here the market maker chooses its quote and cares about its inventory, and whether the trader can see its deviations depends on whether order flow is published. Publishing the flow turns the same market, with the same information on the equilibrium path, from a naive corner into a privy one. The market maker observes the order flow, so its observation loads on the trader’s control rather than on the state, which puts the market outside the baseline (6.2.1) and inside Chapter 4. The origin sets and the blip convention are those of this chapter; the trader’s response kernels come from Chapter 4’s adjoint system.

6.8.1The market

One asset, one informed trader (player ), one market maker (player ), a finite horizon . The fundamental is a Brownian motion with variance rate , the market maker posts a quote , and both players’ payoffs depend on the two only through the mispricing the gap between what the asset is worth and what it is quoted at. The trader observes the signal flow of Chapter 4 and submits an order rate ; noise traders submit ; the market maker sees the total order flow and absorbs it into its inventory, The trader is risk neutral with a quadratic trading cost and earns the mispricing on its orders, and the market maker pays it on the flow it absorbs and pays for inventory, The mispricing is both the market maker’s control and something it does not know. It sets through its quote, but only up to its own estimation error: with , and is its control; is a quote below the belief. The trader likewise acts on its own estimate , which for the trader is a state of its filter. The state is , of which only has dynamics of its own; the shocks are ; and and appear only through . At the first-order condition returns , the competitive rule of Chapter 4; with the market maker quotes off it, and stays fixed across the two cases below.

Two markets.

The market maker’s information is the order flow in both cases, , and it cannot see the trader’s order apart from the noise. Only the trader’s information differs.

  1. Transparent market. The trader sees the quote and the order flow, .

  2. Opaque market. The trader sees the quote only, .

On the equilibrium path the two coincide. The quote is a function of the flow history, and when that function is invertible the trader recovers the flow from the price. Off the path they differ. In the transparent market the trader can compare the posted quote with the equilibrium rule applied to the flow it sees. Any gap is the market maker’s, so the trader is privy to its deviations, . In the opaque market a quote off the rule looks like a quote on the rule after an unusual run of noise trades. The trader filters the gap as noise, . In both markets the trader’s order is hidden in the flow, . For a trader origin, in both markets: the trader keeps the information wedge of Chapter 4 against the market maker’s filter. For a market-maker origin, and in the transparent market, while in the opaque one.

What a quote deviation does.

A quote below the rule draws extra buying now, since the trader sees a larger mispricing. The extra buying lowers the market maker’s inventory, and the inventory outlives the quote: once the market maker is back on its rule, the quote keeps moving with the inventory and the trader keeps trading against it. The market maker therefore weighs the margin it gives up on the flow it already expects against the inventory the extra flow removes. The trader weighs the mispricing it earns against the inventory its order leaves behind and, where the market maker cannot see the order, against the beliefs it moves.

The seed.

A control spike of width is a quote below the rule. It raises the mispricing by during the blip, and the trader’s estimate with it in both markets. All information about reaches the market through the trader, so the quote carries none and a quote seed moves no belief about the fundamental. With the trader’s contemporaneous loading on the mispricing, the trader raises its order by , which lowers the inventory drift by as much and over the blip displaces the inventory by of order and permanent. Every origin- system below has this displacement as its state column.

The market maker’s cost.

The market maker maximizes (6.8.6), so its running cost is the inventory penalty plus the mispricing paid on the flow, . The market maker sees neither factor of the product, so it minimizes the expectation given . That is the product of the expectations plus a covariance, and the expectation of the mispricing is its control, The covariance is the adverse-selection cost, what the trader earns from knowing more. Since the quote is known to the market maker it equals , an equilibrium object that the quote does not move. With the trader’s loadings substituted, the market maker’s estimate of the order is . The running cost is then a quadratic form in the market maker’s own mispricing and its inventory, the block form of (6.2.4), so , , . The trader’s reaction supplies the curvature of the market maker’s cost in its own control. The trader’s running cost is , so . Writing the mispricing through the quote rule, , separates the part the trader knows better from the part the inventory has already moved. It gives the two cross blocks and , with the quote’s inventory loading found below.

6.8.2The transparent market

A quote deviation: the trader’s response.

The trader is privy, so its estimates do not move and it responds through the two coordinates the seed moves, The first loading is the trader’s on-path loading on its own mispricing; a quote shift is, for the trader, a movement of like any other. The second is its loading on the market maker’s inventory with the quote held fixed, where is the value to the trader of a unit of market-maker inventory and is that unit’s effect on the price the trader puts on the market maker’s beliefs. The contraction against the market maker’s filter is , as in Chapter 4. A departure from this loading is ordinary flow to the market maker. Adding the trader’s reaction to the quote’s own response to inventory gives the closed-loop drift of an inventory displacement, negative when the displacement unwinds; is the order’s total loading on inventory. When the market maker’s inventory aversion outweighs the trader’s loading, , and a long market maker quotes below its belief. The two gains solve

with zero terminal conditions. A unit of inventory is worth, to the trader, the orders it will draw times the quote’s inventory loading net of the inventory gain, accumulated along the unwinding. It moves the price the trader puts on the market maker’s beliefs by the same orders times the quote’s lift per unit of belief, , which the market maker then filters away. On the path no belief moves, but the loading the market maker faces depends on how it would misfilter a departure from that loading.

A quote deviation: the forward system.

With there are no naive noise-states, so the filter equations (6.5.6)–(6.5.7) are absent and the displaced state of Section 6.5 is the inventory alone. The seed is the state impulse (6.5.1) through the market maker’s control-to-state map, a quote below the rule and an inventory not yet moved. After the blip the market maker’s control returns to its rule, so the only mispricing left is the quote’s response to the displaced inventory, and the trader answers through both kernels, Substituting into (6.5.3) closes the system on the inventory, so the seed’s whole forward trace is one exponential at the closed-loop drift, and it decays exactly when that drift is negative.

A quote deviation: what the market maker does after its blip.

The market maker’s own continuation is the diagonal on the inventory coordinate, with the blocks of (6.8.10) and its own response kernel substituted, and the row and column maps coincide on the diagonal (Section 6.5.5), both being the closed-loop drift (6.8.13). The covariance in (6.8.10) is dropped, since the quote does not move it. Here is the marginal value of inventory per unit of inventory. The quote’s inventory loading is the trader’s direct loading on inventory less the flow the quote attracts times the inventory price per unit, divided by the curvature of profit in the quote, Read backward from , where inventory has no value, accumulates three things: the cost of holding inventory net of what the trader’s loading on it earns, the drift of inventory at that loading, and the saving from quoting against it, the closed-loop Riccati equation of corner (i) on the inventory coordinate. The inventory price in the first-order condition is , a public quantity. There is no belief row and no belief column, since no seed of the market maker’s moves any filter and its own filter is public. At the pair , solves (6.8.20)–(6.8.21) and gives . The trader loads on inventory only because the market maker quotes against it, and the market maker quotes against inventory only because the trader loads on it, so the competitive rule is a fixed point of that loop. Whether the loop has a second, self-fulfilling fixed point at is not settled analytically; the stationary computation below finds a candidate and rules it out.

An order deviation: the trader’s own condition.

Only the market maker is naive to a trader seed, , so it indexes both the rows and the columns. The diagonal is the coefficient form of Chapter 4’s adjoint system, corner (ii), with the inventory added. Two entries are new. Because the market maker quotes against its inventory, the trader’s cost acquires the cross term , the derivative of in through the quote. The inventory is -measurable, so the market maker’s filter has no coordinate and its observation loads on the trader’s order rather than on . The row of the trader’s state price therefore has no filter coupling and involves only and the on-path order . With a single trader there is no other trader whose reaction to the inventory the deviating player faces, so the row has no drift term. The market maker’s quoting against inventory sits in the trader’s cost, through , not in the drift of . The trader’s inventory price is its own future orders, each weighted by how much the quote it will face moves per unit of inventory, Against the competitive market maker and the row vanishes.

First-order conditions.

The market maker’s condition is corner (i). Its control has no direct map into the inventory, ; the quote moves the inventory only through the trader’s order, so the closed-loop map takes its place. Quoting a unit further below the belief does three things. It gives up a unit of margin on the flow already expected, . It attracts new flow , on which the current margin is lost. And it lowers the inventory by that new flow, which is worth the inventory price. At the optimum the three cancel, There is no wedge. With written out this is (6.8.20): a long market maker quotes below its belief to attract the buying that unwinds the position. The trader’s condition is the deviating player’s condition with the market maker as its naive set, corner (ii), Theorem 4.10 of Chapter 4 on the enlarged state. Besides earning the mispricing, an order moves the market maker’s inventory, priced by the inventory row, and its beliefs, priced by the belief prices as in Chapter 4. The trader sets its estimated mispricing equal to the expected cost of the two, The belief terms are those of (4.5.9) with the adjoints now propagating through the quote rule (6.8.20). Differentiating (6.8.24) in with the inventory and belief terms held fixed, since a privy trader’s filter does not move and the inventory has not yet, gives the contemporaneous loading . In the transparent market only a trader seed is filtered, by the market maker; a quote seed moves no filter.

6.8.3The opaque market

What changes.

The trader now sees the quote but not the flow, so it cannot separate a quote off the rule from a quote on the rule after unusual noise. That holds for a quote shaded smoothly, which is absolutely continuous against the diffusing rule quote. A blip is a jump and would be seen; it enters only as the basis from which the response to a smooth shading is assembled by linearity. For an origin- seed the naive set is . The displaced state acquires a second column beside the inventory displacement, the displacement of the trader’s estimates of every past shock, which evolves by (6.5.3)–(6.5.7) with one difference in where the seed enters. The quote is an observation of the trader’s. Knowing its own order, the trader reads the noise flow exactly from the quote’s martingale part , with the market maker’s loading on order flow, whenever , so a quote seed enters the trader’s filter through that channel rather than through a state displacement. Because is observed directly, its movement in the signal drift cancels from the trader’s innovation, as in Chapter 4. The trader reads a shading as the noise flow that would have produced it, , a density born on the window, which the birth term (6.5.7) carries. Its perceived inventory and its perceived market-maker belief then move by the on-path maps, and , while the actual inventory has not moved: the trader expects an unwinding that is not coming. The trader no longer responds through the response kernel; it applies its policy kernel to the displaced estimates, its estimate of the inventory among them, The market maker’s own gain system acquires the rows dual to the new column, and , computed below on the lag window. Paired with the seed’s belief displacement they give the market maker a wedge of the kind the trader has against it, the value of the belief displacement a shading leaves in the trader’s filter, The market maker’s first-order condition is then the transparent one (6.8.23) with the inventory price projected onto the market maker’s information and the wedge subtracted, The inventory price is projected now, since the market maker’s belief is no longer public. In the last term the market maker shades its quote to make the trader believe noise traders have been active, and is paid for it through the trader’s later orders. The trader’s condition (6.8.24) keeps its form, its origin sets being unchanged, but every kernel in it is different, because the market maker it faces now sets its quote with the wedge in its first-order condition.

6.8.4Computation and comparison

Table 6.2. The stationary transparent market at (, , unit signal gain and noise, , 60 cells). At the loadings on inventory vanish and .
quote’s inventory loading
trader’s quote-fixed inventory loading
inventory’s closed-loop drift
belief correction to the inventory loading
loading on order flow
net price impact
The transparent market, computed.

The stationary form of the transparent market is solved on the lag window of Chapter 4, with the trader’s problem exactly as there and the market maker’s quote loading fixed by the origin- system above in its stationary form. At discount rate the inventory gain is , the market maker’s Riccati equation is algebraic, and the belief row is a linear equation in age with one nonlocal scalar , the correction to the trader’s inventory loading from how the market maker would filter a departure from it; is then a root of a scalar equation, found by tracking the root in upward from the competitive rule. The average-cost limit is degenerate: at the trader’s inventory value equals , its inventory loading responds to with the same slope as the market maker’s rule, and the rule cannot pin . At , , unit signal gain and unit noise, a quarter of the impact a trader faces at is inventory, not information: the market maker’s loading on order flow barely moves, from to , while the net price impact of an order, , rises from to . Table 6.2 collects the equilibrium at ; over the range computed the quote loading is linear in the inventory weight, for (Figure 6.2). At the computation returns the competitive kernels of Chapter 4 to , and refining the grid from to cells moves by percent. The gain equations admit a second self-consistent gain at , with and the trader unwinding inventory it has no reason to touch; the exact best response rules it out, so the competitive rule is the only fixed point of the loop on this branch. Above the trader’s second-order form goes indefinite on this grid at , the same boundary as in Chapter 4, and a larger trading cost is needed to certify larger inventory weights.

The opaque market, computed.

On the lag window the market maker’s problem after its blip, with the trader reading the shading as noise flow, is an ordinary linear-quadratic problem. Its state is the inventory and its own shading history, since the trader’s displaced estimates are a function of the shadings; its control is today’s shading; and the trader’s response to a shading path is its on-path kernel applied to the noise flow it infers by inverting the quote rule, with the signal drift shifted to match. A stationary Riccati equation in unknowns gives the quote’s inventory loading, and and are its inventory–history and history–history blocks. With the trader’s privy response in place of the inference, per unit the quote is raised, the same equation returns the transparent loading to percent at cells. At the opaque loading is against : the market maker loads its quote on inventory forty percent less when the trader cannot see the flow, and the ratio is over the whole range, (Figure 6.2). The reason is the trader’s response to a shading. Read as noise flow, a quote one unit off the rule draws of the order it draws when seen for what it is, and afterwards the trader trades the other way, in total over the window, toward the unwinding it expects and that is not coming. Shading is a weaker tool for managing inventory in the opaque market, so the market maker uses less of it: the trader’s inventory loading is against , the inventory unwinds at against , and the net price impact is against . The market maker also loads its optimal quote on its own past shadings, on the most recent and decaying, a loading that never fires on the equilibrium path.

Figure 6.2. The stationary transparent and opaque markets against the inventory weight (, , unit signal gain and noise, , cells). Left: the quote’s inventory loading . Center: the trader’s quote-fixed inventory loading and the inventory’s closed-loop drift. Right: the market maker’s loading on order flow and the net price impact .

The solvers and figures of this section are the work of Claude Fable 5, a large language model built by Anthropic [7], working under the author’s direction. Interactive versions of the dissertation’s computations are collected at sbabichenko.com/noisestate.

What the pair identifies.

Any difference between the two equilibria is the price of monitoring. Publishing order flow switches the market maker’s wedge off, but it does not discipline the market maker: it loads its quote more on inventory in the transparent market, because a trader who can see the flow can be steered. Most of the gain from being monitored is real inventory efficiency, and the trader’s orders carry the same amount of the fundamental into the price in both markets.

Table 6.3. The two markets at , at the parameters of Table 6.2. Price profits are gross of a trading cost of in both markets. The information concession is paid through the loading on order flow and the inventory concession through the quote’s inventory loading. The market maker’s loss is flow plus inventory penalty.
transparent opaque
trader’s price profits
noise traders’ loss
information concession
inventory concession
market maker’s total loss
expected unwinding rate

Table 6.3 gives the accounts at . The noise traders buy at quotes their own flow has pushed up and sell at quotes it has pushed down. The information concession is the same in the two markets to rounding ( against ), so the inventory concession carries the difference. Transparency transfers about from the noise traders ( against of inventory concession), split evenly between the two strategic players, and separately saves the market maker of real inventory cost, since a trader who can see the flow unwinds the inventory for it.

Which relation holds is an institutional assumption; it does not follow from quotes being actions. A specialist who posts a schedule is monitored; a market maker who prices after seeing the flow is not, and the opaque case is the natural one for the batch auction of Kyle [71]. Existence is known only at the corners; both equilibria above pass the trader’s second-order check of Chapter 4 at each computed , the opaque one is the fixed point of the market maker’s Riccati equation, and both rents vanish continuously as . The nearest benchmarks are the monopolist specialist of Glosten [43] for the transparent market, the transparency comparisons of Madhavan [80] and Pagano and Röell [93], and the manipulation of Allen and Gale [3], with the roles reversed, for the opaque one. The comparison is an old policy fight: the London Stock Exchange moved block-trade publication from immediate to next-day and then to a ninety-minute delay between 1987 and 1992, dealers arguing that delayed publication protected the unwinding of their inventory [42].

6.9Conclusion

Allowing some players to monitor a deviation’s origin extends the Chapter 1 deviation calculus to an augmented forward system with two channels. In naive filtering, a naive player misattributes the deviation to primitive shocks and contributes an information wedge. In privy response, a privy player observes the deviation as a labeled zero-variance coordinate and responds through a new kernel. Under perfect transitive monitoring (Assumption 6.4) the seed responses are linear and superposable (Proposition 6.11), and the seed-dependent forward-backward system reduces to a single backward gain sweep that serves all seed times. Section 6.8 works out what monitoring means in a Kyle–Back market.

One direct extension combines this chapter with the delayed public signals of Chapter 2. Suppose an -origin deviation at time is initially hidden in an aggregate signal but is identified by a reporting rule at time . Before the report, player is naive to the deviating player, so the seed moves ’s ordinary noise-state estimate and is priced through the information wedge. At the report, the new information does more than add another measurement row. It attributes an already propagated seed to player . In the resulting “naive until disclosure, privy afterward” system, origin attribution makes the same delayed transition from hidden influence to disclosure that Chapter 2 derives for signal manipulation.

Two other combinations concern the scale of the calculation. First, translating the rectangular gain into the stationary lag coordinates of Chapter 3 is the bridge to the general strategic-market-maker extension; Section 6.8 does the scalar case. Second, if the dynamics, costs, and transitive monitoring preorder on a finite network are invariant under the graph automorphisms of Chapter 5, the responder–origin gains should inherit an orbit reduction. The stationary rectangular system and the equivariant monitored-response reduction remain open.

The blip convention also makes the chapter a step toward modeling irrationality. A blip is an accidental deviation of variance zero. It is priced, but it occurs with probability zero, so equilibrium behavior never anticipates it. Treating irrationality formally would mean giving blips a variance, which raises the questions of what law the blips follow, whether the naive and privy channels survive when attribution is probabilistic, and how the equilibrium deforms as the blip variance grows from zero.