Projects  /  Decision Mesh

Decision Mesh

The odds of a coin vary over this square. A few thousand sites each flip a coin a handful of times, and each site's share of heads mixes the odds with the coin's own effect and the noise of the flips. The estimator refines a mesh of right triangles or rectangles only where a false-discovery gate finds support in the flips, and stops when no candidate passes.

My post on partial pooling takes up the same trade-off between detail and the data each cell has.

Ready
The odds what the coins really do
The flips each site's share of heads
The fit dots: admitted, by round

Try giving the hill and the pit more flips and watch the mesh fill in around them; pick “nothing at all”, which it leaves uncut at any number of flips; pick “your drawing” and paint on the odds (drag to raise them, tick “lower” or hold shift to lower them).

How the Gate Decides

The drawings below come from one run of the estimator in your browser, on coin flips drawn on this page.

Fitting…

The odds of a coin vary over this square. Six thousand sites each flip a coin about twenty times. Each dot is one site’s share of heads.

Each coin also has an effect of its own on its log-odds, with the standard deviation set by the slider above. Half the sites are held out: the estimator never sees them, and they are used to score the fit at the end.

For a binomial, the variance is determined by the mean. If you know a coin’s odds, you know how much its share of heads should vary, so any extra variation can be measured.

Against the fitted odds, … of sites fall outside the 95% lines. With no coin effects it would be about 5%.

The odds themselves have to be estimated, and an error in the mean also shows up as extra variance. A site’s error has two parts: the fitted surface can be off from the true surface, and the two axes may not capture everything that determines a coin’s odds.

Both parts add to the variance in the same way, growing with the square of the error. Since this page generated the coins, both parts are known. Even the true surface is off by the part the axes don’t capture, which here is the coin effect.

So there are two kinds of overdispersion. Coherent: neighboring sites share it, and the surface should change there. Incoherent: each site has its own, and no surface can fit it.

An error in the surface is shared by neighbors, since they sit under the same part of it. The part the axes miss is specific to each site; anything in it that varied along the axes would belong to the surface. Against a flat surface the mean squared residual is … and the large residuals cluster; against the final fit it is … and they don’t. The mesh is refined for the coherent part. The incoherent part is the coin variance, which the gate uses as the baseline level of disagreement between sites.

The fit starts from a coarse mesh of right triangles, with only its coarsest vertices free: the four corners and the first midpoints, nine in all. A triangle can only be split at the midpoint of its longest edge. If that would leave a vertex in the middle of a neighbor’s edge, the neighbor is split too.

The fit starts from an 8 × 8 grid of rectangles, with the corners and the two midlines free. A rectangle can be split along either midline, so both directions are offered as candidates and the gate picks. A vertex in the middle of a longer edge takes its height from that edge, so the surface stays continuous.

Each segment drawn here has a candidate vertex at its midpoint.

A vertex that isn’t free sits at the average of its two parents, the endpoints of the segment it splits. A free vertex can differ from that average by its surplus.

Admitting a candidate means allowing its surplus to be nonzero. A prior shrinks each surplus toward zero, more strongly for deeper vertices, so fine corrections need more evidence than coarse ones.

Each round scores every candidate against the current fit: would a nonzero surplus there improve it? The score is a weighted sum of the residuals of the sites under the candidate’s tent.

Blue means the data pull the surface up there, orange means down; larger dots are larger scores. Round 0 scored … candidates. Over the whole run, some more had too few sites under them to be scored.

The raw score is biased, and three corrections are applied. First, each residual is centered: shrinkage and the discreteness of the counts give it a nonzero mean even when the surface is right.

Second, the part of the candidate’s tent that the existing coefficients could absorb is removed (a first-order refit), so a candidate isn’t credited for a slope its parents could fit. Third, the bias from the prior’s shrinkage is removed. Dividing by the standard deviation gives the candidate’s z.

The gate looks at all of a round’s scores together. If no candidate mattered, they would follow a single bell curve, the null.

Lindsey’s method fits the density of the scores and estimates the null from its center, so a round whose scores are shifted or spread out is compared with its own null rather than N(0, 1). Here the null has center … and spread …, and … of candidates are estimated to be noise.

From the null, each candidate gets a local false discovery rate: the probability it is noise given its score. Sorted best first, the gate takes the longest run whose average stays at or below 10%.

In round 0 that is … candidates. A cutoff of |z| > 1.96 would have taken ….

The selected candidates are admitted one at a time, best first. Each is re-scored against the surface as updated by the earlier admissions, and admitted only if it still improves the fit.

…

The next round scores the new candidates against the new fit. As the signal is used up, the scores look more like noise and the null gets closer to N(0, 1).

After three rounds in a row that selectadmit nothing, the gate stops. This run took … rounds.

Two variances are refitted between rounds. The coin variance is the incoherent part, the spread around the surface. As the mesh picks up the coherent part, it falls toward the true variance of the coin effects.

The depth variance controls how large a surplus at each depth tends to be. It is estimated by EM from the admitted surpluses, averaging their squares plus their posterior variances by depth. Including the posterior variance keeps a depth’s variance from getting stuck at its floor.

The final surface is flat where the data show nothing and refined where they show a pattern.

On the held-out half: …. (A site’s deviance is twice the log of how much likelier its flips are under its own share of heads than under the fit: zero only for a perfect match.) Changing any setting above refits and redraws every step.

The run: 6,000 sites on the unit square, each flipping a coin about 20 times (between 10 and 30), with odds from the picture chosen above plus a normal effect per coin on the log-odds (standard deviation from the slider), fitted by triangular-decision-mesh or rectangular-decision-mesh compiled to WebAssembly. The scores, nulls, false discovery rates, admissions and variances drawn here are the engine’s own, from a trace of each round’s candidates; the trace is the builds’ one addition, and it changes nothing the fit computes. The fifth drawing is a one-dimensional sketch of the corrections, not the run.

Skip past the story

Why Right Triangles

Below, four fits grow on the same noisy data, each given the same number of free parameters. None of them uses the false-discovery gate. Each grows by the same greedy rule, so the only difference is the geometry.

Ready Starting when the meshes come into view.
Freeform triangles
Right triangles
Rectangles
Regression tree
Truth the surface under the noise

On this one draw of the data, the right triangles get closest to the truth on five of six test surfaces. A new draw can change the order. The freeform engine is a line-for-line JavaScript port of my Decision Mesh. The right-triangle and rectangular engines are written for this page, grown by the same greedy rule as the freeform triangles.

Round-by-Round Results

Each round scores every candidate vertex (the midpoints of the mesh's edges) by how much it would improve the fit. If no candidate mattered, the scores would look like draws from one bell curve, the null. The gate fits the density of all the scores, reads the null's center and spread off the middle of it, estimates the share π0 that are noise, and admits the candidates that stand out from that null at a controlled false-discovery rate. When the scores have no central peak to read a null from, it falls back to the textbook one. Each coin also has an effect of its own on the log-odds, and the coin variance is the estimator's estimate of their spread, fitted alongside the surface. As the mesh learns the odds, what it can't explain shrinks toward the spread the slider set.

Fit Your Own Data

A CSV with a header row: two numeric columns for the position and a column of 0/1 outcomes, one row per trial. If each row already counts several trials, give its trials and its successes instead. The file is read in this browser and never uploaded. The sample file shows the format.

The estimators are triangular-decision-mesh and rectangular-decision-mesh, compiled from their C++ with Emscripten. On my test designs the browser builds reproduce the native meshes to 1e-11. The two share the gate and the model for the coin effects and differ in geometry: right triangles cut only through their hypotenuse, with a plane in each; rectangles cut across either axis, with a bilinear fit in each. Half the sites are held out: the mesh never sees them, and the deviance card scores the fit on them. A site’s deviance is 2[k log(k / n p) + (n−k) log((n−k) / (n(1−p)))], with k heads in n flips and fitted odds p: zero when p is exactly the share of heads, and larger the less likely the flips are under the fit. Even the true surface scores above zero, because each coin has its own effect and flips are random, so it sets the floor. The coins are drawn on this page from the odds you pick, plus a normal effect per coin on the log-odds, with the standard deviation on the slider (0.25 to start). Set it to zero and the coins are plain binomial.

Cite it
@software{triangular_decision_mesh,
  author = {Babichenko, Sam},
  title  = {triangular-decision-mesh: an adaptive right-triangle mesh regression with a false-discovery gate},
  year   = {2026},
  url    = {https://github.com/sbabichenko/triangular-decision-mesh}
}

@software{rectangular_decision_mesh,
  author = {Babichenko, Sam},
  title  = {rectangular-decision-mesh: an adaptive rectangular mesh regression with a false-discovery gate},
  year   = {2026},
  url    = {https://github.com/sbabichenko/rectangular-decision-mesh}
}