BOUNDEDNESS ATLASTHE MURRAY RESEARCH PROGRAMME
Imagined cabinet of luminous specimens and brass instruments

legacy manuscript · 6963958

Diploid boundary identifiability

Rare alleles expose hs; trajectory state dependence supplies information separating h from s.

← Back to the library

Dominance Hides in the Bend: Boundary Identifiability and Sampling Design in Diploid Selection

Exact coefficient inversion; information analysis; simulations

Current scope. Rare-allele limit identifieshs; bend exposes additional coefficient under declared model.

What it adds to the whole

Rare alleles expose hs; trajectory state dependence supplies information separating h from s.

Predictions and research connections

The abstract

Supplied manuscript · PDF page(s) 4. Original wording; read alongside the scope note.

### PDF page 4 Dominance Hides in the Bend: Boundary Identifiability and Sampling Design in Diploid Selection Daniel J. Murray June 18, 2026 Abstract Estimating both the strength s and dominance h of selection from allele-frequency time series is notoriously hard while the favoured allele is rare, because selection then acts almost entirely through heterozygotes and the two parameters enter only through the product hs. We make this exact and show how it resolves. In log-odds coordinates, where the diploid selection update is additive with a state-dependent increment, that increment expands as g(x) = a0 + a1x + O(x2) with a0 = ln(1 + hs) and a1 = s(1 − 2h − h2s)/(1 + hs). We prove three results: the rare-allele boundary identifies only a0, hence only hs (Proposition 1); dominance enters the first state-dependence coefficient a1, whose weak-selection leading term is proportional to (1 − 2h) (Proposition 2); and the pair ( a0, a1) identifies s and h in closed form, with Jacobian determinant −s/(1 + hs)2 and explicit inversion s = ea0 a1 + e2a0 − 1, h = (ea0 − 1)/s (Proposition 3). The statistical counterpart, under a deterministic-trajectory binomial observation model, is that the design information for ( s, h) is nearly rank-deficient near the boundary—not because the coefficient map is singular there (it is not), but because boundary-local data express only a0; we prove the resulting score collinearity analytically (Corollary 1). The information for h after profiling out s then rises by orders of magnitude across the bend. We demonstrate the inversion directly on simulated Wright–Fisher pool- sequencing data, give the closed-form sensitivity of the recovered h to estimation error in a1, and show by simulation that designs confined to the boundary regime cannot recover dominance even at tenfold sequencing depth, whereas mid-frequency designs succeed. The resulting rule is simple: to estimate dominance, sample across the frequency range where the trajectory bends, not where it begins.

Conclusion or closing discussion

Page addresses are retained in the excerpt. These are author claims, not an independent validation certificate.

Open the closing section
PDF page 10 Figure 4. Sampling design, with an increased-depth control. Left: mid-frequency sampling (blue) recovers dominance with quantified uncertainty; sampling the same process near the boundary (red) fails for alleles that remain rare. Right: usable recovery fraction (final sampled frequency > 0.5 for mid, > 0.2 for rare; 60 replicates per condition) for three conditions—mid-frequency at depth 150, rare at depth 150, and rare at depth 1500 (tenfold)—shown per dominance value. Mid-frequency sampling yields high; rare sampling collapses near the boundary, most severely for recessive variants, and the tenfold-depth bar shows that depth does not rescue it. Dominant alleles ( h ≥ 0.5) leave the boundary quickly and are recovered under all three. beyond diploid selection. Relation to existing work. Structural identifiability analysis asks when parameter combinations rather than parameters are determined by data (Walter & Pronzato 1997; Raue et al. 2009); Proposition 1 is a sharp, boundary-localised instance and Proposition 3 its resolution through the state-dependence coefficient. The phenomenon of individually ill-constrained parameters with well-determined combinations is the “sloppy models” programme (Gutenkunst et al. 2007); here the boundary ridge is the fixed- hs direction, and the information that breaks that ridge is localised in the coefficient a1. The question of where to sample is the subject of optimal experimental design, in which D- and A-optimal designs for logistic and Michaelis–Menten models place points away from the asymptotes toward maximum curvature; “sample the bend” is the dominance-specific dynamical case. In population genetics, the difficulty of jointly estimating s and h from allele- frequency trajectories is well documented in time-series inference under Wright–Fisher dynamics (Malaspinas et al. 2012; Tataru et al. 2017; Taus et al. 2017; Paris et al. 2019). Previous likelihood methods can estimate ( s, h) from sufficiently informative trajectories; the present result identifies the local coefficient that contains dominance and explains why early-frequency data fail even when hs is well estimated. The least-squares fit on log-odds used here is itself a simple full-trajectory method; more sophisticated HMM or diffusion-based inference would use the same curvature, and the inversion is offered to make explicit where that curvature lives. 8. Conclusion For diploid selection we prove that the rare-allele boundary identifies only the product hs, that dominance re-enters through the first state-dependence coefficient of the log-odds increment, and that the boundary value and that coefficient identify s and h in closed form. The design information makes the mechanism plain: near the boundary the data express only a0, the score directions are collinear (analytically, by Corollary 1), and dominance is unidentifiable; the bend expresses a1 and separates them. We demonstrate the inversion on noisy data, give its error PDF page 11 propagation, and show that the design implication—sample the bend, not the boundary—cannot be circumvented by sequencing depth. The rule is simple and practical: to estimate dominance, place samples across the frequency range where the trajectory bends. Methods Dynamics and propositions. Deterministic update from (1); Propositions 1–3, Corollary 1, the quadratic coefficient a2, the finite-interval expansion, and the sensitivity ∂h/∂a 1 verified symbolically (SymPy). Design information. Ft = C[∂sxt, ∂hxt]⊤[∂sxt, ∂hxt]/[xt(1 − xt)] summed over sampled times; Schur- complement information Fhh − F 2 hs/Fss; score-collinearity cosines (the cosine between ∂sxt and ∂hxt, as in Corollary 1) within frequency bands; sensitivities by central differences (∆ = 10−5, stable across 10−4–10−6). This is the expected binomial observation information along the deterministic trajectory, conditional on that trajectory; the Wright–Fisher simulations test survival of the prediction under process noise. Inversion on simulated data. Wright–Fisher drift (Ne = 5000), pool-seq depth 200, 120 generations from x0 = 0.02; Jeffreys pseudocount ˜x = (k + 1 2)/(C + 1); a0, a1 by least squares of per-generation log-odds increments on x over 0.02 < x < 0.6 (robustness: window x < 0.3 and weighted least squares with delta-method variances, Section 5); invert by Proposition 3; 200 replicates per h, estimates retained when finite and in [ −0.5, 1.5]. Recovery simulations. Ne = 2000, depth 150 (1500 for the increased-depth control), 80 generations sampled every 8, 60 replicates per h (all conditions, including the depth control); full-trajectory fit of (s, h, x0) by least squares on pseudocount log-odds, multistart over s0 ∈ {0.05, 0.1, 0.2}, h0 ∈ {0, 0.5, 1}, Nelder–Mead, tolerances 10 −4/10−8; usable = final sampled frequency > 0.5 (mid) or > 0.2 (rare); full retention in Table 1. True s = 0.1 throughout; initial frequency 0 .15 (mid) or 0 .02 (rare). Coordinate and null tests (Appendix A). Bending recomputed in probit Φ −1(x), arcsine-root arcsin √x, raw x; null trajectories: shifted logistic, power-law ( t/T )1.5, bounded random walk, all under the same bounds and binomial weight. Reproducibility. All figures and numerical values regenerate from a single self-contained script with fixed random seeds (Python 3, NumPy, SciPy, SymPy), provided as supplementary material with a reviewer copy at submission and to be archived with a permanent DOI on acceptance; it is written to be usable as a standalone sampling-depth calculator. A. Coordinate specificity and null dynamics Coordinate specificity. The information–bending alignment is read in the log-odds, the coordinate in which the update is additive. “Bending” is not coordinate-invariant, so this is explanatory geometry, not an identifiability theorem (which is Propositions 1–3). Recomputing the bending in coordinates that do not linearise the odds-multiplicative composition gives, for the same trajectory, correlations of information with bending of −0.24 (probit), 0.00 (arcsine-root), and 0 .66 (raw frequency), versus 0 .999 in the log-odds. Null dynamics. For the null trajectories the comparison uses the same binomial observation weight x(1 − x) and coordinate bending as an alignment diagnostic, not a dominance-information calculation, since the null flows contain no dominance parameter. Holding the bounds and the binomial weight fixed but replacing the selection trajectory with other bounded trajectories collapses the alignment: a shifted logistic gives −0.36, a power-law rise −0.72, a bounded random walk −0.12. A relationship forced by the coordinate alone could not be destroyed by changing the flow; the mechanism is that for the selection flow the bending region and the high-Fisher-weight region x(1 − x) co-locate at mid-frequency, while the null flows place them apart. Declaration of generative AI and AI-assisted technologies in the writing process During the preparation of this work the author used generative AI and AI-assisted tools to assist with drafting, figure preparation, and editorial revision. Symbolic and numerical checks were performed by author-directed scripts and independently reviewed by the author. After using these tools, the author reviewed and edited the content as needed, verified the mathematics, simulations, and references, and takes full responsibility for the content of the publication. No AI system is listed as an author.

Prediction-bearing source passages

A full-text retrieval aid, including hypotheses, falsifiers, comparisons and mentions of predictions. A matching passage is not automatically a distinct prediction.

PDF page 6
This is the expected binomial observation information along a deterministic trajectory; we use it as a design diagnostic, and the Wright–Fisher simulations below test whether the same geometric prediction survives process noise. The relevant scalar is the information for h after profiling out s, the Schur complement I prof h = Fhh − F 2 hs/Fss. The design matrix treats the initial frequency x0 as fixed; the full-trajectory simulations below estimate x0 as a nuisance parameter, which does four (Figure 1, right; the horizontal axis is the maximum frequency reached, with F accumulated over all observations up to that point). The prediction localises directly. Defining the bending of the path as the second difference of the log-odds, κt = |ρt+1 − 2ρt + ρt−1|, the Fisher information about h and κt peak at the same frequency (x ≈ 0.49 for s = 0.1, h = 0.5), and essentially none of the total information lies where the path does not bend: less than 10 −4 of it sits where κt is below one percent of its maximum (Figure 2). This co-location is a property of the selection flow, not of the log-odds transform:
PDF page 11
Corollary 1) within frequency bands; sensitivities by central differences (∆ = 10−5, stable across 10−4–10−6). This is the expected binomial observation information along the deterministic trajectory, conditional on that trajectory; the Wright–Fisher simulations test survival of the prediction under process noise. Inversion on simulated data. Wright–Fisher drift (Ne = 5000), pool-seq depth 200, 120 generations from x0 = 0.02; Jeffreys pseudocount ˜x = (k + 1 2)/(C + 1); a0, a1 by least squares of per-generation log-odds increments on x over 0.02 < x < 0.6 (robustness: window x < 0.3 and weighted least squares with delta-method variances,