AXOQUANT

AQ-04 · 2026-08 · working paper · self-audit · current

We ran the pipeline on data with nothing in it. Every panel produced a survivor.

A gate that has never been run on data with no edge in it has an unknown false-positive rate. A kill verdict from a search with an unmeasured detection floor is a statement about the search, not about the market. Both were true of ours. This quarter we measured both. Neither answer flattered us.

On the false-positive side: run the whole pipeline — optimise, gate, forward — on synthetic series calibrated to a real market's moments and containing nothing to find. At full search width it returned a best-of-search survivor on every panel it was given, with manufactured terminal returns in the hundreds of per cent after costs.

On the false-negative side: plant a dislocation of known size and sweep it upward until the gate sees it. No cell in the first batch reached the conventional eighty-per-cent power standard at any dislocation we planted. The floor sits above what most contractual obligations pay.

And the disclosure that governs everything below. The generator producing these numbers has not yet passed its own out-of-sample moment-matching test. That table is owed as of this writing, and the campaign resting on it is under re-audit. So what follows is published as a measurement of our pipeline, not as a validated false-positive rate.

§ 1

Every edge-free panel produced a survivor

Until this quarter, every control we ran was a real mechanic on real data: composition-matched, shuffled leg, swapped leader, sector-substituted. Such a control shares the market's genuine structure — trend, volatility clustering, cross-sectional dispersion, the lot. So a control that earns is not evidence of a broken treatment. It is evidence that the control was never a null. A momentum control in a trending quarter earns because momentum works in a trending quarter.

Formally, a matched control answers is this book better than other books like it? It cannot answer is this book better than nothing? We had been reading answers to the first question as though they answered the second. The distinction is set out in full, with the doctrine reversal it forced, in AQ-11. Three incidents in one quarter, one cause:

Control- and placebo-named books occupied the forward right tail at or above their base rate — 7 of one month's top ten against a 35% base rate — and their median forward return beat the real books' in 3 of 3 panels. A composition-matched placebo passed identically to the treatment it was built to null, at p = 0.008. And a set of books we had been treating as nulls turned in strongly positive forward results: they were never nulls, because every instrument in them carried a prospectus-mandated daily rebalance — a named, dated, contractual obligation — and the field that should have caught that was empty.

The last of those has a second half worth publishing. Re-measured properly, those books were artefacts twice over: the pair was mis-specified against an unhedged benchmark, so roughly half the residual variance being "harvested" was uncontrolled currency basis, and one leg carried an unadjusted corporate action. The gate's rejection of them was correct — but for reasons nobody had established, on evidence nobody had checked. Being right by accident is not a working instrument.

So we built the market-level surrogate the literature asks for, and had never built. A regime-switching generator is fitted to a market's real moments and sampled to reproduce them — unconditional volatility, volatility-clustering persistence, fat tails, cross-asset correlation structure, the trading calendar — while containing no exploitable relationship. The surrogate-data requirement is explicit and is the standard we held it to: preserve the nuisance structure, destroy only the hypothesised signal. Twenty-four papers in our corpus bear on that construction.

Then run the pipeline on it. Across 6 independent market panels, the search returned a best-of-search "survivor" on every one. Terminal returns on the manufactured winners ran into the hundreds of per cent after costs, on data with nothing in it. The per-panel magnitudes are withheld: as an attributed table they map our coverage to our search's weakness.

A second pass rebuilt the null on each cell's own forward window rather than a common one, which is the harder test. Of the treatment arms carrying valid forward evidence, 0 of 10 cleared even the median of its own manufactured-survivor distribution. Two of the twelve cells carried no valid forward evidence at all — their winning parameterisation's lookback was longer than the window it was scored on — and their zero rows were placeholders, not measurements. We report that as a defect in the batch rather than as ten of twelve.

Reading

The correct reading, before anyone reaches for the wrong one

Printed beside the nineteen books that survived our specificity gate, the paragraph above invites an obvious inference: that the nineteen are manufactured. We would rather answer that here than leave it to a later page.

The rate is best-of-search on a single panel, measured before the specificity and forward legs run. The downstream nulls are what remove these, and they demonstrably remove some — a matched control ensemble caught a placebo beating its own original on two of three windows, which is exactly the job. What is being reported is not a broken gate.

Nor is it a clean bill. Turned on the surviving set itself, the same instrument left most of those books below their own manufactured-survivor bar: of the nineteen, 2 cleared, 3 were marginal and 14 sat below. Every approximation in that re-score made the bar easier to clear rather than harder, which is what makes the below-bar rows worth taking seriously and the clears worth replicating at full fidelity before anyone leans on them. We publish it as an open question rather than a verdict, because the instrument that produced it has not yet passed §3.

Manufacture happens upstream of the gate. That is a different problem from a broken gate, and it is not fixed by a better gate. The cost we are reporting is optimisation, gate, forward and review capacity consumed before the catch — not, by itself, verdict validity.

It is still the more expensive problem of the two, because it scales with search width and we control search width. That is §2.

FIG. 1 — SURVIVOR MANUFACTURE ON EDGE-FREE DATA
wide search 6 of 6 narrow search withheld 0 — manufactures nothing 100% of panels
Share of edge-free synthetic panels on which the search returned a best-of-search "survivor". At full width, 6 of 6 independent market panels, each generator calibrated to reproduce its market's unconditional volatility, volatility-clustering persistence, tail index, cross-asset correlation structure and trading calendar, and containing no exploitable relationship. The narrow mark is drawn outlined and to rank, not to scale: the mechanism-named narrow search manufactures survivors an order of magnitude less often, and that share, the two configuration counts, and the per-panel manufactured magnitudes are withheld — together they would map our market coverage onto our search's weakness. Manufactured terminal returns ran to the hundreds of per cent after costs. The generator behind both rows has not yet passed its own out-of-sample moment-match test (§3), so this figure is a measurement of our pipeline, not a validated false-positive rate.

§ 2

Width, not depth, drives it

The obvious first hypothesis is that the search is simply too big. It is a good hypothesis and it is only half right. Cutting a panel's configuration count by an order of magnitude cut the manufactured-winner share by an order of magnitude too — and the bar it manufactures shrank without collapsing. Power against the original reference stayed near zero throughout. The failure is selection-limited, not search-size-limited. The two configuration counts are withheld; the ratio between them is the part that generalises.

The narrow instrument had to be repaired before it could be believed, which is itself the finding. On split-unadjusted closes the same narrow search manufactured a triple-digit-per-cent winner out of four corporate actions. Only after back-adjustment at read did its null bar mean anything. A related artefact in the same store: on unadjusted closes, two funds tracking the same index print an ex-dividend sawtooth into their spread, so the crossings a pair strategy trades are dividend-calendar events rather than dislocations. Those pairs outscored the genuine share-class mechanism they were benchmarked against.

So the response is not simply to search less. It is to name the mechanism before searching, and to price every trial you did try. Width is a decision about how many hypotheses you are willing to pay for, and every one of them raises the bar the real edge has to clear. That is the deflated-Sharpe argument stated operationally rather than as a footnote: the correction is not a post-hoc adjustment to a number, it is a constraint on how the number was allowed to be produced.

§ 3

The instrument's own validation is outstanding

One assumption carries this whole apparatus: that a calibrated generator can be edge-free in the way we need. If its regimes are mis-specified — too stationary, wrong correlation dynamics, a missing stylised fact — then a survivor on synthetic data proves nothing about real data, and a failure to survive proves nothing either. The generator becomes a new source of silent-zero, which is this house's dominant failure mode: a component that reports success while doing nothing.

The mitigation was written into the charter in the same paragraph as the instrument, because a mitigation named later is a mitigation not built. Fit the generator, sample it, and confirm out-of-sample that the synthetic series reproduces the real market's moments — volatility-clustering persistence, tail index, cross-sectional dispersion, autocorrelation at the horizons we actually trade. A generator that cannot pass its own moment-matching test is not a null, and any verdict resting on it is void.

That table is owed as of this writing. The first batch of power cells does not publish one, and the campaign resting on those cells is under re-audit. Where the moment-match is missing, the batch is void rather than probably fine.

We publish that rather than wait for it, because the alternative — printing a false-positive rate and quietly leaving its instrument unvalidated — is the failure this entire page exists to describe.

Two smaller instrument findings in the same spirit. Single-draw validation of the generator fails at window scale; the authoritative reference is an ensemble band across twenty-five draws, not one draw, and the single-draw form is retired. And in one re-scoring run a cell returned an explicit validation failure instead of a number: the books resting on it were scored on the remaining cells and the failure was carried loud through the record. An instrument that can say I don't know is worth more than one that always answers.

Which is why the synthetic null is registered in our survivability gate born-informational: it is computed and reported in every verdict, and its failure blocks nothing. Before it earned even that, it had to reproduce 3 of 3 documented verdicts on a book whose answer was already known. The pre-registered flip to a blocking leg waits on its proof against a full re-score of the incumbent book set. We do not let a new instrument set a floor on the strength of its own first results.

§ 4

The other half: a kill without a detection floor is not a finding

The first generator buys a false-positive rate. A second one buys the number nobody measures: it generates the same calibrated series but injects a known relationship with tunable parameters — a cointegrating vector with specified half-life, band width and noise; a bounded process with a known post-lock drift; an episodically defended band of specified width and frequency — and sweeps the planted dislocation upward from zero until the gate starts finding it.

Batch one covered two mechanism classes across two market panels, sweeping residual scale against two mean-reversion half-lives with eight seeds per cell. No cell in the batch reached the conventional eighty-per-cent power standard anywhere in the sweep. Against the dislocation scales those two mechanisms actually offer, the search had little power at all. The power values, the minimum detectable edge and the detection floor in absolute units are withheld; the direction and the fact of measurement are published.

The operational consequence is a rule that now binds every verdict we issue, and it is retrospective rather than prospective-only:

Cite your cell. No fund, forward or kill verdict may be issued without naming the class cell it rests on — mechanism class, market panel, timeframe — and that cell's power at the dislocation actually on offer. A missing cell is a prerequisite batch, not a kill. Verdicts are restated as either a true kill or below our floor. The obligation applies to closed campaigns, not only to new ones.

A harness that judges everything must itself be judged, and by something other than its own output. The four checks that make it credible rather than decorative:

Coverage of the power-cell programme, as of 2026-08-02. Which market panels each cell was measured on is withheld, as are the power values and the detection floor. Batch one is published with its generator's moment-match table outstanding (§3).
Mechanism class Null instrument Power cell
Fund-rebalance mandate (geared and inverse tracking) mean-reverting pair harness measured, batch 1
Defended band, episodic intervention mean-reverting pair harness measured, batch 1
Event-driven classes (cohort events, cascade events, session-open dislocations) own event-study nulls routed out — the pair harness cannot see them by construction
All remaining classes outstanding — a verdict citing one of these is a prerequisite batch, not a kill

§ 5

Our floor sits above what real mechanisms pay

The false-positive side of this gate is well instrumented. Its false-negative side was, until this quarter, unmeasured — and it is the expensive one. A false positive costs capacity. A false negative costs the edge, silently, and leaves a confident kill verdict in the record where an unanswered question should be.

A contractual obligation pays what it pays, and for most of them that is a small number. We enumerated 35 of them from primary documents — fund prospectuses and central-bank pages, each row carrying a verbatim-verified quote, 24 fund-rebalance mandates and 11 defended-band commitments — and measured 33 of them against our own cost model. 17 coverage holes are named in the same record, because a named absence is a finding and an unnamed one is a silent zero. With few exceptions the dislocation on offer sits below what a broad search can resolve. The dislocation sizes and the floor are both withheld in absolute units; printed together they would locate the floor as precisely as printing it.

The consequence for how search is allocated is uncomfortable enough to be worth stating plainly: a tournament is the moonshot channel, not the discovery channel. The corroboration is more uncomfortable still. Our best validated obligation-linked book was found by mechanism-first authoring — reading the obligation, pricing the dislocation, then writing the book — and not by the search that was supposed to find it. It sits below that search's floor and always did.

The floor exists and we have now measured it. Where it sits is proprietary, and is withheld: an absolute detection floor is a map of what we cannot see, which is worth more to a competitor than to a reader.

§ 6

What changed

That last item is the one we underestimated, and it deserves its own numbers. One whole-market daily equity feed in our store carries no delistings whatsoever across four years — not survivorship bias but survivorship absence, which independently blocks merger arbitrage, index deletions and every cross-sectional book on that feed. The corporate-action table behind it is empty, so its prices are unadjusted. The demonstrated cost of that one defect: a screen run against those unadjusted prices returned a positive mean net of costs across 111 trades, with the large majority of individual lines positive and test statistics clearing every conventional threshold — entirely a distribution sawtooth. 65% of its entries fell within five days of an ex-distribution drop against a 17% base rate. Re-run on a distribution-adjusted series, the same signal fires three times in four years and the mean turns negative. The two mean returns and the t-statistics are withheld as an attributable set — published together with the mechanism and the window they locate the screen; the counts, the base rates and the sign flip are the finding and are published.

Nothing about that screen was a research failure. It was a data defect wearing a t-statistic, and the pipeline had no instrument that could tell the difference until we built one.

Stands on

Cited by

Cross-references are hand-maintained; a link check runs before publish.

Revision history

Disclosure

This item publishes the direction, the counts, the apparatus and the policy. Withheld: the two configuration counts, the per-panel manufactured magnitudes, the power values, the minimum detectable edge, the detection floor and the census dislocations in absolute units, which market panels each cell was measured on, and the mean returns and t-statistics of the artefact screen in §6.