AXOQUANT

Method

Three gates, in order. No idea skips one.

A strategy earns capital by passing three independent tests: S — an edge exists; P — the edge is specifically ours, not a factor in costume; L — the size it can actually carry. Ranking is lexicographic, P first: in our forward validation, books that passed both S and P returned +22.2% per window with a 100% win rate (n = 44 windows), while S-passers that failed P managed +4.3%. The top of the in-sample leaderboard is an anti-signal — composite scores ≥ 90 that failed P ran −1.0% forward.

Gate S

Survivability — does an edge exist at all?

Every candidate runs unlevered at fixed size — sizing is an L-question, and compounding artifacts have manufactured +230% mirages from −6% books. The gate reads a block bootstrap of the daily return path:

P( mean return ≤ 0 ) under stationary bootstrap resampling of the daily path

What matters is which statistics actually predict out-of-sample survival. We measured this on our own corpus, and the answer is uncomfortable for leaderboards everywhere:

We ranked every candidate statistic by its forward AUC on our own corpus. The two that predict are consistency-and-emergence measures; every magnitude measure — raw Sharpe, deflated Sharpe, composite score — predicts inversely. The specific ranking is proprietary.

So the gate ranks on consistency and emergence, never on magnitude. A front-loaded book that made half its P&L in the first fifth of the window and then reversed is floored regardless of its headline number: we measured that shape running significantly negative forward across hundreds of cells; its floor parameters are proprietary.

Gate P

Specificity — is the edge yours, or a factor in costume?

The load-bearing filter. Each candidate faces a control ensemble built to null exactly the mechanism it claims. A basket strategy is re-run on random same-size baskets; if random selections do as well, the chosen basket carried no information:

p = ( 1 + #{ controls ≥ treatment } ) / ( N + 1 )   N full re-runs, N pre-registered

Timing strategies face a different null — regression of forward returns on the market factor, testing the residual alpha — because a shuffle null destroys beta and mints false positives. Per-window p-values combine by Fisher's method (−2Σln p ~ χ²), and the family-wise verdict applies Benjamini–Hochberg at a pre-registered false-discovery rate, with consistency requirements across independent windows.

The gate proves itself on planted controls: every market we enter gets a deliberate factor strategy seeded among the candidates, expected to pass S and die at P. When the G10 carry basket arrived with a stable Sharpe of 1.6, the gate refused it at p = 0.62 — the carry factor, correctly named. If a market's planted control ever passes, every verdict from that market is void until the gate is fixed.

Gate L

Leverage — what size can it actually carry?

L* = min( bootstrap-sized, engine-margin ) taken at the worst window

Most survivors cannot carry even 1x — in our corpus, most fail to. Books whose relationship is enforced (a defended band, a rebalance mandate) take scenario-sized leverage instead: a bootstrap cannot contain a break that never occurred in its sample, so the 2015 franc-style gap is imposed on every peg book by hand. Capacity is measured with shared legs netted — overlapping instruments across books net capacity down, not up.

Evidence protocol