Reading list
This is the reference set our house papers cite and the set we check before reaching for a default. It is curated, not comprehensive. Nothing here was added because it is famous, and inclusion is not endorsement — most of what is on these shelves was read for framing and never built.
Shelf sizes are a map of where our problems were, not of the literature's importance. A bibliography is decoration. A bibliography with a verdict against each entry is a record of what we actually did, which is the only version worth publishing.
Five terms, closed set, defined once. Every entry carries exactly one. The verdict is assigned from our own run record, not from our opinion of the paper.
Most entries are Background, and publishing that ratio is the point. A reading list on which everything was adopted is a reading list nobody read; it is a list of citations assembled after the decisions were already made. The distribution is drawn below rather than asserted.
An entry prints what the store holds and nothing more. Where the author field is unresolved it prints an em dash — the ingestion recorded a title and sometimes a year and no author for a large share of this corpus, a defect described below and not papered over here. Public identifiers resolve by hand, one entry at a time, and publish with the entry; we would rather show the gap than fill it with a guess that a reader cannot check.
Entries sort by author where the store holds one, then by year, then by title. Used in appears only where a house item genuinely leans on the paper. The column is sparse on purpose. Cross-references on this site are hand-maintained and link-checked before publish.
Featured shelf
The public literature on machine agents trading, judging and negotiating in markets, together with the older wisdom-of-crowds and diverse-problem-solver results it rests on. We read this shelf because we run such agents: models author strategy code and generate hypotheses at volume across our estate. Naming it costs us nothing, because it is public science and none of it is ours.
The shelf splits. One half reports machine agents performing; the other reports multi-agent teams holding experts back, information leaking into apparent profit, and valid signals failing at regime boundaries. Our own measurement agrees with the sceptical half, and we ran it rather than asserted it.
Two judges from different model lineages were given the same high-scoring candidates and asked to call survive or fail from in-sample information alone. Both rejected everything. They agreed with each other completely and trivially, and they scored exactly the base rate of the sample they were shown. A second model adds no diversity when there is no discrimination to disagree about. The judges were demoted to advisory and consolidation moved into code — nothing a model emits adjudicates anything here. The accuracy figure maps directly onto our own corpus base rate and is withheld; the direction, the agreement and the policy consequence are published. The argument in full is in machines.
One Adopted on this shelf. The diverse-problem-solver result is why our review runs on models of deliberately different lineage rather than on more copies of one — and why, when we measured the ensemble and found no discrimination, we did not reach for a third model.
Featured shelf
Mostly the time-series foundation-model literature of the last four years: Chronos, TimesFM, Moirai and Moirai-2, Lag-Llama, PatchTST, iTransformer, TiRex, TinyTimeMixers, Time-LLM, and the survey that ties them together. We implemented from this shelf, pre-registered what would falsify the implementation, and decommissioned the programme on a dated day.
The deployed model was tested on three separate trading uses across five days and failed all three. As an entry gate it added no win-rate edge at honest trade counts, and the high-scoring configurations were artifacts of nineteen to twenty-eight trades, rejected by a minimum trade count written down before the run. A tabular classifier on the same causal features cleared no out-of-sample bar. And neither track improved the market-regime detector. The floors were pre-registered; their values are withheld, as all our cut-offs are.
One mechanism explains all three. The model sets its forecast band from the realised volatility of the context it is fed, so it reprices the recent past rather than anticipating the future. Measured over 607 four-hour bars, its predicted band width correlated with trailing realised volatility at +0.82, its correlation with forward volatility sat below the plain backward baseline, and its lead-lag profile peaked at zero to one day forward. A quantity that tracks the past coincidentally cannot lead it. The honest caveat, published with the verdict: the label overlap was 84 to 88 days over a single episode that ran mostly in one direction. Small sample — but the same-feature correlation and the lag-zero peak are not marginal, and they are the exact failure the mechanism predicts. The forecast server was stopped and disabled on 2026-07-11.
The distinction a careless reading list destroys: this is a negative about our use, not about the papers. They are competent work on a problem we could not make pay at our horizon, and the one branch left unfalsified — a fine-tune trained on forward realised volatility rather than next-bar price, judged on incremental correlation after the backward-volatility feature — has not been built by anyone here. The account in full is in decommission.
Nine of the twenty-one are Read, not pursued, and the reason is the same for all nine: one model from this family was implemented and the mechanism that killed it is a property of the family, not of the implementation. We did not run eight more experiments to confirm a mechanism we had already identified. That is a stated reason, not a verdict — if the mechanism argument is wrong, these nine are owed a test.
Structure, counts and the problem that sent us there. Entry-level rows publish shelf by shelf as verdict assignment completes; a shelf whose entries would be guesses publishes its count and says so. Each shelf carries an anchor so a house paper can deep-link a single citation.
Microstructure41
Fill models, queue position, order-book dynamics, impact and execution under latency. The largest shelf because execution is where our verdicts die: vectorised screens overstate edge, and our own fills ledger has voided our own results more than once.
Entries publishing.
Volatility33
Realised and implied estimation, stress and uncertainty indices, tail dependence. Volatility is the input to sizing, to regime labels and to every calibrated null we fit, so the shelf is about measuring it honestly rather than predicting it.
Entries publishing.
Regime18
Regime-switching and hidden-Markov dynamics, changepoint detection, state-space estimation. Read first for labelling exposure, then re-read for a different job entirely: our edge-free generator is specified as a regime-switching model fitted to a market's real moments. Specified, not yet demonstrated — the moment-match table that would show the fit holds is owed, and until it lands the generator is a design rather than a validated null.
Entries publishing.
Cointegration18
Long-run relationships, error correction, Hurst and Granger machinery. Read because a pair that cointegrates in-sample and not out of it is the most common false positive we generate, and the tests do not warn you.
Entries publishing.
Risk management11
Stops, trailing exits, drawdown control and restart rules. Read hard, because every consensus-conservative default we tested empirically moved in the deployable direction, and the stop-loss bundle was the largest single case.
Entries publishing.
Portfolio10
Construction, higher moments, hierarchical and risk-parity methods, and the ways funds fail. Read against our own evidence that de-correlating a negative-mean distribution compresses it toward its mean and destroys the right tail that carries the edge.
Entries publishing.
Outliers9
Anomaly detection in multivariate series, conformal methods, root-cause attribution. Read for the data-quality problem before the trading one: an unadjusted corporate action looks exactly like an anomaly worth trading, and has been mistaken for one.
Entries publishing.
Crypto9
Perpetual funding, fee determination, equilibrium dynamics and pair mechanics. A small shelf, because the tradeable literature is much thinner than the volume of writing on the asset class suggests. Shelf size here measures the literature, not our exposure.
Entries publishing.
Tabular deep learning8
Prior-fitted networks, transformers for tables, embeddings for numerical features. Read because a learned scorer out-ranked our hand-built gate, which opened the governance question rather than settling it. It advises. It does not set a floor.
Entries publishing.
Metrics8
Ranking and discrimination measures, information coefficients, multi-class AUC. Read because which in-sample statistic predicts out-of-sample survival is a question our own corpus answers uncomfortably. That the answer is counterintuitive is published; the ranking itself is a calibration of our own screen and is withheld.
Entries publishing.
Features7
Automated generation, pruning, causal selection for multivariate series. A small shelf because most feature machinery we tried added leakage faster than it added signal, and the causal-window constraint rules out a good deal of it before we start.
Entries publishing.
Geared funds6
Daily rebalance mechanics, compounding drag beyond volatility decay, long-horizon behaviour. A mechanism class with a named obligated party and a dated, contractual trigger, which is the shape we look for before a pattern is worth authoring.
Entries publishing.
Pipelines4
Automated machine-learning pipelines and their operations. Read for the engineering, kept for the failure modes — the dominant one in an estate this size is a component that reports success while doing nothing, which is why ours are built to fail loudly.
Entries publishing.
Entropy4
Approximate and sample entropy, complexity measures for physiological and financial series. Read for complexity features. The shelf is small because the payoff was small, and we stopped rather than kept reading.
Entries publishing.
Machine learning, general3
Scale and stability challenges, double descent. Framing for everything else on these shelves; nothing here is a method we run.
Entries publishing.
Imbalance3
Class imbalance, weighted losses, synthetic minority oversampling. Read because tradeable events are rare by construction, and a classifier that never fires scores extremely well on the metric nobody should be using.
Entries publishing.
Labeling2
Triple-barrier and event-based labelling. Two papers, because the no-leakage constraint decides most of this for us before the literature gets a say: any label whose definition reads data after the timestamp is out, whatever it is called.
Entries publishing.
Cross-domain2
Cointegration and monitoring methods borrowed from structural and mechanical engineering. Kept as a standing reminder that the technique is rarely the novel part, and that a method's home field usually validated it more carefully than we will.
Entries publishing.
On 2026-07-16 we ingested an external strategy corpus and scanned it with one question: does it hold a mechanism we are not already running? It does not. The scan is published because of what it found about our own ingestion, not because of what it found in the literature.
One ingested collection of 244 documents broke down roughly as follows. About a quarter was pure non-finance noise — immunology, DNA nanotechnology, weather models, EEG, high-entropy alloys, telescopes, RNA sequencing — pulled in by title resolution matching against reference lists. About a third was asset-pricing and market-structure theory rather than tradeable mechanism. Most of the remainder named mechanism families we already run. A separate compendium of 151 published strategies was, against the mandate as it stood that day, roughly four-fifths out of scope — the option-spread, fixed-income, credit, convertible, structured-product, real-estate and tax-arbitrage families — and every in-scope entry in it named a family already in our sweep corpus.
One candidate family came out genuinely unexplored. It is not named here. Naming it would say where we are about to look, and that is the one thing a reading list is not obliged to disclose.
A coincidence worth flagging so it is not read as an error: that external collection also contained 244 documents. It is a different collection, ingested separately for a different purpose, and it is not the corpus charted in Fig. 1.
The process failure and the fix are both published. The resolver matched titles against reference lists with no finance-relevance filter, so it imported whatever the citation graph handed it; a relevance filter was added at ingestion, and the instruction not to re-scan the collection expecting gold was written down with the reason. The same class of defect explains the metadata gap on this page: a large share of the store's author fields are unresolved, which is why many entries above print a title and a year and no author. Fixing that is a manual pass, and it is in progress rather than done.
Papers read for a currently live edge are counted in the shelf totals in Fig. 1 but are not listed as entries. Naming which shelf a live edge reads from narrows the search for anyone looking, and the shelf total is the honest place to carry them — present in the count, absent from the list, and declared as such rather than left to be noticed.
Two further omissions, stated for the same reason. Verdict assignment is a manual pass, complete on two shelves and continuing on the rest; a shelf publishes counts until its entries are real, because forty-one rows of guessed metadata would be worse than an honest empty section. And public identifiers resolve one entry at a time, so an entry prints what the store holds and no more.
Each of those omissions is declared where it occurs rather than left to be noticed.
If we have misread a paper, tell us. Verdicts are the most correctable claim on this site — they are our judgement of our own record, and both halves of that can be wrong. Reports are acted on and logged in the library's corrections register like any other error.
Items here are never silently edited. A wrong item is corrected by a new item that cites it, so the record shows the correction as well as the correct answer.