The methodology, stated plainly
Every design decision below is downstream of one objective: maximising the probability that observed historical performance generalises to unseen market conditions.
Markets are adaptive, noisy and non-stationary. That has three consequences most systematic approaches under-price. Relationships decay: a feature that predicted returns in one monetary regime may be uninformative or inverted in another. Backtests are easy to win: with enough features, horizons and parameterisations, an apparently significant strategy can always be found. Edge is fragile to implementation: a genuine statistical edge can be entirely consumed by spread, impact and turnover.
Look-ahead contamination is structurally prevented.
What data architecture lets us ask historical questions without contaminating the experiment with information that was unavailable at the time?
Most quantitative work is conducted on datasets that quietly encode the answer. Standard vendor universes contain only companies that survived. Fundamental data is commonly aligned to the reporting period rather than the filing date, giving a model information weeks before the market received it. Macroeconomic series are supplied in revised form rather than as first published.
The research environment ingests approximately 400 source datasets spanning more than 45 years: market data, SEC EDGAR filings for every company in the universe, and official macroeconomic statistics from the Federal Reserve (FRED). The filings are annual and quarterly reports (10-K, 10-Q and their foreign and older equivalents 20-F, 40-F, 10-K405, 10-KSB, 10-QSB), current reports (8-K), insider transactions (Form 4) and large-holder disclosures (Schedules 13D and 13G): about 21 form types, each read as of its filing date. Filings are extracted period-matched and point-in-time, so a model sees a disclosure when the market saw it, not when the quarter ended.
The historical universe tracks 1,767 US-listed companies from 1980, retaining bankruptcies, mergers, acquisitions and delistings: 422 of them no longer trade. Companies that failed remain in the record with the returns they actually delivered.
Of that panel, 1,345 companies are listed today. A Book selects a name only if it clears the Book's liquidity floor — by default a $10M median daily dollar volume — and never a name without a measured price or volume, so every position can be traded at swing-horizon sizes.
| Definition | Companies | What it counts |
|---|---|---|
| Universe | 1,767 | The US companies the platform tracks from 1980, including those that failed, merged or delisted. |
| Listed today | 1,345 | Companies in the universe that still trade. |
| No longer trade | 422 | Delisted, acquired, merged or failed, and kept with the returns they actually delivered. |
| Fair value today | 1,301 | Companies with a fair value on the Oct 9, 2026 close, in the latest daily fair-value run. |
| Study sample | 1,763 | Companies the studies measured every Wednesday from 1995; a median of 1,148 per Wednesday. |
Observations are partitioned into distinct market eras, and models train only on information available at that point in history — look-ahead contamination is structurally prevented rather than avoided by discipline.
A question asked of this dataset — would this decision have worked? — returns an answer that is not quietly pre-informed by the future.
Twenty technical indicators computed from one price history are not twenty pieces of evidence.
What representation is rich enough to describe a security's condition across independent domains, without collapsing into redundant restatements of price?
Every security, every trading day, is represented as a 1,715-dimensional state vector constructed across orthogonal domains: price kinematics, order-flow and microstructure behaviour; multi-horizon trend persistence; corporate and cross-asset sentiment parsed with domain-specific NLP — FinBERT and the Loughran-McDonald financial dictionaries — rather than general-purpose scoring, which systematically misreads financial language; multi-scale Daubechies-4 wavelet decomposition and geometric pattern structure; period-matched fundamentals from filings; explicit Fama-French factor-exposure decomposition; macro-financial regime conditioning; and intrinsic-value dislocation measures.
Volatility structure is modelled with higher-order estimators — realised semivariance, bipower jump share, realised skew and kurtosis — rather than a single dispersion measure. Unsupervised anomaly detection (Isolation Forest) flags statistically unusual states that no supervised label anticipates. The macro layer contributes 347 engineered indicators built from 78 FRED series — levels, 1-, 3- and 6-month changes, 2-year z-scores and accelerations, plus category composites.
Breadth across independent domains means the system is not repeatedly confirming the same observation. A conviction supported by valuation dislocation, insider accumulation, factor positioning and regime alignment is a different object from one supported by four momentum indicators.
A forecast's reliability is itself a measured quantity.
Can statistically significant forward-looking signal be extracted from a high-dimensional, non-stationary state space — and can the uncertainty around it be quantified with finite-sample guarantees?
Point forecasts are the standard output of the category, and they are close to useless for capital allocation. “Expected return of 4.2%” contains no statement about dispersion, and therefore no basis on which to size a position.
Forecasts are produced at 1-day, 5-day, 21-day and 63-day horizons, matched to the swing structure of the platform.
The predictive layer uses deliberately diverse model families rather than an ensemble of near-identical gradient-boosting variants. Diversity in inductive bias is the point of an ensemble; correlated learners provide none.
Model outputs are consolidated via Bernstein Online Aggregation — an online-learning aggregator with formal second-order regret guarantees — which re-weights each model as its realised performance accumulates, rather than fixing weights at training time. This is the mechanism behind the homepage's claim: the system re-weights what it believes, every trading day.
Prediction intervals are produced through Conformal Quantile Regression — quantile bundles combined with split conformal prediction under Adaptive Conformal Inference — yielding distribution-free intervals with finite-sample coverage guarantees that hold without assuming the return distribution takes any particular form. Every model is continuously monitored on conformal interval coverage, calibration slope and CRPS (the continuous ranked probability score), so a forecast's reliability is itself a measured quantity.
Coverage guarantees are what allow uncertainty to be an input to position sizing rather than a caveat in a footnote.
Regime is a routing and interaction term — not a flat additive predictor.
Can changes in the macro-financial environment be detected early enough, and reliably enough, to alter model behaviour and portfolio risk before the change is priced?
Most systematic strategies are regime-blind. They apply one set of learned relationships across credit expansions and credit crises alike, and discover the regime change through drawdown.
The regime engine models independent axes of market behaviour — credit conditions, monetary policy, market liquidity, inflation, real interest rates, recession risk, market sentiment, volatility structure and the fundamental cycle — and classifies each one from where its indicators sit in their own history, against percentile thresholds set for that axis. A hidden Markov model and a gradient-boosted classifier are trained for each axis as cross-checks; they do not set the published states. A Hidden Semi-Markov Model then fuses the nine readings into one market state. The semi-Markov formulation models state duration explicitly, which a standard HMM cannot — and regime persistence is economically meaningful.
The resulting state vector is the primary router for the platform: it conditions quantile optimisation, scales risk premia inside the daily valuation calculation, and adjusts ensemble weighting to the prevailing environment. Regime is a routing and interaction term — not a flat additive predictor.
Today's reading, published daily: the Regime Radar →An unqualified value signal is a value-trap generator.
Can intrinsic-value estimation adapt to changing discount rates, macroeconomic conditions and market dynamics, rather than relying on assumptions fixed at the time of the analysis?
Conventional DCF fixes WACC, terminal growth and margins at a point in time and holds them for months — precisely the inputs that move most when the environment shifts. A subtler failure: fundamental valuation runs on a slower clock than the market. A security can be genuinely undervalued and continue falling for a year; for a swing-horizon strategy, an unqualified value signal is a value-trap generator.
Guided by the HSMM state probabilities, the valuation engine re-calibrates discount structures, margin parameters and growth risk premia when a macro or liquidity transition is detected — not at the analyst's next review.
Valuation structures are fused with asset kinematics — price velocity, acceleration, localised frequency volatility — separating undervalued and turning from undervalued and still falling.
The full valuation vector is recomputed every trading day against macroeconomic releases, TIPS-implied real yields, incoming filings and current sentiment.
The difference between a research programme and a backtest.
How do we distinguish a real edge from the best of many things we tried?
If a hundred strategy variants are tested, several will show a t-statistic above 2.0 by construction. A conventional significance threshold applied to a mined search space is not evidence. Candidate strategies must therefore clear a battery designed specifically for mined search spaces:
| Gate | Standard |
|---|---|
| Multiple-testing correction | The winning variant must survive a Romano-Wolf step-down at a 5% family-wise error rate across every variant the search compared, and its Sharpe ratio is deflated for the number of trials actually run. Bars: Deflated Sharpe ≥ 0.95, probability of backtest overfitting ≤ 0.30, Hansen SPA p ≤ 0.05 |
| Data-snooping tests | White's Reality Check, Hansen's SPA and Romano-Wolf stepdown across the full trial family |
| Sharpe deflation & track sufficiency | Probabilistic and Deflated Sharpe Ratio, minimum track-record and minimum backtest length — with the number of trials counted honestly, never assumed to be one |
| Overfitting probability | Combinatorially symmetric cross-validation (CSCV) estimate of the Probability of Backtest Overfitting |
| Resampling done right | Stationary block bootstrap at the Politis-White optimal block length; combinatorial purged cross-validation with embargo against leakage across folds |
| Search discipline | Hyperparameter search (Optuna TPE) is scored on a lower confidence bound — the mean walk-forward fold Sharpe minus one standard error — so a variant that holds up in every fold ranks above one that did well in a single fold |
Kill-switch criteria — a required out-of-sample floor — are declared before the result is seen; declaring the threshold in advance is what makes it a test rather than a rationalisation. Failing any gate returns the component to research.
The edge is evaluated net, not gross.
How should security-level forecasts, each with its own uncertainty, be transformed into a coherent portfolio under competing constraints — and does the edge survive the cost of acting on it?
The most common failure in systematic investing is treating a portfolio as a list of good ideas. Ten high-conviction positions sharing a factor exposure are one position. The second most common failure is transaction-cost blindness: a backtest on closing prices can report a strong edge for a strategy that loses money in production.
Construction is organised as template → policy → portfolio: a single strategy template, fully parameterised by policy, can power thousands of independently configured portfolios. The platform publishes three model Books, one per risk profile; every client Book is built from its own risk profile, mandate and preferences, and every parameter lives in the policy, versioned and reproducible.
Portfolios are built on the belief distributions published by the predictive layer — not point forecasts — and resolved against portfolio state: sizing by conviction, equal weight or the Kelly criterion as each Book's policy sets it, with each position's uncertainty drawn from its own forecast interval; horizon conflict resolution; binding bounds on gross and net exposure; tiered drawdown circuit breakers; a regime overlay that can only reduce risk, never add it; and risk-profile presets, so the same research output produces different portfolios for different mandates.
Trading costs are modelled per stock: a bid-ask spread estimated from daily prices (never below one cent, capped at 60 basis points), a 0.5-basis-point commission, and an Almgren-Chriss market-impact charge for orders above 0.5% of a stock's daily volume, at standard default parameters. There are no live fills yet, so none of these is fitted to real executions. Where estimated cost exceeds the predicted statistical edge, the instruction is suppressed. The edge is evaluated net, not gross. The $10M median daily dollar-volume floor exists specifically so that impact costs remain modelable at realistic sizes.
A portfolio decision, not a ranked list — with concentration, correlation, mandate constraints and implementation cost resolved before the recommendation arrives.
A control plane, not a report.
Should this portfolio be allowed to take this change — and at what size?
Risk is not a stage at the end of the pipeline. It is a synchronous control plane: every allocation passes through it before any instruction exists, and it is the only layer whose objective function is not return.
VaR is not subadditive — a VaR budget can be exhausted by a diversifying trade. CVaR is subadditive and positively homogeneous, which yields an exact Euler decomposition of the portfolio budget into per-position contributions: “this order spent 12 basis points of risk budget” is an arithmetic fact, not an attribution heuristic. Portfolio tail risk is estimated by filtered historical simulation — chosen for identifiability at the platform's actual sample sizes, because a measured statistic with a limit can be backtested and an unidentifiable dynamic model cannot.
Covariance comes from a factor model estimated across the whole universe, not from the few names in one book, because a book of fifteen names on a year of data cannot identify a full correlation matrix. Tail risk is measured by filtered historical simulation, which keeps the real shape of past losses and rescales them to today's volatility, and fat tails are fitted only where there is enough data to fit them. The risk budget is set in CVaR rather than VaR, because CVaR adds up exactly across positions. Concentration is measured as correlation-adjusted effective N, and liquidity horizons are re-estimated under stress. Reverse stress testing is closed-form: the smallest shock that breaches a limit is computed, not guessed.
Limits live in a versioned, dated, approved limit book with four-eyes change control; regime overlays are tighten-only by construction — no override can ever widen a limit; kill-switch governance, restricted-list pre-trade compliance and an append-only audit trail are enforced in software. Risk evaluation is fully deterministic — no randomness, no wall-clock caches — so every verdict is exactly replayable.
Constraints that bind at construction produce a different portfolio from risk reports that describe one.
No single intelligence grades its own homework.
The platform decomposes the investment problem into five specialised intelligence services, fed by market data and prediction services upstream, so each function can be independently validated before recombination into a decision system:
| Intelligence | The question it owns |
|---|---|
| Investment | What do we believe? Publishes belief distributions, with uncertainty, per security and horizon. |
| Strategy | Which investment policy is optimal? Searches the policy space and validates strategy templates. |
| Portfolio | How do we express and implement? Owns construction, position lifecycle, rebalancing and implementation cost. |
| Risk | Should we allow this? A synchronous control plane operating pre-trade, intra-horizon and post-trade. |
| Experiment | How does the system improve? Adjudicates evidence and emits recommendations — and never acts. |
A service may publish a quantity only if it also validates it.
The layer that evaluates whether an edge is real holds no position in the answer — it will never promote a model, change a weight, or size a position. Two doctrines follow: gating controls consumption, never evaluation — blocked predictions are still scored, which is precisely what makes every gate falsifiable; and limit changes are themselves gated through pre-registered hypotheses, so even the risk framework's own evolution requires declared, testable evidence.
Because each stage is independently validated, a failure can be localised to a layer — a selection problem is distinguishable from an allocation problem from an implementation problem — rather than diagnosed as “the model is underperforming.”
The platform runs a pre-registered experiment on itself, every trading day.
Model skill is evaluated on a Model × Context lattice — roughly 4,000 evidence cells per day — across sectors, liquidity buckets and horizons, with empirical-Bayes shrinkage so that a thin cell is pooled toward the prior and cannot over-claim: an apparently spectacular result on eight observations shrinks to nearly nothing, and says so.
Rolling statistics flatter themselves: a rolling IC series can carry lag-1 autocorrelation near 0.97, so months of daily values contain only a handful of independent observations. The evidence layer therefore derives non-rolling, per-day metrics and reports effective sample sizes computed from them — a deliberately harder standard.
Every monitored question is a standing, pre-registered hypothesis: metric, direction and minimum effect are declared before any evidence is read, and pre-registration is enforced in software, not advisory. Verdicts are corrected for multiplicity (Benjamini-Hochberg across each day's hypothesis family), inference uses block-bootstrap methods that respect serial dependence, and robustness is tested with the Deflated Sharpe Ratio in both directions — because a low probability of skill is evidence of skill not established, not of harm demonstrated, and the two must never be conflated.
Page-Hinkley change detectors monitor every model-horizon pair for degradation, and distribution shift is vetoed with two-sample Kolmogorov-Smirnov tests — chosen over population-stability heuristics whose thresholds were calibrated on samples three orders of magnitude larger than a trading day provides.
The system that evaluates whether an edge is real never trades on the answer — and every claim it makes was specified before the data arrived.
The evolution is real — and it is disciplined.
The homepage says the platform evolves. This section states, precisely, what that means — because “self-learning” is the most abused claim in this category, and precision is the difference between an adaptive system and an unaccountable one.
| Layer | Update mechanism | Cadence |
|---|---|---|
| Feature matrix | Full recomputation across the production universe | Every trading day |
| Valuation estimates | Re-estimation against current macro, yields, filings, sentiment | Every trading day |
| Regime state | Re-inference of the market state vector | Every trading day |
| Belief consolidation | Bernstein Online Aggregation across model outputs | Continuous, per observation |
| Model weights & architecture | Scheduled retraining and promotion | Periodic; the newest model is promoted, its skill measured live |
| Strategy templates | Authored, validated, then promoted or rejected | Periodic, gated |
Beliefs, regime state and valuations adapt continuously as outcomes arrive. Strategies change only through the explicit statistical gates of the Promotion Gate above; models are retrained on a schedule, the newest is promoted, and its skill is measured live. The loop is continuous in evaluation; for strategies it is deliberately discontinuous in deployment.
One component of an evidence framework, not the argument itself.
Performance is presented here — at the end of the technical case, not the beginning — because it is one component of an evidence framework, not the argument itself.
A deliberately mechanical track that isolates raw predictive selection: the week's ten highest-ranked names by model belief, each held long or short by the sign of its forecast, rebalanced after Monday's close, with no portfolio optimisation. Because nothing downstream can flatter it, it is the cleanest public measure of whether the models select well.
| Track | What it isolates | Construction | Result |
|---|---|---|---|
| Golden Tickets | Raw predictive selection | Weekly top 10, long or short by forecast sign, no optimisation | Reading the published baskets… |
Transaction costs are deliberately not modelled here, and that is what makes the track readable. Golden Tickets exists to answer one question — do the models select well — so anything layered on top of the selection would blend prediction quality with an execution assumption, and a reader could no longer tell which of the two moved the figure. Implementation cost is priced where it changes a decision: inside portfolio construction, against the specific size and liquidity of a real book, where the edge is evaluated net.
The platform publishes three model Books, one per risk profile; every client Book is built from its own profile and inherits the same evidence chain: deterministic walk-forward validation of its policy, factor attribution against its benchmark, and exact CVaR budget attribution per order.
A months-long live window is sufficient to demonstrate that a system operates as designed under real frictions. It is not sufficient to establish a statistically significant edge, and we do not present it as if it were.
Any published performance figure ships with: the exact observation window (start and end dates); gross and net of all costs, fees, commissions and slippage; annualised volatility and maximum drawdown; Sharpe or Sortino ratio; average and maximum gross and net exposure; beta and factor attribution versus benchmark; turnover; and the benchmark stated as total or price return.
The honest formulation — and the only one we use — is: “observed during our live validation period, under the methodology and disclosures set out here.” Never “proven to outperform.”
Past performance does not guarantee future results. Nothing on this page is investment advice. All investing involves risk of loss.
What popular signals actually did next
Chart patterns, RSI and MACD, short interest, insider buying, analyst coverage, volume and liquidity — each measured on the same 1,767 US stocks every week since 1995 against the same day's median stock, with the method and every number stated, and 24 popular trading beliefs tested one by one. Most signals are context, not forecasts; a few carry a small, steady edge.
Read the studiesInvestment intelligence begins where prediction ends.
No credit card required · You keep final authority over every action.