Methodology Results Explore Research Updates

Systematic Futures Research

A systematic approach
to ES futures.

Meridian Research studies the CME E-mini S&P 500 futures market. XGBoost models forecast the next hour’s return every 15 minutes, and a defined policy turns those forecasts into long, short or flat exposure. The question the research asks is simple: do those decisions add value beyond holding the market?

ES Futures Market
15m Decision Interval
1h Forecast Horizon
8.6y Simulated History
Scroll to explore
01 What We Test

Skill beyond the market, or just the market?

Holding ES outright has paid well since 2018, so profit on its own proves nothing. Every result on this page is set against a passive long benchmark over the same dates, with the same costs and the same account rules. What the research measures is the part of the return that comes from timing: being long, short or flat at the right moments rather than simply being in the market.

From Data to Trading Decisions

Three stages turn market data into a forecast, and the forecast into a position.

Phase 1

Market Features

Inputs are built from CME ES price, volume and volatility at minute, 15-minute, hourly and daily resolution, plus session timing and the 2-year, 10-year and bond Treasury futures as context. Every input uses only information available at the moment of the decision.

Output Price, volume, session & rates context
Purpose Describe the market as it stood at decision time
Phase 2

Return Forecasts

XGBoost models forecast the return over the next hour. They are retrained every six months on all prior history, and three models are trained each time and averaged to steady the forecast.

Output One forward-return score
Purpose Estimate the next hour’s return
Phase 3

Trading Policy

The policy acts on rank, not on the raw number. An unusually high forecast relative to recent history opens a long, an unusually low one reverses it into a short, and a recovering forecast reverses back. Time limits and a daily flat window close positions whatever the forecast says.

Output Long / Short / Flat
Purpose Defined entry, reversal and exit rules

Defined exposure rules. Explicit trading costs.

The strategy trades one contract at a time and never adds to a position. Hold limits and a daily flat window cap how long any position can run, regardless of the forecast. Every simulated fill pays the bid-ask spread and commission. Account sizing is kept separate from the signal research: the signal is judged on a single contract first, then replayed as an account in whole contracts.

01 One contract at a time, no pyramiding
02 Hold limits of 16 hours long and 24 hours short
03 Flat every day from 14:00 to 17:00 CT
04 Spread and commission paid on every fill

Strategy versus passive long, 2018 to 2026

Simulated MES account from 2 January 2018 to 14 August 2026, starting at $100,000 with a 1× notional exposure target, after spread and commission. The strategy was configured while looking at this period, so read these figures as an upper bound on what the approach has done, not as an expectation of what it will do.

Account summary

Metric Strategy Passive long
CAGR 24.9% 11.2%
Sharpe, annualised 1.50 0.65
Max drawdown 20.4% 34.3%
Final equity $679k $250k
Positive years 7 of 9 7 of 9

Strategy figures are from the best of three training runs that differ only in the random seed used to fit the models. 2026 is a partial year. Passive long holds the same notional exposure throughout and pays the same costs.

Account equity, log scale
Strategy Passive long

Weekly account equity in USD. Hover or use the arrow keys to read values; the annual table below carries the same data.

Beta to ES 0.40
Timing alpha, t-stat 6.4
Time in market 38%
Profitable months 78 of 104

Beta measures how much of the strategy’s return moves with the market; a value near one would mean the strategy is mostly just long. The timing t-statistic tests whether the return that remains after removing market exposure is distinguishable from zero. Both are measured on a single-contract book over the full window.

Annual returns

Year Strategy Passive long
2018−1.3%−6.4%
201910.3%27.0%
202098.6%15.1%
202138.9%26.6%
2022−3.1%−19.4%
202337.8%19.2%
202418.5%18.0%
202522.6%13.0%
202616.5%12.4%

Strategy is the best of three runs, seed 8. 2026 covers 1 January to 14 August. The policy needs 8,000 prior forecasts before it trades, so the first trade is in May 2018. Returns before 2020 were modest and the account compounds, so the later years carry most of the dollar gain.

By direction

Long trades make up about 70% of all trades and about 85% of the profit. Short trades are fewer, shorter and less profitable per trade, but they made money in seven of nine years, including 2022, when they offset most of the losses on long trades.

After costs

Every fill pays the spread and commission. Those costs consumed about 20% of gross account profit. The figures above are net of them.

Where the edge shows up

The timing component is strongest when volatility is high and weakest when it is low, and in 2022 the strategy lost 3% while the market fell 19%. Both are what a timing edge, rather than a leveraged market bet, should look like.

Two Ways to Explore the Research

Start with the plain-language overview, or go straight to the specification behind the results.

One market. Three possible positions.

The system studies whether forecasts can help time exposure to the S&P 500 futures market. Long means it profits if prices rise, short means it profits if prices fall, and flat means no position at all. It goes long when the forecast is unusually strong, flips direction when the forecast turns, and closes when a time limit or the daily flat window arrives.

A Regular Decision Cycle A fresh decision every 15 minutes, using only information available at that moment
Rank-Based Decisions It acts only when a forecast is unusually high or low compared with recent forecasts, not on the raw number
Time Limits on Every Position No position runs longer than 16 hours long or 24 hours short, and everything is flat from 14:00 to 17:00 CT
What It Looks Like in Practice In the market a little over a third of the time, roughly 480 trades a year, and a typical hold of around four and a half hours

What the Research Covers

  • ES market data with Treasury futures as context
  • Forward-return forecasting
  • Long, short and flat decisions
  • Session and holding limits
  • Account replay in whole contracts
  • Spread, commission and roll costs

Specification.

The current configuration: a selected feature subset, a three-member XGBoost ensemble and a rank-based reversal policy, evaluated walk-forward with purge and embargo. These are the parameters that produce the results above.

Data & Target

Data CME ES, ZT, ZN and ZB minute bars via Databento, aggregated to 15-minute, hourly and daily bars
Price Basis Roll-adjusted continuous series for features and labels; raw active-contract prices for fills
Target Forward log return over four 15-minute steps, one hour
Features 38 inputs selected from a larger registry: ATR-scaled returns from 5 minutes to 20 days, price location against the session, prior-session levels and VWAP, realised-volatility and volume ratios, RSI and EMA spreads, session clock, ZN returns, ES–ZN correlation and divergence, and Treasury curve twist and butterfly from ZT, ZN and ZB

Model & Validation

XGBoost Three regressors averaged; pseudo-Huber objective; depth 8; learning rate 0.008; row and column subsampling
Early Stopping Boosting rounds chosen on the six months before each forecast window
Walk-Forward Expanding window from November 2015; 17 six-month folds from 2018 to 2026; one-day embargo; purge where labels cross a boundary
Skill Mean out-of-sample rank correlation 0.058, positive in 17 of 17 folds; mean directional hit rate 52.8%

Policy

Long Entry From flat, when the score exceeds the 80th percentile of the prior 24,000 scores, about one year
Reversals A long falling below the 10th percentile closes straight into a short; a short rising above the 70th percentile of its own 72,000-score window, about three years, closes straight into a long
Limits Maximum hold 16 hours long and 24 hours short; flat 14:00 to 17:00 CT; invalid inputs and limits close to flat without reversing

Execution & Account

Fills Next raw minute open of the active contract; 0.125 points spread cost plus $2.50 commission per side on ES; a roll is booked as two fills
Account $100,000 in whole MES contracts sized to 1× notional at entry and held fixed for the trade; $0.75 commission per side
Costs Spread and commission consumed about 20% of gross account profit

Temporal Integrity

Features Available at decision
Split boundaries Purge & embargo
Percentiles Prior scores only
Retraining Prior history only

Attribution, Unit Book

Beta to ES 0.40
Alpha t-stat 6.2
Timing t, low vol 2.4
Timing t, mid vol 2.9
Timing t, high vol 5.5
Exposure fraction 0.38

Common Questions

Isn’t this just exposure to a rising market?

That is the first thing the research checks. Beta to ES is about 0.40, the strategy is in the market only about 38% of the time, and in 2022 it lost 3% while the market fell 19%. The timing component that remains after removing market exposure has a t-statistic above 6. Profit in a rising market alone would not establish an edge; these are the tests that do.

What markets do you cover?

ES, the CME E-mini S&P 500 future, is the only traded market. ZN, the 10-year Treasury note future, and the 2-year (ZT) and bond (ZB) futures are inputs and are never traded. The account replay sizes the same decisions in Micro E-mini (MES) contracts and can use full ES contracts instead.

What is walk-forward validation?

Models are trained on history up to a cut-off and then forecast the following six months, which they have never seen. The cut-off rolls forward every six months, so every forecast in the results comes from a model with no access to that period. The configuration itself was still chosen using these results, which is why they are an upper bound rather than a forecast.

How does account sizing work?

The replay starts with $100,000 and sizes each trade to about 1× notional in whole MES contracts, fixed for the life of the trade. This is a research convention, not a recommendation for a minimum account size or a loss limit. Smaller accounts are constrained by whole-contract sizing.

Follow the Research

Telegram, Discord, Collective2 and hosted automation are planned; destinations and launch details are to be confirmed. The methodology and results above describe the current state of the research.

Explore the Methodology