Evaluation Engine · research preview

Test whether your edge survives before you risk real capital.

Almost any strategy looks profitable in a backtest. ApexQuant doesn't tell you if a strategy makes money — it stress-tests assumptions to expose how fragile an edge really is. We stress-test. We do not optimize.

→ Robustness, not profitability→ Out-of-sample first→ No fake verdicts
What the engine stress-tests

Six ways we try to break your strategy

When the engine runs your submitted hypothesis, it applies the same battery we use on our own research. These analyses require real historical data — they are computed by the backtest engine, never guessed.

In-sample / out-of-sample split

Calibrate on history, evaluate frozen on unseen data. The test that breaks most ideas.

Random universe test

Swap chosen assets for random ones. If the edge only lived in cherry-picked symbols, it surfaces here.

Concentration risk

Does the result lean on one or two assets? We measure cross-asset PnL correlation and contribution.

Regime stability

Split out-of-sample into bull, bear and flat. A real edge survives more than one regime.

Cost sensitivity

Re-run with rising realistic costs. Many strategies die the moment fees enter.

Parameter robustness

Perturb parameters slightly. A robust edge is stable; an overfit one collapses.

The report you'll receive

What a Robustness Profile looks like

Below is the format of the institutional report the engine produces. The figures are illustrative — your real evaluation is computed on historical data when you submit a hypothesis.

Sample report — illustrative format only, not your evaluation

Robustness Profile

Example: Momentum L/S · Top-30 · 1y window

Fragile
Fig. 1 — Cumulative edge, IS vs OOSUNI/USDT · representative
highzeroentry threshold (1 bp)IN-SAMPLE · 2022–2024OUT-OF-SAMPLE · 2025–2026edge peaks, then collapsestime →
Figure 1. Cumulative edge accumulates steadily through the in-sample period, then flattens and reverses out-of-sample. The same parameters, applied to unseen data, produce no exploitable signal. Curve is illustrative; underlying metrics below are exact.
IS Sharpe6.94
OOS Sharpe-0.71
VerdictInvalidated
TestResultReading
IS → OOS degradation-48%Edge halves
Spearman ρ (IS→OOS)0.41Unstable
Concentration0.622 assets dominate
Cost-to-death0.28%Dies under fees
Regime consistency1 / 3Bull only
Illustrative figures. Real evaluations are computed by the backtest engine on historical data.

This is a mock-up of the report format you will receive when the engine runs your submitted hypothesis. The numbers above are invented for illustration and describe no real strategy.

What the MVP does — and doesn't — do

Deliberately limited. On purpose.

A focused engine that does one thing reproducibly is worth more than one that pretends to validate everything. The first version evaluates a single, well-understood strategy family.

Supported
  • · Cross-sectional ranking strategies
  • · Momentum, reversion, carry, volatility, RSI signals
  • · Fixed crypto universes and hand-picked assets
  • · Structured composition — no code required
Not yet (by design)
  • · Single-asset market timing
  • · Arbitrary user code or custom datasets
  • · Automatic parameter optimization
  • · Equities, FX, options

Automatic optimization is exactly what we refuse to build. An engine that searches for the best-looking parameters is a machine for manufacturing overfit. We stress the configuration you bring — we never tune it for you.