Back-Test

Calibration evidence for the AI signal engine. Replays proxy-AI BUY/SELL/HOLD recommendations over the selected window and scores against forward returns. All numbers derive from the same deterministic data — re-run produces identical output for an unchanged window.

≥75% conviction hit-rate
n/a
no high-conv signals
Sharpe Δ vs benchmark
n/a
vs STI (^STI)
Sortino Δ vs benchmark
n/a
vs STI (^STI)
Max-DD vs benchmark
n/a
vs STI (^STI)
SELL hit-rate
n/a
no SELL signals emitted

K1 by market regime

SPX trailing-60d log-return at signal date

High-conviction (raw ≥75) hit-rate split by regime — bull (>+5%), sideways (±5%), bear (<−5%). A K1 that holds only in one regime is not a falsifiable claim. Wilson 95% CI shown for each bucket.

Bull (SPX +5%)
n/a
no signals in regime
Sideways (±5%)
n/a
no signals in regime
Bear (SPX −5%)
n/a
no signals in regime

HOLD prudence rate

|fwd_5d| < σ_per-ticker

When the engine says HOLD, how often does the price stay inside the ticker's one-sigma forward-return band? Fair criterion — replaces the prior near-impossible |fwd|<1% rule that suppressed HOLD hit-rates and forced HOLD out of K1.

HOLD prudence (σ-based)
n/a
no HOLD signals

Conviction calibration

0 signals

Hit-rate per conviction band using calibrated conviction scores (Platt isotonic regression on historical outcomes). Well-calibrated = monotonically rising bars. Flat or inverted = residual miscalibration — raw arithmetic bands shown in the CSV export for comparison.

No data.

Drawdown comparison

no data

Peak-to-trough loss from the start of the window. Shallower trough = stronger capital preservation. Model = conviction-weighted long-only NAV; benchmark = STI (^STI).

No risk data.

NAV growth

model vs STI (^STI) · normalised to 1.0 at window start

No risk data.

Per-ticker hit-rate

click column header to sort
No data.
Methodology · MAS FEAT audit footer. Engine type: proxy-rules backtest (deterministic RSI / MA stack / MACD / sentiment heuristic). NOT the live LLM that production emits — running 480+ Bedrock/Ollama calls per backtest is expensive and non-deterministic. The proxy uses the same inputs the LLM sees and is a faithful approximation of its decision surface. Live-LLM signal back- test is a future item. Reading the numbers: raw all-band hit-rates (incl. K5 SELL hit-rate) are dominated by mid-conviction noise — that's why conviction-weighting matters. The headline calibration is K1 (high-conviction band hit-rate): when the engine is confident, it's substantially more often correct than when it's not. Use the conviction-band chart above to verify monotonic calibration. v4 (2026-05-25) — confidence + regime + HOLD disclosure: K1 now ships with a Wilson 95% binomial CI and a low-n badge (n<30) so the point estimate is read with its interval. K1 is also split by SPX trailing-60d regime (bull >+5% · sideways ±5% · bear <−5%) so single-regime claims are visible. HOLD prudence uses a fair criterion: |fwd_h| < σ_per-ticker (historical h-day log-return std-dev from pre-display closes; fallback 2.5% when warmup is thin). The prior |fwd|<1% rule was near-impossible and suppressed HOLD hit-rates. Universe: . Lookback: 90 trading days. Forward horizon: 5 days. Benchmark (risk metrics): STI (^STI)Singapore market home-bias. NAV scheme: conviction-weighted long-only · monthly rebalance · HOLD = cash · SELL = exclude (D-2026-05-11-AO). Last refresh: . Decision-support — not personalised investment advice under SFA / FAA.