Agentic AI Engineering · Case Study

The trading platform that publishes its own negative result

Four research agents, live data across four markets, a verified backtesting engine. Then a walk-forward test: the strategies cut drawdown roughly in half, and still lost to buy-and-hold on risk-adjusted return. Both halves of that result are printed on the platform's front page.

The finding

Walk-forward validation, run once, on data the strategies had never seen.

What the test showed

Five portfolio variants were ranked on a tuning window. The winner showed a +0.27 Sharpe edge over buy-and-hold. Run once on a later window it had never touched, that edge became −0.10. Replicated independently on crypto: +0.09 → −0.21, with zero of five variants holding any edge at all.

UniverseTuned onTested onIn-sample edgeOut-of-sampleDrawdown saved
10 US equities2006–20152016–2026+0.27−0.10+38.3pp
7 crypto pairs2018–20222023–2026+0.09−0.21+16.7pp

Full write-up of the walk-forward test, with every variant and the method →

The correlation between in-sample and out-of-sample edge across all variants was +0.19 — essentially noise. In-sample ranking carried almost no information about which strategy would perform next. That is what overfitting looks like when you measure it instead of assuming it away.

What did generalise

Drawdown reduction, without exception. Every variant roughly halved it in both asset classes — 4.6–23.7% against buy-and-hold's 50.0% on equities, 16.0–36.4% against 53.1% on crypto. The strategies are a risk-management tool. They are not an alpha source, and any single-window backtest suggesting otherwise is measuring its own tuning data.

Why this is the interesting part

There are thousands of AI trading projects. Nearly all publish a win rate that was never measured, or one measured on the same data used to build the strategy. Mine started that way too — the first backtest returned +220% on AAPL and looked excellent.

I didn't believe it, because it came from the same window I'd spent hours tuning against. So I defined the walk-forward protocol first, ran it once, and published what came back. The platform's strategies page now leads with that result rather than the flattering one.

Engineering judgement is mostly about which numbers you're willing to trust. That's the capability this project is evidence of — not the trading.

What's actually built

Next.js 16 · Supabase · Groq · deployed and live

TypeScript18,030lines
Routes22app screens
Edge Functions4Supabase Deno
Backtests75verified runs

Multi-agent research

Four agents — news, technical, smart-money, risk — run on a Supabase Edge Function and synthesise a market view with an LLM fallback chain.

AI signal pipeline

Live OHLCV → RSI/MACD/EMA/ATR computed in pure TypeScript → Finnhub news → LLM → structured JSON signal persisted to Postgres.

Backtesting engine

Client-side, CORS-safe, runs in the browser. Verified against an independent implementation to the cent before any result was trusted.

Portfolio backtester

Equal-weight basket, each sleeve trend-gated, per-sleeve P&L attribution, benchmarked against equal-weight buy-and-hold on the identical window.

Live market data

US, India, UAE and crypto via an Edge Function proxy — Finnhub, Yahoo and Binance — sidestepping CORS and centralising rate limits.

Commodities layer

Real dollars-per-barrel WTI and Brent series back to 1986, gold spot with technicals, 19 live oil and gold instruments.

Engineering decisions worth naming

Same skill, different market

The forecasting stack under AlphaOS is the one I use in demand planning: build a signal from noisy live data, quantify the uncertainty honestly, and separate what you measured from what you assumed. A trading strategy that looks profitable on its tuning window is the same failure as a forecast that looks accurate on the history it was fitted to.

The reason I built it in markets is that markets are unforgiving. A demand forecast can be wrong for a quarter before anyone notices. A trading backtest lies immediately and convincingly, and only a protocol you commit to in advance catches it.

See the same approach applied to supply chain →