Multi-agent research
Four agents — news, technical, smart-money, risk — run on a Supabase Edge Function and synthesise a market view with an LLM fallback chain.
Four research agents, live data across four markets, a verified backtesting engine. Then a walk-forward test: the strategies cut drawdown roughly in half, and still lost to buy-and-hold on risk-adjusted return. Both halves of that result are printed on the platform's front page.
Walk-forward validation, run once, on data the strategies had never seen.
Five portfolio variants were ranked on a tuning window. The winner showed a +0.27 Sharpe edge over buy-and-hold. Run once on a later window it had never touched, that edge became −0.10. Replicated independently on crypto: +0.09 → −0.21, with zero of five variants holding any edge at all.
| Universe | Tuned on | Tested on | In-sample edge | Out-of-sample | Drawdown saved |
|---|---|---|---|---|---|
| 10 US equities | 2006–2015 | 2016–2026 | +0.27 | −0.10 | +38.3pp |
| 7 crypto pairs | 2018–2022 | 2023–2026 | +0.09 | −0.21 | +16.7pp |
Full write-up of the walk-forward test, with every variant and the method →
The correlation between in-sample and out-of-sample edge across all variants was +0.19 — essentially noise. In-sample ranking carried almost no information about which strategy would perform next. That is what overfitting looks like when you measure it instead of assuming it away.
Drawdown reduction, without exception. Every variant roughly halved it in both asset classes — 4.6–23.7% against buy-and-hold's 50.0% on equities, 16.0–36.4% against 53.1% on crypto. The strategies are a risk-management tool. They are not an alpha source, and any single-window backtest suggesting otherwise is measuring its own tuning data.
There are thousands of AI trading projects. Nearly all publish a win rate that was never measured, or one measured on the same data used to build the strategy. Mine started that way too — the first backtest returned +220% on AAPL and looked excellent.
I didn't believe it, because it came from the same window I'd spent hours tuning against. So I defined the walk-forward protocol first, ran it once, and published what came back. The platform's strategies page now leads with that result rather than the flattering one.
Engineering judgement is mostly about which numbers you're willing to trust. That's the capability this project is evidence of — not the trading.
Next.js 16 · Supabase · Groq · deployed and live
Four agents — news, technical, smart-money, risk — run on a Supabase Edge Function and synthesise a market view with an LLM fallback chain.
Live OHLCV → RSI/MACD/EMA/ATR computed in pure TypeScript → Finnhub news → LLM → structured JSON signal persisted to Postgres.
Client-side, CORS-safe, runs in the browser. Verified against an independent implementation to the cent before any result was trusted.
Equal-weight basket, each sleeve trend-gated, per-sleeve P&L attribution, benchmarked against equal-weight buy-and-hold on the identical window.
US, India, UAE and crypto via an Edge Function proxy — Finnhub, Yahoo and Binance — sidestepping CORS and centralising rate limits.
Real dollars-per-barrel WTI and Brent series back to 1986, gold spot with technicals, 19 live oil and gold instruments.
The forecasting stack under AlphaOS is the one I use in demand planning: build a signal from noisy live data, quantify the uncertainty honestly, and separate what you measured from what you assumed. A trading strategy that looks profitable on its tuning window is the same failure as a forecast that looks accurate on the history it was fitted to.
The reason I built it in markets is that markets are unforgiving. A demand forecast can be wrong for a quarter before anyone notices. A trading backtest lies immediately and convincingly, and only a protocol you commit to in advance catches it.