# Backtest Desk > Take one finished trading-strategy backtest - the strategy's description or the code that ran > it, the equity curve, return series or trade log it produced, and what the author claims about > it - and decide in one sitting whether the result can be believed and how much risk it actually > carries. A free in-browser engine reads the series and computes every metric first; the audit > lane names the biases with evidence and discounts the reported Sharpe; the risk lane turns the > real drawdowns and tails into limits, sizing and a monitoring plan. Derived from > @wshobson/backtesting-frameworks and @wshobson/risk-metrics-calculation. https://backtest-desk.skillsafe.ai/ ## What it is Two lanes over one work object. The work object is a backtest: `strategy` (description and/or code), `series` (equity curve, returns or trade log), `claims` (what the author says), `context` (asset class, capital, costs, constraints) and three parameters (risk-free rate, cost per side in basis points, capital). The lanes are `audit` and `risk`, chosen by a `task` field, and one system prompt routes on it. The risk memo takes the audit's haircut as `prior_audit` - the handoff is a button on the audit result. ## The free lane - no model, no account - **A real table reader.** CSV, TSV, semicolon, pipe or whitespace tables; the date column and the value column are found from the header or, failing that, from the shape of the values; an equity curve, a return series and a trade log (entry, exit, pnl or pnl%) are told apart; percent versus decimal returns are detected; the sampling frequency is read from the median gap between dates. - **The metrics engine**, over every row: CAGR, annualised volatility, Sharpe and Sortino against the risk-free rate, maximum drawdown with its dates, duration and recovery, longest time under water, the top five drawdown episodes, Calmar, historical VaR and CVaR at 95 and 99 percent, skew, excess kurtosis, hit rate, best and worst period, average win and loss, profit factor, tail ratio, the probabilistic Sharpe ratio (Bailey and Lopez de Prado, against zero, skew- and kurtosis-adjusted), the minimum track record length, the rolling one-year Sharpe range, a monthly return table, and - for a trade log - trades per year, annual cost drag at the stated bps and the Sharpe net of it. - **The code lint**, deterministic, with stable ids the audit must reconcile: `LA-NEG-SHIFT`, `LA-BFILL`, `LA-CENTER`, `LA-FULL-SCALE`, `LA-NO-LAG`, `LA-PINE-LOOKAHEAD`, `LA-ILOC-NEXT`, `LA-COC`, `EX-SAME-BAR`, `SV-CURRENT-UNIVERSE`, `SV-DROPNA`, `TC-NO-COSTS`, `TC-ZERO`, `OF-INSAMPLE-OPT`, `OF-NO-OOS`, `OF-MANY-PARAMS`, `DQ-NO-SEED`, `DQ-ADJ-LEVEL`. - **Series flags**: `SR-NONE`, `SR-SHORT`, `SR-TOO-GOOD`, `SR-NO-LOSSES`, `SR-DD-TINY`, `SR-PSR-LOW`, `SR-TRACK-SHORT`, `SR-FAT-TAILS`, `SR-REGIME`, `SR-BAD-DATES`, `SR-UNSORTED`, `SR-DUPES`, `SR-NONPOSITIVE`. - **Claims reconciliation**: the Sharpe, drawdown, CAGR and hit rate the author states are parsed and compared with the computed ones: `CL-SHARPE`, `CL-DRAWDOWN`, `CL-CAGR`, `CL-WINRATE`. Only an excerpt of the series (first and last rows plus the monthly table) is sent to the model. The metrics were computed here over every row and travel as `facts.metrics`. ## The audit lane Verdict `credible`, `discounted` or `unreliable`, with one sentence naming what decided it. Then: an eight-bias status table in a fixed order - look-ahead, survivorship, overfitting, selection, transaction costs, data quality, execution realism, capacity - each `clear`, `suspected`, `confirmed` or `unknown` with evidence and a fix; numbered findings `BT-001`... most damaging first, each with severity, category, the place in the paste it cites, problem, impact, fix and a pasteable snippet; a haircut (reported Sharpe, expected Sharpe range, discount percent, the arithmetic); a validation plan (out-of-sample split, walk-forward, parameter stability, cost sensitivity, Monte Carlo on trade order, paper trading) made concrete for this code; and questions for the author. Invariants the page enforces and contradicts on screen when broken: a critical finding forbids `credible`; two confirmed biases forbid `credible`; `unreliable` needs a critical or high finding; a confirmed look-ahead means `unreliable`; a discount above 50 percent is inconsistent with `credible`; ids are sequential; every prescan flag is reconciled exactly once. ## The risk lane Verdict `deployable`, `size-down` or `not-yet`. Then: every metric read, with the memo's figure next to the browser's own and a badge when they disagree; the tail (VaR 95, CVaR 95, worst period, a reading); the drawdown episodes narrated; limits as numbers with actions; sizing (volatility target, Kelly fraction, maximum leverage, the arithmetic); stress scenarios with an expected loss at the intended scale; and a monitoring plan. Invariants: a max drawdown of 40 percent or worse forbids `deployable`; a Sharpe below 0.5 forbids `deployable`; a prior audit of `unreliable` forces `not-yet`; no series forces `not-yet` with an empty metrics table; every quoted metric must match the computed one. ## What it will not do It reads. It never runs your code, never fetches a price, never reaches an exchange or a broker, and never recommends buying or selling anything. It is a review of one backtest and a risk reading of one series, not investment advice, not a portfolio optimiser and not a live risk system. A credential in the paste is named and flagged for rotation, never repeated. Everything it returns is information, not investment advice. Past performance - and a backtest is a simulation of past performance - does not predict future results, and nothing on the page promises or projects a return. Verdicts such as "deployable" or "size-down" describe the risk evidence in one series, not a recommendation to trade. ## API `https://backtest-desk.skillsafe.ai/api.html` documents the `task` field, both lanes' input and output contracts and worked examples in eight languages. Model: `gpt-terra`. Runs are metered; the estimate is free; both bundled examples replay saved runs for free in both lanes. ## Sources Derived from two agent skills of the quantitative-trading plugin in wshobson/agents (MIT): `@wshobson/backtesting-frameworks` (look-ahead, survivorship and transaction-cost handling, walk-forward and Monte Carlo validation) and `@wshobson/risk-metrics-calculation` (VaR, CVaR, Sharpe, Sortino, drawdown analysis, stress testing). Not affiliated with their author.