← Backtest Desk / API
Tokens

Drive Backtest Desk from your own code

Everything the web page does is available over HTTP: send a strategy's backtest in - its code or description, an excerpt of the series it produced, the author's claims and the metrics you computed - and get back the same structured audit or risk memo. The natural uses are a research pipeline that gates a strategy's promotion to paper trading on the audit verdict, a batch job that audits every notebook in a repository and fails on a critical finding, and a risk memo generated per candidate before the allocation meeting.

Base URL and the envelope

Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses the same envelope, so one helper covers the whole API:

{ "ok": true,  "data":  { ... } }
{ "ok": false, "error": { "code": "...", "message": "...", "status": 402, "details": { ... } } }

Send your token as Authorization: Bearer … and Content-Type: application/json on every call. The app is identified by the token, which is minted for backtest-desk; there is no separate app header.

Error codes

codestatuswhat to do
unauthorized401The token is missing, malformed or expired. Get a new one from the token page.
payment_required402The balance is below min_credits. Call /estimate first and top up.
forbidden403A guest token tried a metered run. Sign in for a personal token; guests may call /me and /estimate only.
not_found404Unknown job id, or the app slug the token was minted for no longer exists.
conflict409The same Idempotency-Key was replayed with a different body. Change the key or send the original input.
validation_error400 / 422The body is not a JSON object, or task or strategy is missing or the wrong type. /estimate applies the same rule: a bare string, number, null or array is a 400.
rate_limited429Too many requests. Back off and retry; do not tight-loop.
internal5xxA server-side failure. Retry with the SAME Idempotency-Key so you are not billed twice.

1. Get a token

The easiest route is the token page: it shows the token this browser already holds, with Copy token and Copy shell export buttons, and a sign-in button for a personal token. Nothing on that page needs a developer tool.

A guest token can call /me and /estimate. Running an audit or a risk memo is metered, so it needs a personal token from signing in.

# The token page is the shortest path - it hands you a ready-made shell export:
#   https://backtest-desk.skillsafe.ai/tokens.html
#   export SKILLSAFE_TOKEN="aut_..."
#
# To mint a guest token from the command line instead (enough for /me and /estimate):
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
  -H "Content-Type: application/json" -d '{"slug":"backtest-desk"}'
# {"ok":true,"data":{"token":"gst_...","subject_type":"guest"}}

2. A tiny client

Every call is the same three things: the base URL, your bearer token, and a JSON body. One helper covers all of them.

# A shell function: call <path> [json-body]
BASE="https://api.skillsafe.ai/v1/app-api"
call() {
  if [ -n "$2" ]; then
    curl -sS -X POST "$BASE/$1" -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
      -H "Content-Type: application/json" --data-binary "$2"
  else
    curl -sS "$BASE/$1" -H "Authorization: Bearer $SKILLSAFE_TOKEN"
  fi
}

3. Check the session and the balance

GET /me tells you what the token is - subject_type is user for a personal token and guest otherwise - and the credits balance a run will be billed against. Compare it with the estimate's hold_credits before you run: a balance below min_credits is a 402, a balance between the two runs with a reduced output cap and returns truncated: true.

call me
# {"ok":true,"data":{"subject_type":"user","subject_id":"usr_...","credits":48210}}

4. Price the run — free

POST /estimate takes the exact input you are about to run - the same object, as the whole body - and returns what a run would reserve. It creates no job and costs nothing. Assert its model, model_alias and markup_bps once in CI: they prove the app is bound to the model and the markup it advertises.

Estimate the lane you will run. The two lanes carry different prompt sections and different output shapes, so hold_credits differs between task: "audit" and task: "risk". Never quote one lane's hold for the other. The hold is a reservation, not a price - charged_credits on the finished job is what you actually pay, usually far less.

The input, and the facts, honestly

The web page runs a free in-browser engine (btscan.js) before every run: it reads the pasted series, computes every metric over every row, lints the code for the classic pitfalls, and compares the author's claims with the numbers. Those results travel as facts and the prompt treats them as the only figures it may quote. Over the API there is no browser, so you fill facts. Two honest options:

The series itself never has to travel in full: series_excerpt is the first and last rows plus a monthly return table, and series_kind says what it was (equity, returns, trades or none).

# A shortened audit input. The real page sends the whole strategy and ~40 metrics.
INPUT='{"task":"audit","strategy":"# file: momentum.py\nimport pandas as pd\ntickers = pd.read_csv(\"sp500_current_constituents.csv\")[\"Symbol\"].tolist()\nprices = pd.read_csv(\"prices_adj_close.csv\", index_col=0, parse_dates=True)[tickers].dropna()\nrets = prices.pct_change()\nmom = prices / prices.shift(252) - prices / prices.shift(21)\nsignal = (mom.rank(axis=1, pct=True) > 0.8).astype(float).shift(-1)\nstrategy = (signal.div(signal.sum(axis=1), axis=0) * rets).sum(axis=1)","series_kind":"equity","series_excerpt":"date,equity\n2016-01-04,102229.33\n2016-01-05,101988.10\n... 2341 rows profiled in the browser and not sent ...\n2024-12-30,1948120.55\n2024-12-31,1951377.02","frequency":"daily","claims":"Sharpe 2.4, max drawdown 9%, CAGR 34% over 2016-2024 on the S&P 500 universe.","context":"US large-cap equities, long only. About 20 million. No costs modelled yet.","params":{"rf_rate":0.02,"cost_bps":5,"capital":100000},"facts":{"metrics":{"obs":2347,"years":9.31,"periods_per_year":252,"sharpe":2.39,"sortino":3.81,"cagr":0.4123,"ann_vol":0.1401,"max_drawdown":-0.1425,"max_dd_from":"2020-02-17","max_dd_to":"2020-11-16","calmar":2.89,"var_95":0.0133,"cvar_95":0.017,"skew":-0.02,"kurtosis":0.07,"win_rate":0.568,"profit_factor":1.49,"psr":1,"min_track_record":122},"flags":[{"id":"LA-NEG-SHIFT","label":"a negative shift reads a future row"},{"id":"SV-CURRENT-UNIVERSE","label":"the universe is today'"'"'s constituent list"},{"id":"SV-DROPNA","label":"dropna on the price panel removes delisted names"},{"id":"TC-NO-COSTS","label":"no commission, slippage, fee or spread anywhere in the code"},{"id":"OF-NO-OOS","label":"no out-of-sample, walk-forward or holdout period is mentioned"},{"id":"CL-DRAWDOWN","label":"the author claims a 9% max drawdown; the series shows -14.2%"}],"resources":[{"id":"RES-CODE","label":"python code, 8 lines"},{"id":"RES-SERIES","label":"2347 daily equity points, 2016-01-04 to 2024-12-31"}],"claims_parsed":{"sharpe":2.4,"max_drawdown":-0.09,"cagr":0.34,"period":"2016-2024"}}}'
call estimate "$INPUT"
# {"ok":true,"data":{"hold_credits":9310,"min_credits":1480,"model":"gpt-5.6-terra",
#   "model_alias":"gpt-terra","markup_bps":1000,"sponsor_enabled":false}}
# Then the same input with "task":"risk" - a different hold.

5. Run it, then poll

The task field comes first

Backtest Desk is one app with two lanes and one system prompt that routes on task. Both lanes take the same input - the same strategy, the same series facts, the same claims - and return the same outer envelope, so one client handles both.

taskquestionverdictsbody
auditCan this backtest result be believed?credible · discounted · unreliablebiases, findings, haircut, validation_plan, questions_for_author
riskHow much risk does it carry, and how should it be sized and limited?deployable · size-down · not-yetmetrics, tail_risk, drawdowns, limits, sizing, stress, monitoring

The natural order is audit first, risk second on the same input: pass the audit's verdict, haircut and top findings as prior_audit (section 8) and the memo sizes against the haircut Sharpe. If task is missing or unrecognised the model picks the closer lane, sets lane to the one it chose and says so in the first assumption - route on the reply's lane, not on what you asked for.

POST /run with the same input object as the body creates a job and returns job_id. Poll GET /jobs/{job_id} until status is succeeded or failed. The model's reply is the string at output.output; charged_credits is what you paid. Always send an Idempotency-Key derived from the lane and the input - a retried request with the same key returns the same job instead of billing twice.

# Key = lane + a hash of the input + an attempt counter. Same key, same body -> same job, no second charge.
KEY="backtest-desk:audit:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"
JOB=$(curl -sS -X POST "$BASE/run" -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
  -H "Content-Type: application/json" -H "Idempotency-Key: $KEY" --data-binary "$INPUT" \
  | python3 -c 'import sys,json; print(json.load(sys.stdin)["data"]["job_id"])')

while :; do
  STATUS=$(call "jobs/$JOB" | python3 -c 'import sys,json; print(json.load(sys.stdin)["data"]["status"])')
  [ "$STATUS" = "succeeded" ] || [ "$STATUS" = "failed" ] && break
  sleep 2
done
call "jobs/$JOB" | python3 -c 'import sys,json; d=json.load(sys.stdin)["data"]; print(d["charged_credits"], d.get("truncated")); print(d["output"]["output"][:400])'

6. Or stream it

POST /run-stream takes the same body and headers and answers with Server-Sent Events: an event: job frame carrying {"job_id": ...}, a series of event: delta frames each carrying {"text": "..."} - a chunk of the JSON reply - and a final event: done whose data is the same payload a finished job returns. Concatenate the deltas or read output.output from the done frame; they are the same text.

A script sees the deltas. A browser page does not: against a page, the endpoint sends heartbeat event: tick frames and only the final done, which is why the web app's progress card advances on elapsed time. Read the SSE yourself from a script if you want the live text.

curl -sN -X POST "$BASE/run-stream" -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
  -H "Content-Type: application/json" -H "Idempotency-Key: $KEY" --data-binary "$INPUT"
# event: job
# data: {"job_id":"job_..."}
#
# event: delta
# data: {"text":"{\"lane\":\"audit\",\"title\":\"12-1 momentum"}
# ...
# event: done
# data: {"job_id":"job_...","status":"succeeded","charged_credits":2140,"output":{"output":"{...}"}}

7. Parse the result

The reply is one JSON object as a string inside the envelope, so unwrap it twice. The model is told to return no prose and no code fences; be defensive anyway - strip a leading fence, take the first { to the last }, and parse. Then route on lane.

What the web app's normalizer does to it

Invariants worth asserting in CI

Audit

Risk

# Unwrap twice: the envelope, then the reply string.
call "jobs/$JOB" | python3 -c '
import sys, json
raw = json.load(sys.stdin)["data"]["output"]["output"].strip()
raw = raw[raw.find("{"): raw.rfind("}") + 1]
r = json.loads(raw)
print(r["lane"], r["verdict"], "-", r["verdict_reason"])
if r["lane"] == "audit":
    for f in r["findings"]: print(f["id"], f["severity"], f["category"], "|", f["target"])
    print("haircut", r["haircut"]["reported_sharpe"], "->", r["haircut"]["expected_sharpe"], r["haircut"]["discount_pct"], "%")
else:
    for m in r["metrics"]: print(m["metric"], m["value"], m["concern"])
    for l in r["limits"]: print("limit:", l["limit"], l["value"])
'

The output contract

One JSON object. The common envelope, then the body for the lane answered. Every key is present; "" or [] rather than an omission.

{
  "lane": "audit" | "risk",
  "title": "12-1 momentum, top quintile, S&P 500, 2016-2024",
  "verdict": "credible" | "discounted" | "unreliable"      // audit
           | "deployable" | "size-down" | "not-yet",       // risk
  "verdict_reason": "one sentence naming what decided it",
  "exec_summary": "two to five sentences",
  "assumptions": ["..."], "open_questions": ["..."],
  "coverage_check": [{ "id": "LA-NEG-SHIFT", "addressed": true, "note": "..." }],   // one per facts.flags id
  "quick_wins": ["..."], "summary": "one paragraph",

  // audit body
  "biases": [ { "bias": "look-ahead", "status": "clear|suspected|confirmed|unknown", "evidence": "...", "fix": "..." }, ... 8 in fixed order ],
  "findings": [ { "id": "BT-001", "severity": "critical|high|medium|low", "category": "<a bias>|other",
                  "target": "momentum.py line 22, signal = signal.shift(-1)", "problem": "...", "impact": "...", "fix": "...", "snippet": "..." } ],
  "haircut": { "reported_sharpe": "2.4", "expected_sharpe": "0.0 to 0.5", "discount_pct": 90, "reasoning": "..." },
  "validation_plan": [ { "step": "...", "why": "...", "how": "..." } ],
  "questions_for_author": ["..."],

  // risk body
  "metrics": [ { "metric": "Sharpe", "value": "0.59", "reading": "...", "concern": "low|medium|high" } ],
  "tail_risk": { "var_95": "1.93% per trade", "cvar_95": "2.45% per trade", "worst_period": "-4.1%", "reading": "..." },
  "drawdowns": [ { "rank": 1, "depth": "-11.74%", "from": "2022-02-14", "to": "not recovered", "recovery": "98 trades and counting", "reading": "..." } ],
  "limits": [ { "limit": "Daily VaR 95", "value": "0.6% of the book", "rationale": "..." } ],
  "sizing": { "vol_target": "6% annualised", "kelly_fraction": "0.25", "max_leverage": "1.5x", "reasoning": "..." },
  "stress": [ { "scenario": "2022 rates repricing repeat", "expected_loss": "-9% at target vol", "reasoning": "..." } ],
  "monitoring": [ { "signal": "Rolling 12m Sharpe", "threshold": "below 0.3", "action": "halve size and review" } ]
}

The enums

fieldvalues
laneaudit, risk
verdict (audit)credible, discounted, unreliable
verdict (risk)deployable, size-down, not-yet
biases[].statusclear, suspected, confirmed, unknown
findings[].severitycritical, high, medium, low
findings[].categoryone of the eight bias names, or other
metrics[].concernlow, medium, high
series_kind (input)equity, returns, trades, none
frequency (input)daily, weekly, monthly, quarterly, hourly, hourly-24, trades, unknown

The eight biases, in order

look-ahead, survivorship, overfitting, selection, transaction costs, data quality, execution realism, capacity. Always eight, always this order, always these exact strings.

The facts.metrics keys

Returns and drawdowns are decimals (0.124 is 12.4 percent, -0.31 a 31 percent drawdown); VaR and CVaR are positive losses per period. Send what you have; every key is optional and the contract says "not computable from the paste" for the rest.

obs, years, periods_per_year, frequency, kind, start, end,
total_return, cagr, ann_return, ann_vol, sharpe, sortino, rf_rate,
max_drawdown, max_dd_from, max_dd_trough, max_dd_to, max_dd_periods, longest_underwater, calmar,
var_95, cvar_95, var_99, cvar_99, skew, kurtosis,
win_rate, best, worst, avg_win, avg_loss, profit_factor, tail_ratio,
psr, min_track_record, rolling_sharpe_min, rolling_sharpe_max,
months, positive_months, worst_month, best_month,
trades, trades_per_year, cost_bps, cost_drag, sharpe_net,
drawdowns: [{ rank, depth, from, trough, to, periods }]

The flag ids the web app's lint can send

code:    LA-NEG-SHIFT LA-BFILL LA-CENTER LA-FULL-SCALE LA-NO-LAG LA-PINE-LOOKAHEAD LA-ILOC-NEXT LA-COC EX-SAME-BAR
         SV-CURRENT-UNIVERSE SV-DROPNA TC-NO-COSTS TC-ZERO OF-INSAMPLE-OPT OF-NO-OOS OF-MANY-PARAMS DQ-NO-SEED DQ-ADJ-LEVEL
series:  SR-NONE SR-SHORT SR-TOO-GOOD SR-NO-LOSSES SR-DD-TINY SR-PSR-LOW SR-TRACK-SHORT SR-FAT-TAILS SR-REGIME
         SR-BAD-DATES SR-UNSORTED SR-DUPES SR-NONPOSITIVE
claims:  CL-SHARPE CL-DRAWDOWN CL-CAGR CL-WINRATE

You may send your own ids too; the contract only requires that every id you send comes back exactly once in coverage_check.

8. The risk memo, carrying the audit

The second lane on the same backtest. Copy the audit's verdict, haircut and a few top findings into prior_audit: the memo then sizes against the haircut Sharpe rather than the reported one, and an unreliable audit forces not-yet. Everything else in the input is the same object with task switched to risk. This example is the second bundled backtest - a walk-forward futures carry book, 240 trades, real costs, a 2022 drawdown.

RISK='{"task":"risk","strategy":"# file: carry_wf.py\n\"\"\"Futures carry, long high roll yield, short low, weekly rebalance. Walk-forward: 3-year train, 6-month test, 2015-2024. Commission 2.50 per contract per side plus one tick of slippage.\"\"\"\nimport backtrader as bt\nnp.random.seed(7)\nclass Carry(bt.Strategy):\n    params = dict(lookback=20, legs=2)\n    def next(self):\n        self.order_target_percent(d, target)   # fills at the next open","series_kind":"trades","series_excerpt":"entry_date,exit_date,symbol,pnl\n2018-01-02,2018-01-08,FGBL,-7217\n2018-01-15,2018-01-23,ZF,-19031\n... 228 rows profiled in the browser and not sent ...\n2024-12-09,2024-12-16,6E,4120","frequency":"trades","claims":"Sharpe about 0.6 before costs, max drawdown about 12% in 2022, roughly 35 trades a year, 2018-2024 out-of-sample.","context":"Rates and FX futures on CME and Eurex, a 1,000,000 book, mid-day execution. We could size up to 5,000,000 if the risk memo supports it.","params":{"rf_rate":0.02,"cost_bps":3,"capital":1000000},"facts":{"metrics":{"obs":240,"years":6.94,"periods_per_year":34.6,"kind":"trades","sharpe":0.59,"sortino":0.9,"cagr":0.0544,"ann_vol":0.0586,"max_drawdown":-0.1174,"max_dd_from":"2022-02-14","max_dd_to":"not recovered","longest_underwater":98,"calmar":0.46,"var_95":0.0193,"cvar_95":0.0245,"cvar_99":0.0312,"skew":0.05,"kurtosis":1.02,"win_rate":0.55,"profit_factor":1.13,"psr":0.939,"min_track_record":272,"trades_per_year":34.6,"cost_drag":0.0208,"sharpe_net":0.24,"drawdowns":[{"rank":1,"depth":-0.1174,"from":"2022-02-14","trough":"2023-11-06","to":"not recovered","periods":98}]},"flags":[{"id":"SR-PSR-LOW","label":"probabilistic Sharpe ratio 0.939 - the sample cannot distinguish this Sharpe from zero at 95%"},{"id":"SR-TRACK-SHORT","label":"240 observations against a minimum track record of 272"}],"resources":[{"id":"RES-CODE","label":"python code, 33 lines"},{"id":"RES-SERIES","label":"240 trades over 6.9 years"}],"claims_parsed":{"sharpe":0.6,"max_drawdown":-0.12,"period":"2018-2024"}},"prior_audit":{"verdict":"discounted","haircut":{"reported_sharpe":"0.59","expected_sharpe":"0.25 to 0.45","discount_pct":42,"reasoning":"..."},"top_findings":["BT-001 (high, overfitting): the grid search inside each train window is unreported"]}}'
KEY2="backtest-desk:risk:$(printf '%s' "$RISK" | shasum -a 256 | cut -c1-16):a1"
curl -sS -X POST "$BASE/run" -H "Authorization: Bearer $SKILLSAFE_TOKEN" -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY2" --data-binary "$RISK"
# poll jobs/{job_id} as in step 5; expect lane "risk", verdict "size-down" or "not-yet"

9. Use it in a research pipeline

The audit verdict is a gate. A shell step that fails the pipeline on unreliable or on any critical finding is the smallest useful integration; a stricter shop also fails on a discount above 50 percent, and generates the risk memo only for what passes.

#!/bin/sh
# gate.sh - fail the pipeline when the audit says the backtest cannot be believed.
set -eu
: "${SKILLSAFE_TOKEN:?set SKILLSAFE_TOKEN from the token page}"
INPUT=$(python3 build_input.py strategies/momentum.py results/equity.csv)   # your metrics + lint -> the input object
KEY="backtest-desk:audit:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"
JOB=$(curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/run" -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
  -H "Content-Type: application/json" -H "Idempotency-Key: $KEY" --data-binary "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')
while :; do
  J=$(curl -sS "https://api.skillsafe.ai/v1/app-api/jobs/$JOB" -H "Authorization: Bearer $SKILLSAFE_TOKEN")
  S=$(printf '%s' "$J" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
  [ "$S" = succeeded ] && break
  [ "$S" = failed ] && { echo "run failed"; exit 2; }
  sleep 3
done
printf '%s' "$J" | python3 -c '
import sys, json
raw = json.load(sys.stdin)["data"]["output"]["output"]
r = json.loads(raw[raw.find("{"): raw.rfind("}") + 1])
crit = [f["id"] for f in r["findings"] if f["severity"] == "critical"]
print(r["verdict"], "-", r["verdict_reason"])
for f in r["findings"]: print(" ", f["id"], f["severity"], f["target"])
bad = r["verdict"] == "unreliable" or crit or r["haircut"]["discount_pct"] > 50
sys.exit(1 if bad else 0)
'

Truncation and partial results

If the balance sits between min_credits and hold_credits, the run still executes with a reduced output cap and the finished job carries truncated: true. The reply is then cut mid-JSON. Do what the web app does: try to close the object (append "}]}-style tails until one parses), keep every section that arrived, count them against the lane's section list (title, verdict, exec_summary, biases, findings, haircut, validation_plan, questions_for_author, coverage_check, summary for an audit; title, verdict, exec_summary, metrics, tail_risk, drawdowns, limits, sizing, stress, monitoring, coverage_check, summary for a memo), report "N of M sections", and top up before re-running with the same Idempotency-Key and the attempt counter incremented.

Backtest Desk reads a backtest and a series; it never runs code, fetches prices, reaches a broker or recommends a trade. Nothing it returns is investment advice, and past performance - simulated or real - does not predict future results. Derived from @wshobson/backtesting-frameworks and @wshobson/risk-metrics-calculation (wshobson/agents, MIT).