Drive Backtest Desk from your own code
Everything the web page does is available over HTTP: send a strategy's backtest in - its code or description, an excerpt of the series it produced, the author's claims and the metrics you computed - and get back the same structured audit or risk memo. The natural uses are a research pipeline that gates a strategy's promotion to paper trading on the audit verdict, a batch job that audits every notebook in a repository and fails on a critical finding, and a risk memo generated per candidate before the allocation meeting.
Base URL and the envelope
Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses
the same envelope, so one helper covers the whole API:
{ "ok": true, "data": { ... } }
{ "ok": false, "error": { "code": "...", "message": "...", "status": 402, "details": { ... } } }
Send your token as Authorization: Bearer … and Content-Type: application/json
on every call. The app is identified by the token, which is minted for backtest-desk;
there is no separate app header.
Error codes
| code | status | what to do |
|---|---|---|
unauthorized | 401 | The token is missing, malformed or expired. Get a new one from the token page. |
payment_required | 402 | The balance is below min_credits. Call /estimate first and top up. |
forbidden | 403 | A guest token tried a metered run. Sign in for a personal token; guests may call /me and /estimate only. |
not_found | 404 | Unknown job id, or the app slug the token was minted for no longer exists. |
conflict | 409 | The same Idempotency-Key was replayed with a different body. Change the key or send the original input. |
validation_error | 400 / 422 | The body is not a JSON object, or task or strategy is missing or the wrong type. /estimate applies the same rule: a bare string, number, null or array is a 400. |
rate_limited | 429 | Too many requests. Back off and retry; do not tight-loop. |
internal | 5xx | A server-side failure. Retry with the SAME Idempotency-Key so you are not billed twice. |
1. Get a token
The easiest route is the token page: it shows the token this browser already holds, with Copy token and Copy shell export buttons, and a sign-in button for a personal token. Nothing on that page needs a developer tool.
A guest token can call /me and /estimate. Running an
audit or a risk memo is metered, so it needs a personal token from signing in.
# The token page is the shortest path - it hands you a ready-made shell export:
# https://backtest-desk.skillsafe.ai/tokens.html
# export SKILLSAFE_TOKEN="aut_..."
#
# To mint a guest token from the command line instead (enough for /me and /estimate):
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
-H "Content-Type: application/json" -d '{"slug":"backtest-desk"}'
# {"ok":true,"data":{"token":"gst_...","subject_type":"guest"}}
# Open https://backtest-desk.skillsafe.ai/tokens.html and press "Copy token",
# or mint a guest token here. A guest token cannot run a metered lane.
import json, urllib.request
req = urllib.request.Request(
"https://api.skillsafe.ai/v1/app-api/guest",
data=json.dumps({"slug": "backtest-desk"}).encode(), method="POST")
req.add_header("Content-Type", "application/json")
with urllib.request.urlopen(req) as r:
TOKEN = json.load(r)["data"]["token"]
// Open https://backtest-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token cannot run a metered lane.
const res = await fetch("https://api.skillsafe.ai/v1/app-api/guest", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ slug: "backtest-desk" }),
});
const TOKEN = (await res.json()).data.token;
// Open https://backtest-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token cannot run a metered lane.
guestReq, _ := http.NewRequest(http.MethodPost,
"https://api.skillsafe.ai/v1/app-api/guest",
bytes.NewReader([]byte(`{"slug":"backtest-desk"}`)))
guestReq.Header.Set("Content-Type", "application/json")
guestRes, err := http.DefaultClient.Do(guestReq)
if err != nil {
panic(err)
}
defer guestRes.Body.Close()
var guest struct{ Data struct{ Token string `json:"token"` } `json:"data"` }
json.NewDecoder(guestRes.Body).Decode(&guest)
token := guest.Data.Token
// Open https://backtest-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token cannot run a metered lane.
var client = HttpClient.newHttpClient();
var guestReq = HttpRequest.newBuilder(URI.create("https://api.skillsafe.ai/v1/app-api/guest"))
.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString("{\"slug\":\"backtest-desk\"}"))
.build();
var guestBody = client.send(guestReq, HttpResponse.BodyHandlers.ofString()).body();
// parse with your JSON library of choice; the token is at data.token
String token = new ObjectMapper().readTree(guestBody).at("/data/token").asText();
# Open https://backtest-desk.skillsafe.ai/tokens.html and press "Copy token",
# or mint a guest token here. A guest token cannot run a metered lane.
require "net/http"; require "json"
uri = URI("https://api.skillsafe.ai/v1/app-api/guest")
res = Net::HTTP.post(uri, { slug: "backtest-desk" }.to_json, "Content-Type" => "application/json")
TOKEN = JSON.parse(res.body)["data"]["token"]
<?php
// Open https://backtest-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token cannot run a metered lane.
$ch = curl_init("https://api.skillsafe.ai/v1/app-api/guest");
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_HTTPHEADER => ["Content-Type: application/json"],
CURLOPT_POSTFIELDS => json_encode(["slug" => "backtest-desk"]),
CURLOPT_RETURNTRANSFER => true,
]);
$token = json_decode(curl_exec($ch), true)["data"]["token"];
// Open https://backtest-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token cannot run a metered lane.
var http = new HttpClient();
var guestRes = await http.PostAsync("https://api.skillsafe.ai/v1/app-api/guest",
new StringContent("{\"slug\":\"backtest-desk\"}", Encoding.UTF8, "application/json"));
using var guestDoc = JsonDocument.Parse(await guestRes.Content.ReadAsStringAsync());
var token = guestDoc.RootElement.GetProperty("data").GetProperty("token").GetString();
2. A tiny client
Every call is the same three things: the base URL, your bearer token, and a JSON body. One helper covers all of them.
# A shell function: call <path> [json-body]
BASE="https://api.skillsafe.ai/v1/app-api"
call() {
if [ -n "$2" ]; then
curl -sS -X POST "$BASE/$1" -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
-H "Content-Type: application/json" --data-binary "$2"
else
curl -sS "$BASE/$1" -H "Authorization: Bearer $SKILLSAFE_TOKEN"
fi
}
import json, urllib.request, urllib.error
BASE = "https://api.skillsafe.ai/v1/app-api"
def call(path, body=None, method=None, headers=None):
data = None if body is None else json.dumps(body).encode()
req = urllib.request.Request(f"{BASE}/{path}", data=data, method=method or ("POST" if data else "GET"))
req.add_header("Authorization", f"Bearer {TOKEN}")
req.add_header("Content-Type", "application/json")
for k, v in (headers or {}).items():
req.add_header(k, v)
try:
with urllib.request.urlopen(req) as r:
env = json.load(r)
except urllib.error.HTTPError as e:
env = json.load(e)
if not env.get("ok"):
raise RuntimeError(f"{env['error']['code']}: {env['error']['message']}")
return env["data"]
const BASE = "https://api.skillsafe.ai/v1/app-api";
async function call(path, body, headers = {}) {
const res = await fetch(`${BASE}/${path}`, {
method: body === undefined ? "GET" : "POST",
headers: { Authorization: `Bearer ${TOKEN}`, "Content-Type": "application/json", ...headers },
body: body === undefined ? undefined : JSON.stringify(body),
});
const env = await res.json();
if (!env.ok) throw new Error(`${env.error.code}: ${env.error.message}`);
return env.data;
}
const base = "https://api.skillsafe.ai/v1/app-api"
type envelope struct {
OK bool `json:"ok"`
Data json.RawMessage `json:"data"`
Error *struct {
Code string `json:"code"`
Message string `json:"message"`
} `json:"error"`
}
func call(token, path string, body any, headers map[string]string) (json.RawMessage, error) {
var rdr io.Reader
method := http.MethodGet
if body != nil {
b, _ := json.Marshal(body)
rdr = bytes.NewReader(b)
method = http.MethodPost
}
req, _ := http.NewRequest(method, base+"/"+path, rdr)
req.Header.Set("Authorization", "Bearer "+token)
req.Header.Set("Content-Type", "application/json")
for k, v := range headers {
req.Header.Set(k, v)
}
res, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
defer res.Body.Close()
var env envelope
if err := json.NewDecoder(res.Body).Decode(&env); err != nil {
return nil, err
}
if !env.OK {
return nil, fmt.Errorf("%s: %s", env.Error.Code, env.Error.Message)
}
return env.Data, nil
}
static final String BASE = "https://api.skillsafe.ai/v1/app-api";
static final HttpClient HTTP = HttpClient.newHttpClient();
static final ObjectMapper JSON = new ObjectMapper();
static JsonNode call(String token, String path, Object body, Map<String, String> headers) throws Exception {
var b = HttpRequest.newBuilder(URI.create(BASE + "/" + path))
.header("Authorization", "Bearer " + token)
.header("Content-Type", "application/json");
headers.forEach(b::header);
if (body == null) b.GET(); else b.POST(HttpRequest.BodyPublishers.ofString(JSON.writeValueAsString(body)));
var env = JSON.readTree(HTTP.send(b.build(), HttpResponse.BodyHandlers.ofString()).body());
if (!env.path("ok").asBoolean()) {
throw new RuntimeException(env.at("/error/code").asText() + ": " + env.at("/error/message").asText());
}
return env.get("data");
}
require "net/http"; require "json"
BASE = "https://api.skillsafe.ai/v1/app-api"
def call(path, body = nil, headers = {})
uri = URI("#{BASE}/#{path}")
req = body ? Net::HTTP::Post.new(uri) : Net::HTTP::Get.new(uri)
req["Authorization"] = "Bearer #{TOKEN}"
req["Content-Type"] = "application/json"
headers.each { |k, v| req[k] = v }
req.body = body.to_json if body
env = JSON.parse(Net::HTTP.start(uri.host, uri.port, use_ssl: true) { |h| h.request(req) }.body)
raise "#{env['error']['code']}: #{env['error']['message']}" unless env["ok"]
env["data"]
end
<?php
const BASE = "https://api.skillsafe.ai/v1/app-api";
function call(string $token, string $path, ?array $body = null, array $headers = []): array {
$ch = curl_init(BASE . "/" . $path);
$h = ["Authorization: Bearer $token", "Content-Type: application/json"];
foreach ($headers as $k => $v) $h[] = "$k: $v";
curl_setopt_array($ch, [CURLOPT_HTTPHEADER => $h, CURLOPT_RETURNTRANSFER => true]);
if ($body !== null) {
curl_setopt($ch, CURLOPT_POST, true);
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode($body));
}
$env = json_decode(curl_exec($ch), true);
if (!($env["ok"] ?? false)) throw new RuntimeException($env["error"]["code"] . ": " . $env["error"]["message"]);
return $env["data"];
}
const string Base = "https://api.skillsafe.ai/v1/app-api";
static readonly HttpClient Http = new HttpClient();
static async Task<JsonElement> Call(string token, string path, object? body = null, Dictionary<string, string>? headers = null) {
var req = new HttpRequestMessage(body is null ? HttpMethod.Get : HttpMethod.Post, $"{Base}/{path}");
req.Headers.Authorization = new AuthenticationHeaderValue("Bearer", token);
if (body is not null) req.Content = new StringContent(JsonSerializer.Serialize(body), Encoding.UTF8, "application/json");
foreach (var (k, v) in headers ?? new()) req.Headers.TryAddWithoutValidation(k, v);
var env = JsonDocument.Parse(await (await Http.SendAsync(req)).Content.ReadAsStringAsync()).RootElement;
if (!env.GetProperty("ok").GetBoolean()) {
var e = env.GetProperty("error");
throw new Exception($"{e.GetProperty("code")}: {e.GetProperty("message")}");
}
return env.GetProperty("data");
}
3. Check the session and the balance
GET /me tells you what the token is - subject_type is user for a
personal token and guest otherwise - and the credits balance a run will be
billed against. Compare it with the estimate's hold_credits before you run: a balance
below min_credits is a 402, a balance between the two runs with a reduced output cap and
returns truncated: true.
call me
# {"ok":true,"data":{"subject_type":"user","subject_id":"usr_...","credits":48210}}
me = call("me")
print(me["subject_type"], me["credits"]) # user 48210
const me = await call("me");
console.log(me.subject_type, me.credits); // user 48210
raw, err := call(token, "me", nil, nil)
if err != nil {
panic(err)
}
var me struct {
SubjectType string `json:"subject_type"`
Credits int `json:"credits"`
}
json.Unmarshal(raw, &me)
fmt.Println(me.SubjectType, me.Credits)
var me = call(token, "me", null, Map.of());
System.out.println(me.get("subject_type").asText() + " " + me.get("credits").asInt());
me = call("me")
puts "#{me['subject_type']} #{me['credits']}"
$me = call($token, "me");
echo $me["subject_type"], " ", $me["credits"], "\n";
var me = await Call(token, "me");
Console.WriteLine($"{me.GetProperty("subject_type")} {me.GetProperty("credits")}");
4. Price the run — free
POST /estimate takes the exact input you are about to run - the same object, as the
whole body - and returns what a run would reserve. It creates no job and costs nothing. Assert its
model, model_alias and markup_bps once in CI: they prove the app
is bound to the model and the markup it advertises.
Estimate the lane you will run. The two lanes carry different prompt sections and
different output shapes, so hold_credits differs between task: "audit" and
task: "risk". Never quote one lane's hold for the other. The hold is a reservation, not a
price - charged_credits on the finished job is what you actually pay, usually far less.
The input, and the facts, honestly
The web page runs a free in-browser engine (btscan.js) before every run: it reads the
pasted series, computes every metric over every row, lints the code for the classic pitfalls, and
compares the author's claims with the numbers. Those results travel as facts and the
prompt treats them as the only figures it may quote. Over the API there is no browser, so
you fill facts. Two honest options:
- Compute the metrics yourself (Sharpe, Sortino, drawdowns, VaR, CVaR and the rest - the key names are listed under the output contract) and lint your own code, then send them in the same shape. This is what makes the audit sharp and the risk memo numerical.
- Send
facts: {"metrics": {}, "flags": [], "resources": [], "claims_parsed": {}}. The audit then works from the code and the description alone and marks most biasesunknownorsuspected; the risk lane, by contract, returnsverdict: "not-yet"with an emptymetricstable, because a risk memo without numbers is a request for the numbers.
The series itself never has to travel in full: series_excerpt is the first and last rows
plus a monthly return table, and series_kind says what it was (equity,
returns, trades or none).
# A shortened audit input. The real page sends the whole strategy and ~40 metrics.
INPUT='{"task":"audit","strategy":"# file: momentum.py\nimport pandas as pd\ntickers = pd.read_csv(\"sp500_current_constituents.csv\")[\"Symbol\"].tolist()\nprices = pd.read_csv(\"prices_adj_close.csv\", index_col=0, parse_dates=True)[tickers].dropna()\nrets = prices.pct_change()\nmom = prices / prices.shift(252) - prices / prices.shift(21)\nsignal = (mom.rank(axis=1, pct=True) > 0.8).astype(float).shift(-1)\nstrategy = (signal.div(signal.sum(axis=1), axis=0) * rets).sum(axis=1)","series_kind":"equity","series_excerpt":"date,equity\n2016-01-04,102229.33\n2016-01-05,101988.10\n... 2341 rows profiled in the browser and not sent ...\n2024-12-30,1948120.55\n2024-12-31,1951377.02","frequency":"daily","claims":"Sharpe 2.4, max drawdown 9%, CAGR 34% over 2016-2024 on the S&P 500 universe.","context":"US large-cap equities, long only. About 20 million. No costs modelled yet.","params":{"rf_rate":0.02,"cost_bps":5,"capital":100000},"facts":{"metrics":{"obs":2347,"years":9.31,"periods_per_year":252,"sharpe":2.39,"sortino":3.81,"cagr":0.4123,"ann_vol":0.1401,"max_drawdown":-0.1425,"max_dd_from":"2020-02-17","max_dd_to":"2020-11-16","calmar":2.89,"var_95":0.0133,"cvar_95":0.017,"skew":-0.02,"kurtosis":0.07,"win_rate":0.568,"profit_factor":1.49,"psr":1,"min_track_record":122},"flags":[{"id":"LA-NEG-SHIFT","label":"a negative shift reads a future row"},{"id":"SV-CURRENT-UNIVERSE","label":"the universe is today'"'"'s constituent list"},{"id":"SV-DROPNA","label":"dropna on the price panel removes delisted names"},{"id":"TC-NO-COSTS","label":"no commission, slippage, fee or spread anywhere in the code"},{"id":"OF-NO-OOS","label":"no out-of-sample, walk-forward or holdout period is mentioned"},{"id":"CL-DRAWDOWN","label":"the author claims a 9% max drawdown; the series shows -14.2%"}],"resources":[{"id":"RES-CODE","label":"python code, 8 lines"},{"id":"RES-SERIES","label":"2347 daily equity points, 2016-01-04 to 2024-12-31"}],"claims_parsed":{"sharpe":2.4,"max_drawdown":-0.09,"cagr":0.34,"period":"2016-2024"}}}'
call estimate "$INPUT"
# {"ok":true,"data":{"hold_credits":9310,"min_credits":1480,"model":"gpt-5.6-terra",
# "model_alias":"gpt-terra","markup_bps":1000,"sponsor_enabled":false}}
# Then the same input with "task":"risk" - a different hold.
INPUT = {
"task": "audit",
"strategy": open("momentum.py").read(),
"series_kind": "equity",
"series_excerpt": "date,equity\n2016-01-04,102229.33\n...\n2024-12-31,1951377.02",
"frequency": "daily",
"claims": "Sharpe 2.4, max drawdown 9%, CAGR 34% over 2016-2024 on the S&P 500 universe.",
"context": "US large-cap equities, long only. About 20 million. No costs modelled yet.",
"params": {"rf_rate": 0.02, "cost_bps": 5, "capital": 100000},
"facts": {
"metrics": {"obs": 2347, "years": 9.31, "periods_per_year": 252, "sharpe": 2.39, "sortino": 3.81,
"cagr": 0.4123, "ann_vol": 0.1401, "max_drawdown": -0.1425, "calmar": 2.89,
"var_95": 0.0133, "cvar_95": 0.017, "skew": -0.02, "kurtosis": 0.07,
"win_rate": 0.568, "profit_factor": 1.49, "psr": 1, "min_track_record": 122},
"flags": [{"id": "LA-NEG-SHIFT", "label": "a negative shift reads a future row"},
{"id": "SV-CURRENT-UNIVERSE", "label": "the universe is today's constituent list"},
{"id": "TC-NO-COSTS", "label": "no commission, slippage, fee or spread anywhere in the code"}],
"resources": [{"id": "RES-CODE", "label": "python code, 8 lines"}],
"claims_parsed": {"sharpe": 2.4, "max_drawdown": -0.09, "cagr": 0.34, "period": "2016-2024"},
},
}
est = call("estimate", INPUT)
assert est["model_alias"] == "gpt-terra" and est["markup_bps"] == 1000
print(est["hold_credits"], est["min_credits"]) # e.g. 9310 1480
estRisk = call("estimate", {**INPUT, "task": "risk"}) # a different hold
const INPUT = {
task: "audit",
strategy: await fs.readFile("momentum.py", "utf8"),
series_kind: "equity",
series_excerpt: "date,equity\n2016-01-04,102229.33\n...\n2024-12-31,1951377.02",
frequency: "daily",
claims: "Sharpe 2.4, max drawdown 9%, CAGR 34% over 2016-2024 on the S&P 500 universe.",
context: "US large-cap equities, long only. About 20 million. No costs modelled yet.",
params: { rf_rate: 0.02, cost_bps: 5, capital: 100000 },
facts: {
metrics: { obs: 2347, years: 9.31, periods_per_year: 252, sharpe: 2.39, sortino: 3.81, cagr: 0.4123,
ann_vol: 0.1401, max_drawdown: -0.1425, calmar: 2.89, var_95: 0.0133, cvar_95: 0.017,
skew: -0.02, kurtosis: 0.07, win_rate: 0.568, profit_factor: 1.49, psr: 1, min_track_record: 122 },
flags: [{ id: "LA-NEG-SHIFT", label: "a negative shift reads a future row" },
{ id: "SV-CURRENT-UNIVERSE", label: "the universe is today's constituent list" },
{ id: "TC-NO-COSTS", label: "no commission, slippage, fee or spread anywhere in the code" }],
resources: [{ id: "RES-CODE", label: "python code, 8 lines" }],
claims_parsed: { sharpe: 2.4, max_drawdown: -0.09, cagr: 0.34, period: "2016-2024" },
},
};
const est = await call("estimate", INPUT);
console.assert(est.model_alias === "gpt-terra" && est.markup_bps === 1000);
console.log(est.hold_credits, est.min_credits); // e.g. 9310 1480
const riskEst = await call("estimate", { ...INPUT, task: "risk" });
input := map[string]any{
"task": "audit",
"strategy": string(mustRead("momentum.py")),
"series_kind": "equity",
"series_excerpt": "date,equity\n2016-01-04,102229.33\n...\n2024-12-31,1951377.02",
"frequency": "daily",
"claims": "Sharpe 2.4, max drawdown 9%, CAGR 34% over 2016-2024 on the S&P 500 universe.",
"context": "US large-cap equities, long only. About 20 million. No costs modelled yet.",
"params": map[string]any{"rf_rate": 0.02, "cost_bps": 5, "capital": 100000},
"facts": map[string]any{
"metrics": map[string]any{"obs": 2347, "years": 9.31, "periods_per_year": 252, "sharpe": 2.39,
"cagr": 0.4123, "ann_vol": 0.1401, "max_drawdown": -0.1425, "var_95": 0.0133, "cvar_95": 0.017,
"skew": -0.02, "kurtosis": 0.07, "win_rate": 0.568, "psr": 1, "min_track_record": 122},
"flags": []map[string]string{{"id": "LA-NEG-SHIFT", "label": "a negative shift reads a future row"},
{"id": "TC-NO-COSTS", "label": "no commission, slippage, fee or spread anywhere in the code"}},
"resources": []map[string]string{{"id": "RES-CODE", "label": "python code, 8 lines"}},
"claims_parsed": map[string]any{"sharpe": 2.4, "max_drawdown": -0.09, "cagr": 0.34},
},
}
raw, err := call(token, "estimate", input, nil)
if err != nil {
panic(err)
}
var est struct {
Hold int `json:"hold_credits"`
Min int `json:"min_credits"`
ModelAlias string `json:"model_alias"`
MarkupBps int `json:"markup_bps"`
}
json.Unmarshal(raw, &est)
fmt.Println(est.Hold, est.Min, est.ModelAlias, est.MarkupBps)
Map<String, Object> input = new LinkedHashMap<>();
input.put("task", "audit");
input.put("strategy", Files.readString(Path.of("momentum.py")));
input.put("series_kind", "equity");
input.put("series_excerpt", "date,equity\n2016-01-04,102229.33\n...\n2024-12-31,1951377.02");
input.put("frequency", "daily");
input.put("claims", "Sharpe 2.4, max drawdown 9%, CAGR 34% over 2016-2024 on the S&P 500 universe.");
input.put("context", "US large-cap equities, long only. About 20 million. No costs modelled yet.");
input.put("params", Map.of("rf_rate", 0.02, "cost_bps", 5, "capital", 100000));
input.put("facts", Map.of(
"metrics", Map.of("obs", 2347, "sharpe", 2.39, "cagr", 0.4123, "ann_vol", 0.1401, "max_drawdown", -0.1425,
"var_95", 0.0133, "cvar_95", 0.017, "win_rate", 0.568, "psr", 1, "min_track_record", 122),
"flags", List.of(Map.of("id", "LA-NEG-SHIFT", "label", "a negative shift reads a future row"),
Map.of("id", "TC-NO-COSTS", "label", "no commission, slippage, fee or spread anywhere in the code")),
"resources", List.of(Map.of("id", "RES-CODE", "label", "python code, 8 lines")),
"claims_parsed", Map.of("sharpe", 2.4, "max_drawdown", -0.09, "cagr", 0.34)));
var est = call(token, "estimate", input, Map.of());
System.out.println(est.get("hold_credits").asInt() + " " + est.get("model_alias").asText());
INPUT = {
task: "audit",
strategy: File.read("momentum.py"),
series_kind: "equity",
series_excerpt: "date,equity\n2016-01-04,102229.33\n...\n2024-12-31,1951377.02",
frequency: "daily",
claims: "Sharpe 2.4, max drawdown 9%, CAGR 34% over 2016-2024 on the S&P 500 universe.",
context: "US large-cap equities, long only. About 20 million. No costs modelled yet.",
params: { rf_rate: 0.02, cost_bps: 5, capital: 100_000 },
facts: {
metrics: { obs: 2347, sharpe: 2.39, cagr: 0.4123, ann_vol: 0.1401, max_drawdown: -0.1425,
var_95: 0.0133, cvar_95: 0.017, win_rate: 0.568, psr: 1, min_track_record: 122 },
flags: [{ id: "LA-NEG-SHIFT", label: "a negative shift reads a future row" },
{ id: "TC-NO-COSTS", label: "no commission, slippage, fee or spread anywhere in the code" }],
resources: [{ id: "RES-CODE", label: "python code, 8 lines" }],
claims_parsed: { sharpe: 2.4, max_drawdown: -0.09, cagr: 0.34 }
}
}
est = call("estimate", INPUT)
raise "wrong binding" unless est["model_alias"] == "gpt-terra" && est["markup_bps"] == 1000
puts est["hold_credits"]
estRisk = call("estimate", INPUT.merge(task: "risk"))
$input = [
"task" => "audit",
"strategy" => file_get_contents("momentum.py"),
"series_kind" => "equity",
"series_excerpt" => "date,equity\n2016-01-04,102229.33\n...\n2024-12-31,1951377.02",
"frequency" => "daily",
"claims" => "Sharpe 2.4, max drawdown 9%, CAGR 34% over 2016-2024 on the S&P 500 universe.",
"context" => "US large-cap equities, long only. About 20 million. No costs modelled yet.",
"params" => ["rf_rate" => 0.02, "cost_bps" => 5, "capital" => 100000],
"facts" => [
"metrics" => ["obs" => 2347, "sharpe" => 2.39, "cagr" => 0.4123, "ann_vol" => 0.1401, "max_drawdown" => -0.1425,
"var_95" => 0.0133, "cvar_95" => 0.017, "win_rate" => 0.568, "psr" => 1, "min_track_record" => 122],
"flags" => [["id" => "LA-NEG-SHIFT", "label" => "a negative shift reads a future row"],
["id" => "TC-NO-COSTS", "label" => "no commission, slippage, fee or spread anywhere in the code"]],
"resources" => [["id" => "RES-CODE", "label" => "python code, 8 lines"]],
"claims_parsed" => ["sharpe" => 2.4, "max_drawdown" => -0.09, "cagr" => 0.34],
],
];
$est = call($token, "estimate", $input);
echo $est["hold_credits"], " ", $est["model_alias"], "\n";
$riskEst = call($token, "estimate", array_merge($input, ["task" => "risk"]));
var input = new Dictionary<string, object> {
["task"] = "audit",
["strategy"] = File.ReadAllText("momentum.py"),
["series_kind"] = "equity",
["series_excerpt"] = "date,equity\n2016-01-04,102229.33\n...\n2024-12-31,1951377.02",
["frequency"] = "daily",
["claims"] = "Sharpe 2.4, max drawdown 9%, CAGR 34% over 2016-2024 on the S&P 500 universe.",
["context"] = "US large-cap equities, long only. About 20 million. No costs modelled yet.",
["params"] = new { rf_rate = 0.02, cost_bps = 5, capital = 100000 },
["facts"] = new {
metrics = new { obs = 2347, sharpe = 2.39, cagr = 0.4123, ann_vol = 0.1401, max_drawdown = -0.1425,
var_95 = 0.0133, cvar_95 = 0.017, win_rate = 0.568, psr = 1, min_track_record = 122 },
flags = new[] { new { id = "LA-NEG-SHIFT", label = "a negative shift reads a future row" },
new { id = "TC-NO-COSTS", label = "no commission, slippage, fee or spread anywhere in the code" } },
resources = new[] { new { id = "RES-CODE", label = "python code, 8 lines" } },
claims_parsed = new { sharpe = 2.4, max_drawdown = -0.09, cagr = 0.34 },
},
};
var est = await Call(token, "estimate", input);
Console.WriteLine($"{est.GetProperty("hold_credits")} {est.GetProperty("model_alias")}");
input["task"] = "risk";
var riskEst = await Call(token, "estimate", input);
5. Run it, then poll
The task field comes first
Backtest Desk is one app with two lanes and one system prompt that routes on task. Both
lanes take the same input - the same strategy, the same series facts, the same claims - and return
the same outer envelope, so one client handles both.
task | question | verdicts | body |
|---|---|---|---|
audit | Can this backtest result be believed? | credible · discounted · unreliable | biases, findings, haircut, validation_plan, questions_for_author |
risk | How much risk does it carry, and how should it be sized and limited? | deployable · size-down · not-yet | metrics, tail_risk, drawdowns, limits, sizing, stress, monitoring |
The natural order is audit first, risk second on the same input: pass the audit's verdict,
haircut and top findings as prior_audit (section 8) and the memo sizes against the
haircut Sharpe. If task is missing or unrecognised the model picks the closer lane, sets
lane to the one it chose and says so in the first assumption - route on the reply's
lane, not on what you asked for.
POST /run with the same input object as the body creates a job and returns job_id.
Poll GET /jobs/{job_id} until status is succeeded or failed. The
model's reply is the string at output.output; charged_credits is what you paid.
Always send an Idempotency-Key derived from the lane and the input - a retried
request with the same key returns the same job instead of billing twice.
# Key = lane + a hash of the input + an attempt counter. Same key, same body -> same job, no second charge.
KEY="backtest-desk:audit:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"
JOB=$(curl -sS -X POST "$BASE/run" -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
-H "Content-Type: application/json" -H "Idempotency-Key: $KEY" --data-binary "$INPUT" \
| python3 -c 'import sys,json; print(json.load(sys.stdin)["data"]["job_id"])')
while :; do
STATUS=$(call "jobs/$JOB" | python3 -c 'import sys,json; print(json.load(sys.stdin)["data"]["status"])')
[ "$STATUS" = "succeeded" ] || [ "$STATUS" = "failed" ] && break
sleep 2
done
call "jobs/$JOB" | python3 -c 'import sys,json; d=json.load(sys.stdin)["data"]; print(d["charged_credits"], d.get("truncated")); print(d["output"]["output"][:400])'
import hashlib, time
def idem_key(inp, attempt=1):
h = hashlib.sha256(json.dumps(inp, sort_keys=True).encode()).hexdigest()[:16]
return f"backtest-desk:{inp['task']}:{h}:a{attempt}"
job = call("run", INPUT, headers={"Idempotency-Key": idem_key(INPUT)})
while True:
j = call(f"jobs/{job['job_id']}")
if j["status"] in ("succeeded", "failed"):
break
time.sleep(2)
if j["status"] == "failed":
raise SystemExit(j.get("error"))
raw = j["output"]["output"] # the JSON audit, as a string
print(j["charged_credits"], j.get("truncated"))
import { createHash } from "node:crypto";
const idemKey = (inp, attempt = 1) =>
`backtest-desk:${inp.task}:${createHash("sha256").update(JSON.stringify(inp)).digest("hex").slice(0, 16)}:a${attempt}`;
const job = await call("run", INPUT, { "Idempotency-Key": idemKey(INPUT) });
let j;
for (;;) {
j = await call(`jobs/${job.job_id}`);
if (j.status === "succeeded" || j.status === "failed") break;
await new Promise(r => setTimeout(r, 2000));
}
if (j.status === "failed") throw new Error(JSON.stringify(j.error));
const raw = j.output.output; // the JSON audit, as a string
console.log(j.charged_credits, j.truncated);
b, _ := json.Marshal(input)
sum := sha256.Sum256(b)
key := fmt.Sprintf("backtest-desk:%s:%x:a1", input["task"], sum[:8])
raw, err := call(token, "run", input, map[string]string{"Idempotency-Key": key})
if err != nil {
panic(err)
}
var job struct{ JobID string `json:"job_id"` }
json.Unmarshal(raw, &job)
var j struct {
Status string `json:"status"`
Charged int `json:"charged_credits"`
Output struct{ Output string `json:"output"` } `json:"output"`
}
for {
raw, _ = call(token, "jobs/"+job.JobID, nil, nil)
json.Unmarshal(raw, &j)
if j.Status == "succeeded" || j.Status == "failed" {
break
}
time.Sleep(2 * time.Second)
}
audit := j.Output.Output // the JSON audit, as a string
String body = JSON.writeValueAsString(input);
String hash = HexFormat.of().formatHex(MessageDigest.getInstance("SHA-256").digest(body.getBytes())).substring(0, 16);
String key = "backtest-desk:" + input.get("task") + ":" + hash + ":a1";
var job = call(token, "run", input, Map.of("Idempotency-Key", key));
JsonNode j;
while (true) {
j = call(token, "jobs/" + job.get("job_id").asText(), null, Map.of());
String s = j.get("status").asText();
if (s.equals("succeeded") || s.equals("failed")) break;
Thread.sleep(2000);
}
String raw = j.at("/output/output").asText(); // the JSON audit, as a string
System.out.println(j.get("charged_credits").asInt());
require "digest"
key = "backtest-desk:#{INPUT[:task]}:#{Digest::SHA256.hexdigest(INPUT.to_json)[0, 16]}:a1"
job = call("run", INPUT, "Idempotency-Key" => key)
loop do
@j = call("jobs/#{job['job_id']}")
break if %w[succeeded failed].include?(@j["status"])
sleep 2
end
raw = @j["output"]["output"] # the JSON audit, as a string
puts @j["charged_credits"]
$key = "backtest-desk:" . $input["task"] . ":" . substr(hash("sha256", json_encode($input)), 0, 16) . ":a1";
$job = call($token, "run", $input, ["Idempotency-Key" => $key]);
do {
$j = call($token, "jobs/" . $job["job_id"]);
if (in_array($j["status"], ["succeeded", "failed"])) break;
sleep(2);
} while (true);
$raw = $j["output"]["output"]; // the JSON audit, as a string
echo $j["charged_credits"], "\n";
var bodyJson = JsonSerializer.Serialize(input);
var hash = Convert.ToHexString(SHA256.HashData(Encoding.UTF8.GetBytes(bodyJson)))[..16].ToLower();
var key = $"backtest-desk:{input["task"]}:{hash}:a1";
var job = await Call(token, "run", input, new() { ["Idempotency-Key"] = key });
JsonElement j;
while (true) {
j = await Call(token, $"jobs/{job.GetProperty("job_id").GetString()}");
var s = j.GetProperty("status").GetString();
if (s == "succeeded" || s == "failed") break;
await Task.Delay(2000);
}
var raw = j.GetProperty("output").GetProperty("output").GetString(); // the JSON audit, as a string
Console.WriteLine(j.GetProperty("charged_credits"));
6. Or stream it
POST /run-stream takes the same body and headers and answers with Server-Sent Events: an
event: job frame carrying {"job_id": ...}, a series of event: delta
frames each carrying {"text": "..."} - a chunk of the JSON reply - and a final
event: done whose data is the same payload a finished job returns. Concatenate the deltas
or read output.output from the done frame; they are the same text.
A script sees the deltas. A browser page does not: against a page, the endpoint sends heartbeat
event: tick frames and only the final done, which is why the web app's progress
card advances on elapsed time. Read the SSE yourself from a script if you want the live text.
curl -sN -X POST "$BASE/run-stream" -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
-H "Content-Type: application/json" -H "Idempotency-Key: $KEY" --data-binary "$INPUT"
# event: job
# data: {"job_id":"job_..."}
#
# event: delta
# data: {"text":"{\"lane\":\"audit\",\"title\":\"12-1 momentum"}
# ...
# event: done
# data: {"job_id":"job_...","status":"succeeded","charged_credits":2140,"output":{"output":"{...}"}}
req = urllib.request.Request(f"{BASE}/run-stream", data=json.dumps(INPUT).encode(), method="POST")
for k, v in {"Authorization": f"Bearer {TOKEN}", "Content-Type": "application/json",
"Idempotency-Key": idem_key(INPUT)}.items():
req.add_header(k, v)
event, buf, done = None, [], None
with urllib.request.urlopen(req) as r:
for line in r:
line = line.decode().rstrip("\n")
if line.startswith("event: "):
event = line[7:]
elif line.startswith("data: "):
data = json.loads(line[6:])
if event == "delta":
buf.append(data.get("text", ""))
elif event == "done":
done = data
raw = done["output"]["output"] if done else "".join(buf)
const res = await fetch(`${BASE}/run-stream`, {
method: "POST",
headers: { Authorization: `Bearer ${TOKEN}`, "Content-Type": "application/json", "Idempotency-Key": idemKey(INPUT) },
body: JSON.stringify(INPUT),
});
const reader = res.body.getReader(), dec = new TextDecoder();
let buffer = "", event = "message", text = "", done = null;
for (;;) {
const { value, done: end } = await reader.read();
if (end) break;
buffer += dec.decode(value, { stream: true });
let idx;
while ((idx = buffer.indexOf("\n\n")) >= 0) {
const frame = buffer.slice(0, idx); buffer = buffer.slice(idx + 2);
for (const line of frame.split("\n")) {
if (line.startsWith("event: ")) event = line.slice(7);
else if (line.startsWith("data: ")) {
const data = JSON.parse(line.slice(6));
if (event === "delta") text += data.text || "";
else if (event === "done") done = data;
}
}
}
}
const raw = done ? done.output.output : text;
req, _ := http.NewRequest(http.MethodPost, base+"/run-stream", bytes.NewReader(b))
req.Header.Set("Authorization", "Bearer "+token)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", key)
res, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer res.Body.Close()
sc := bufio.NewScanner(res.Body)
sc.Buffer(make([]byte, 1<<20), 1<<20)
event, text := "", strings.Builder{}
var done map[string]any
for sc.Scan() {
line := sc.Text()
switch {
case strings.HasPrefix(line, "event: "):
event = line[7:]
case strings.HasPrefix(line, "data: "):
var data map[string]any
json.Unmarshal([]byte(line[6:]), &data)
if event == "delta" {
text.WriteString(fmt.Sprint(data["text"]))
} else if event == "done" {
done = data
}
}
}
_ = done // done["output"].(map[string]any)["output"] is the full reply; text.String() is the same
var req = HttpRequest.newBuilder(URI.create(BASE + "/run-stream"))
.header("Authorization", "Bearer " + token)
.header("Content-Type", "application/json")
.header("Idempotency-Key", key)
.POST(HttpRequest.BodyPublishers.ofString(body)).build();
var res = HTTP.send(req, HttpResponse.BodyHandlers.ofLines());
String event = "";
var text = new StringBuilder();
JsonNode done = null;
for (String line : (Iterable<String>) res.body()::iterator) {
if (line.startsWith("event: ")) event = line.substring(7);
else if (line.startsWith("data: ")) {
var data = JSON.readTree(line.substring(6));
if (event.equals("delta")) text.append(data.path("text").asText(""));
else if (event.equals("done")) done = data;
}
}
String raw = done != null ? done.at("/output/output").asText() : text.toString();
uri = URI("#{BASE}/run-stream")
req = Net::HTTP::Post.new(uri, "Authorization" => "Bearer #{TOKEN}", "Content-Type" => "application/json", "Idempotency-Key" => key)
req.body = INPUT.to_json
event, text, done = nil, +"", nil
Net::HTTP.start(uri.host, uri.port, use_ssl: true) do |h|
h.request(req) do |res|
res.read_body do |chunk|
chunk.each_line do |line|
if line.start_with?("event: ") then event = line[7..].strip
elsif line.start_with?("data: ")
data = JSON.parse(line[6..])
text << data["text"].to_s if event == "delta"
done = data if event == "done"
end
end
end
end
end
raw = done ? done["output"]["output"] : text
$ch = curl_init(BASE . "/run-stream");
$event = ""; $text = ""; $done = null;
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_HTTPHEADER => ["Authorization: Bearer $token", "Content-Type: application/json", "Idempotency-Key: $key"],
CURLOPT_POSTFIELDS => json_encode($input),
CURLOPT_WRITEFUNCTION => function ($ch, $chunk) use (&$event, &$text, &$done) {
foreach (explode("\n", $chunk) as $line) {
if (str_starts_with($line, "event: ")) $event = trim(substr($line, 7));
elseif (str_starts_with($line, "data: ")) {
$data = json_decode(substr($line, 6), true);
if ($event === "delta") $text .= $data["text"] ?? "";
elseif ($event === "done") $done = $data;
}
}
return strlen($chunk);
},
]);
curl_exec($ch);
$raw = $done ? $done["output"]["output"] : $text;
var req = new HttpRequestMessage(HttpMethod.Post, $"{Base}/run-stream") {
Content = new StringContent(bodyJson, Encoding.UTF8, "application/json") };
req.Headers.Authorization = new AuthenticationHeaderValue("Bearer", token);
req.Headers.Add("Idempotency-Key", key);
using var res = await Http.SendAsync(req, HttpCompletionOption.ResponseHeadersRead);
using var reader = new StreamReader(await res.Content.ReadAsStreamAsync());
string evt = "", text = ""; JsonElement? done = null;
while (await reader.ReadLineAsync() is { } line) {
if (line.StartsWith("event: ")) evt = line[7..];
else if (line.StartsWith("data: ")) {
var data = JsonDocument.Parse(line[6..]).RootElement;
if (evt == "delta") text += data.TryGetProperty("text", out var t) ? t.GetString() : "";
else if (evt == "done") done = data;
}
}
var raw = done is { } d ? d.GetProperty("output").GetProperty("output").GetString() : text;
7. Parse the result
The reply is one JSON object as a string inside the envelope, so unwrap it twice. The model is told to
return no prose and no code fences; be defensive anyway - strip a leading fence, take the first
{ to the last }, and parse. Then route on lane.
What the web app's normalizer does to it
- Reads
lane; if it is notauditorrisk, decides from the body shape (abiasesarray means audit, alimitsarray means risk), else the lane asked for. - Coerces every enum to its allowed set with a default: an unknown verdict becomes the middle one
(
discounted/size-down), an unknown severitymedium, an unknown bias statusunknown, an unknown concernmedium. - Always renders exactly the eight biases in the fixed order; one the reply omitted shows as
unknownwith "not answered". - Assigns
BT-nnnids to findings that lack one; drops a finding with neither a problem nor a fix; sorts by severity for display. - Clamps
haircut.discount_pctto 0..100. - Throws (and the page retries once with a
retry_note) when an audit has neither findings nor a bias table, or a risk memo has neither metrics nor limits nor a summary.
Invariants worth asserting in CI
Audit
- A
criticalfinding forbidsverdict: "credible". - Two or more biases
confirmedforbidcredible. unreliablerequires at least onecriticalorhighfinding.- A
confirmedlook-ahead meansunreliable. haircut.discount_pct > 50is inconsistent withcredible.findings[i].id == "BT-%03d" % (i+1); the set ofcoverage_check[].idequals the set offacts.flags[].idyou sent, no duplicates, none invented.
Risk
facts.metrics.max_drawdown <= -0.40forbidsdeployable; so doesfacts.metrics.sharpe < 0.5.prior_audit.verdict == "unreliable"forcesnot-yet.- Empty
facts.metricsforcesnot-yetwithmetrics: []. - Every
metrics[].valuere-parses to the number you sent infacts.metrics(the web app compares each within 3 percent or 0.011, sign-insensitive, and badges a mismatch).
# Unwrap twice: the envelope, then the reply string.
call "jobs/$JOB" | python3 -c '
import sys, json
raw = json.load(sys.stdin)["data"]["output"]["output"].strip()
raw = raw[raw.find("{"): raw.rfind("}") + 1]
r = json.loads(raw)
print(r["lane"], r["verdict"], "-", r["verdict_reason"])
if r["lane"] == "audit":
for f in r["findings"]: print(f["id"], f["severity"], f["category"], "|", f["target"])
print("haircut", r["haircut"]["reported_sharpe"], "->", r["haircut"]["expected_sharpe"], r["haircut"]["discount_pct"], "%")
else:
for m in r["metrics"]: print(m["metric"], m["value"], m["concern"])
for l in r["limits"]: print("limit:", l["limit"], l["value"])
'
def parse_reply(raw):
t = raw.strip()
if t.startswith("```"):
t = t.split("\n", 1)[1].rsplit("```", 1)[0]
return json.loads(t[t.find("{"): t.rfind("}") + 1])
r = parse_reply(raw)
assert r["lane"] in ("audit", "risk")
if r["lane"] == "audit":
ids = [f["id"] for f in r["findings"]]
assert ids == [f"BT-{i+1:03d}" for i in range(len(ids))], ids
if any(f["severity"] == "critical" for f in r["findings"]):
assert r["verdict"] != "credible"
sent = {f["id"] for f in INPUT["facts"]["flags"]}
got = [c["id"] for c in r["coverage_check"]]
assert set(got) == sent and len(got) == len(sent), (sent, got)
else:
if INPUT.get("prior_audit", {}).get("verdict") == "unreliable":
assert r["verdict"] == "not-yet"
for m in r["metrics"]:
assert m["concern"] in ("low", "medium", "high")
function parseReply(raw) {
let t = raw.trim().replace(/^```[a-z]*\s*/i, "").replace(/```\s*$/, "");
return JSON.parse(t.slice(t.indexOf("{"), t.lastIndexOf("}") + 1));
}
const r = parseReply(raw);
if (r.lane === "audit") {
r.findings.forEach((f, i) => { if (f.id !== `BT-${String(i + 1).padStart(3, "0")}`) throw new Error("ids not sequential"); });
if (r.findings.some(f => f.severity === "critical") && r.verdict === "credible") throw new Error("credible with a critical finding");
const sent = new Set(INPUT.facts.flags.map(f => f.id)), got = r.coverage_check.map(c => c.id);
if (got.length !== sent.size || !got.every(id => sent.has(id))) throw new Error("coverage_check does not reconcile the sent flags");
} else {
if (INPUT.prior_audit?.verdict === "unreliable" && r.verdict !== "not-yet") throw new Error("risk ignored an unreliable audit");
}
type Finding struct {
ID, Severity, Category, Target, Problem, Impact, Fix, Snippet string
}
type Reply struct {
Lane, Verdict, VerdictReason string `json:"lane"`
Findings []Finding
Biases []struct{ Bias, Status, Evidence, Fix string }
Haircut struct {
ReportedSharpe string `json:"reported_sharpe"`
ExpectedSharpe string `json:"expected_sharpe"`
DiscountPct float64 `json:"discount_pct"`
}
CoverageCheck []struct{ ID string `json:"id"`; Addressed bool; Note string } `json:"coverage_check"`
Metrics []struct{ Metric, Value, Reading, Concern string }
Limits []struct{ Limit, Value, Rationale string }
}
func parseReply(raw string) (Reply, error) {
t := strings.TrimSpace(raw)
t = t[strings.Index(t, "{") : strings.LastIndex(t, "}")+1]
var r Reply
err := json.Unmarshal([]byte(t), &r)
return r, err
}
static JsonNode parseReply(String raw) throws Exception {
String t = raw.trim().replaceFirst("^```[a-z]*\\s*", "").replaceFirst("```\\s*$", "");
return JSON.readTree(t.substring(t.indexOf('{'), t.lastIndexOf('}') + 1));
}
var r = parseReply(raw);
if (r.get("lane").asText().equals("audit")) {
var findings = r.get("findings");
for (int i = 0; i < findings.size(); i++) {
if (!findings.get(i).get("id").asText().equals(String.format("BT-%03d", i + 1))) throw new IllegalStateException("ids not sequential");
if (findings.get(i).get("severity").asText().equals("critical") && r.get("verdict").asText().equals("credible"))
throw new IllegalStateException("credible with a critical finding");
}
}
def parse_reply(raw)
t = raw.strip.sub(/\A```[a-z]*\s*/i, "").sub(/```\s*\z/, "")
JSON.parse(t[t.index("{")..t.rindex("}")])
end
r = parse_reply(raw)
if r["lane"] == "audit"
r["findings"].each_with_index { |f, i| raise "ids" unless f["id"] == format("BT-%03d", i + 1) }
raise "credible with a critical finding" if r["findings"].any? { |f| f["severity"] == "critical" } && r["verdict"] == "credible"
sent = INPUT[:facts][:flags].map { |f| f[:id] }.sort
raise "coverage" unless r["coverage_check"].map { |c| c["id"] }.sort == sent
end
function parseReply(string $raw): array {
$t = trim($raw);
$t = preg_replace('/^```[a-z]*\s*/i', '', $t);
$t = preg_replace('/```\s*$/', '', $t);
return json_decode(substr($t, strpos($t, "{"), strrpos($t, "}") - strpos($t, "{") + 1), true);
}
$r = parseReply($raw);
if ($r["lane"] === "audit") {
foreach ($r["findings"] as $i => $f) {
if ($f["id"] !== sprintf("BT-%03d", $i + 1)) throw new RuntimeException("ids not sequential");
if ($f["severity"] === "critical" && $r["verdict"] === "credible") throw new RuntimeException("credible with a critical finding");
}
}
static JsonElement ParseReply(string raw) {
var t = raw.Trim();
t = Regex.Replace(t, @"^```[a-z]*\s*", "", RegexOptions.IgnoreCase);
t = Regex.Replace(t, @"```\s*$", "");
return JsonDocument.Parse(t[t.IndexOf('{')..(t.LastIndexOf('}') + 1)]).RootElement;
}
var r = ParseReply(raw);
if (r.GetProperty("lane").GetString() == "audit") {
var i = 0;
foreach (var f in r.GetProperty("findings").EnumerateArray()) {
if (f.GetProperty("id").GetString() != $"BT-{++i:000}") throw new Exception("ids not sequential");
if (f.GetProperty("severity").GetString() == "critical" && r.GetProperty("verdict").GetString() == "credible")
throw new Exception("credible with a critical finding");
}
}
The output contract
One JSON object. The common envelope, then the body for the lane answered. Every key is present; "" or [] rather than an omission.
{
"lane": "audit" | "risk",
"title": "12-1 momentum, top quintile, S&P 500, 2016-2024",
"verdict": "credible" | "discounted" | "unreliable" // audit
| "deployable" | "size-down" | "not-yet", // risk
"verdict_reason": "one sentence naming what decided it",
"exec_summary": "two to five sentences",
"assumptions": ["..."], "open_questions": ["..."],
"coverage_check": [{ "id": "LA-NEG-SHIFT", "addressed": true, "note": "..." }], // one per facts.flags id
"quick_wins": ["..."], "summary": "one paragraph",
// audit body
"biases": [ { "bias": "look-ahead", "status": "clear|suspected|confirmed|unknown", "evidence": "...", "fix": "..." }, ... 8 in fixed order ],
"findings": [ { "id": "BT-001", "severity": "critical|high|medium|low", "category": "<a bias>|other",
"target": "momentum.py line 22, signal = signal.shift(-1)", "problem": "...", "impact": "...", "fix": "...", "snippet": "..." } ],
"haircut": { "reported_sharpe": "2.4", "expected_sharpe": "0.0 to 0.5", "discount_pct": 90, "reasoning": "..." },
"validation_plan": [ { "step": "...", "why": "...", "how": "..." } ],
"questions_for_author": ["..."],
// risk body
"metrics": [ { "metric": "Sharpe", "value": "0.59", "reading": "...", "concern": "low|medium|high" } ],
"tail_risk": { "var_95": "1.93% per trade", "cvar_95": "2.45% per trade", "worst_period": "-4.1%", "reading": "..." },
"drawdowns": [ { "rank": 1, "depth": "-11.74%", "from": "2022-02-14", "to": "not recovered", "recovery": "98 trades and counting", "reading": "..." } ],
"limits": [ { "limit": "Daily VaR 95", "value": "0.6% of the book", "rationale": "..." } ],
"sizing": { "vol_target": "6% annualised", "kelly_fraction": "0.25", "max_leverage": "1.5x", "reasoning": "..." },
"stress": [ { "scenario": "2022 rates repricing repeat", "expected_loss": "-9% at target vol", "reasoning": "..." } ],
"monitoring": [ { "signal": "Rolling 12m Sharpe", "threshold": "below 0.3", "action": "halve size and review" } ]
}
The enums
| field | values |
|---|---|
lane | audit, risk |
verdict (audit) | credible, discounted, unreliable |
verdict (risk) | deployable, size-down, not-yet |
biases[].status | clear, suspected, confirmed, unknown |
findings[].severity | critical, high, medium, low |
findings[].category | one of the eight bias names, or other |
metrics[].concern | low, medium, high |
series_kind (input) | equity, returns, trades, none |
frequency (input) | daily, weekly, monthly, quarterly, hourly, hourly-24, trades, unknown |
The eight biases, in order
look-ahead, survivorship, overfitting, selection,
transaction costs, data quality, execution realism, capacity.
Always eight, always this order, always these exact strings.
The facts.metrics keys
Returns and drawdowns are decimals (0.124 is 12.4 percent, -0.31 a 31 percent drawdown); VaR and CVaR
are positive losses per period. Send what you have; every key is optional and the contract says
"not computable from the paste" for the rest.
obs, years, periods_per_year, frequency, kind, start, end,
total_return, cagr, ann_return, ann_vol, sharpe, sortino, rf_rate,
max_drawdown, max_dd_from, max_dd_trough, max_dd_to, max_dd_periods, longest_underwater, calmar,
var_95, cvar_95, var_99, cvar_99, skew, kurtosis,
win_rate, best, worst, avg_win, avg_loss, profit_factor, tail_ratio,
psr, min_track_record, rolling_sharpe_min, rolling_sharpe_max,
months, positive_months, worst_month, best_month,
trades, trades_per_year, cost_bps, cost_drag, sharpe_net,
drawdowns: [{ rank, depth, from, trough, to, periods }]
The flag ids the web app's lint can send
code: LA-NEG-SHIFT LA-BFILL LA-CENTER LA-FULL-SCALE LA-NO-LAG LA-PINE-LOOKAHEAD LA-ILOC-NEXT LA-COC EX-SAME-BAR
SV-CURRENT-UNIVERSE SV-DROPNA TC-NO-COSTS TC-ZERO OF-INSAMPLE-OPT OF-NO-OOS OF-MANY-PARAMS DQ-NO-SEED DQ-ADJ-LEVEL
series: SR-NONE SR-SHORT SR-TOO-GOOD SR-NO-LOSSES SR-DD-TINY SR-PSR-LOW SR-TRACK-SHORT SR-FAT-TAILS SR-REGIME
SR-BAD-DATES SR-UNSORTED SR-DUPES SR-NONPOSITIVE
claims: CL-SHARPE CL-DRAWDOWN CL-CAGR CL-WINRATE
You may send your own ids too; the contract only requires that every id you send comes back exactly once in coverage_check.
8. The risk memo, carrying the audit
The second lane on the same backtest. Copy the audit's verdict, haircut and a few
top findings into prior_audit: the memo then sizes against the haircut Sharpe rather than the
reported one, and an unreliable audit forces not-yet. Everything else in the input is
the same object with task switched to risk. This example is the second bundled
backtest - a walk-forward futures carry book, 240 trades, real costs, a 2022 drawdown.
RISK='{"task":"risk","strategy":"# file: carry_wf.py\n\"\"\"Futures carry, long high roll yield, short low, weekly rebalance. Walk-forward: 3-year train, 6-month test, 2015-2024. Commission 2.50 per contract per side plus one tick of slippage.\"\"\"\nimport backtrader as bt\nnp.random.seed(7)\nclass Carry(bt.Strategy):\n params = dict(lookback=20, legs=2)\n def next(self):\n self.order_target_percent(d, target) # fills at the next open","series_kind":"trades","series_excerpt":"entry_date,exit_date,symbol,pnl\n2018-01-02,2018-01-08,FGBL,-7217\n2018-01-15,2018-01-23,ZF,-19031\n... 228 rows profiled in the browser and not sent ...\n2024-12-09,2024-12-16,6E,4120","frequency":"trades","claims":"Sharpe about 0.6 before costs, max drawdown about 12% in 2022, roughly 35 trades a year, 2018-2024 out-of-sample.","context":"Rates and FX futures on CME and Eurex, a 1,000,000 book, mid-day execution. We could size up to 5,000,000 if the risk memo supports it.","params":{"rf_rate":0.02,"cost_bps":3,"capital":1000000},"facts":{"metrics":{"obs":240,"years":6.94,"periods_per_year":34.6,"kind":"trades","sharpe":0.59,"sortino":0.9,"cagr":0.0544,"ann_vol":0.0586,"max_drawdown":-0.1174,"max_dd_from":"2022-02-14","max_dd_to":"not recovered","longest_underwater":98,"calmar":0.46,"var_95":0.0193,"cvar_95":0.0245,"cvar_99":0.0312,"skew":0.05,"kurtosis":1.02,"win_rate":0.55,"profit_factor":1.13,"psr":0.939,"min_track_record":272,"trades_per_year":34.6,"cost_drag":0.0208,"sharpe_net":0.24,"drawdowns":[{"rank":1,"depth":-0.1174,"from":"2022-02-14","trough":"2023-11-06","to":"not recovered","periods":98}]},"flags":[{"id":"SR-PSR-LOW","label":"probabilistic Sharpe ratio 0.939 - the sample cannot distinguish this Sharpe from zero at 95%"},{"id":"SR-TRACK-SHORT","label":"240 observations against a minimum track record of 272"}],"resources":[{"id":"RES-CODE","label":"python code, 33 lines"},{"id":"RES-SERIES","label":"240 trades over 6.9 years"}],"claims_parsed":{"sharpe":0.6,"max_drawdown":-0.12,"period":"2018-2024"}},"prior_audit":{"verdict":"discounted","haircut":{"reported_sharpe":"0.59","expected_sharpe":"0.25 to 0.45","discount_pct":42,"reasoning":"..."},"top_findings":["BT-001 (high, overfitting): the grid search inside each train window is unreported"]}}'
KEY2="backtest-desk:risk:$(printf '%s' "$RISK" | shasum -a 256 | cut -c1-16):a1"
curl -sS -X POST "$BASE/run" -H "Authorization: Bearer $SKILLSAFE_TOKEN" -H "Content-Type: application/json" \
-H "Idempotency-Key: $KEY2" --data-binary "$RISK"
# poll jobs/{job_id} as in step 5; expect lane "risk", verdict "size-down" or "not-yet"
audit = parse_reply(raw) # from step 7
top = sorted(audit["findings"], key=lambda f: ["critical", "high", "medium", "low"].index(f["severity"]))[:5]
RISK = {**INPUT, "task": "risk",
"prior_audit": {"verdict": audit["verdict"], "haircut": audit["haircut"],
"top_findings": [f"{f['id']} ({f['severity']}, {f['category']}): {f['problem']}" for f in top]}}
job = call("run", RISK, headers={"Idempotency-Key": idem_key(RISK)})
# poll as in step 5, then:
memo = parse_reply(j["output"]["output"])
assert memo["lane"] == "risk"
if audit["verdict"] == "unreliable":
assert memo["verdict"] == "not-yet"
for row in memo["metrics"]:
print(f"{row['metric']:28s} {row['value']:>16s} {row['concern']}")
print(memo["sizing"]["vol_target"], memo["sizing"]["kelly_fraction"], memo["sizing"]["max_leverage"])
const audit = parseReply(raw); // from step 7
const order = ["critical", "high", "medium", "low"];
const top = [...audit.findings].sort((a, b) => order.indexOf(a.severity) - order.indexOf(b.severity)).slice(0, 5);
const RISK = { ...INPUT, task: "risk",
prior_audit: { verdict: audit.verdict, haircut: audit.haircut,
top_findings: top.map(f => `${f.id} (${f.severity}, ${f.category}): ${f.problem}`) } };
const riskJob = await call("run", RISK, { "Idempotency-Key": idemKey(RISK) });
// poll as in step 5, then:
const memo = parseReply(riskDone.output.output);
if (audit.verdict === "unreliable" && memo.verdict !== "not-yet") throw new Error("memo ignored an unreliable audit");
for (const row of memo.metrics) console.log(row.metric, row.value, row.concern);
console.log(memo.sizing.vol_target, memo.sizing.kelly_fraction, memo.sizing.max_leverage);
audit, _ := parseReply(raw) // from step 7
top := []string{}
for _, f := range audit.Findings[:min(5, len(audit.Findings))] {
top = append(top, fmt.Sprintf("%s (%s, %s): %s", f.ID, f.Severity, f.Category, f.Problem))
}
risk := map[string]any{}
for k, v := range input {
risk[k] = v
}
risk["task"] = "risk"
risk["prior_audit"] = map[string]any{"verdict": audit.Verdict, "haircut": audit.Haircut, "top_findings": top}
rb, _ := json.Marshal(risk)
rsum := sha256.Sum256(rb)
raw, err = call(token, "run", risk, map[string]string{"Idempotency-Key": fmt.Sprintf("backtest-desk:risk:%x:a1", rsum[:8])})
// poll jobs/{job_id} as in step 5; then parseReply and read Metrics, Limits
var audit = parseReply(raw); // from step 7
var top = new ArrayList<String>();
for (var f : audit.get("findings")) {
if (top.size() == 5) break;
top.add(f.get("id").asText() + " (" + f.get("severity").asText() + ", " + f.get("category").asText() + "): " + f.get("problem").asText());
}
Map<String, Object> risk = new LinkedHashMap<>(input);
risk.put("task", "risk");
risk.put("prior_audit", Map.of("verdict", audit.get("verdict").asText(), "haircut", JSON.convertValue(audit.get("haircut"), Map.class), "top_findings", top));
String rbody = JSON.writeValueAsString(risk);
String rkey = "backtest-desk:risk:" + HexFormat.of().formatHex(MessageDigest.getInstance("SHA-256").digest(rbody.getBytes())).substring(0, 16) + ":a1";
var riskJob = call(token, "run", risk, Map.of("Idempotency-Key", rkey));
// poll as in step 5, then parseReply(...) and read "metrics", "limits", "sizing"
audit = parse_reply(raw) # from step 7
order = %w[critical high medium low]
top = audit["findings"].sort_by { |f| order.index(f["severity"]) }.first(5)
.map { |f| "#{f['id']} (#{f['severity']}, #{f['category']}): #{f['problem']}" }
risk = INPUT.merge(task: "risk", prior_audit: { verdict: audit["verdict"], haircut: audit["haircut"], top_findings: top })
rkey = "backtest-desk:risk:#{Digest::SHA256.hexdigest(risk.to_json)[0, 16]}:a1"
jobRisk = call("run", risk, "Idempotency-Key" => rkey)
# poll as in step 5, then:
memo = parse_reply(@j["output"]["output"])
raise "memo ignored an unreliable audit" if audit["verdict"] == "unreliable" && memo["verdict"] != "not-yet"
memo["metrics"].each { |m| puts "#{m['metric']} #{m['value']} #{m['concern']}" }
$audit = parseReply($raw); // from step 7
$order = ["critical", "high", "medium", "low"];
usort($audit["findings"], fn($a, $b) => array_search($a["severity"], $order) <=> array_search($b["severity"], $order));
$top = array_map(fn($f) => "{$f['id']} ({$f['severity']}, {$f['category']}): {$f['problem']}", array_slice($audit["findings"], 0, 5));
$risk = array_merge($input, ["task" => "risk", "prior_audit" => ["verdict" => $audit["verdict"], "haircut" => $audit["haircut"], "top_findings" => $top]]);
$rkey = "backtest-desk:risk:" . substr(hash("sha256", json_encode($risk)), 0, 16) . ":a1";
$riskJob = call($token, "run", $risk, ["Idempotency-Key" => $rkey]);
// poll as in step 5, then:
$memo = parseReply($j["output"]["output"]);
if ($audit["verdict"] === "unreliable" && $memo["verdict"] !== "not-yet") throw new RuntimeException("memo ignored an unreliable audit");
foreach ($memo["metrics"] as $m) echo $m["metric"], " ", $m["value"], " ", $m["concern"], "\n";
var audit = ParseReply(raw); // from step 7
var order = new[] { "critical", "high", "medium", "low" };
var top = audit.GetProperty("findings").EnumerateArray()
.OrderBy(f => Array.IndexOf(order, f.GetProperty("severity").GetString())).Take(5)
.Select(f => $"{f.GetProperty("id")} ({f.GetProperty("severity")}, {f.GetProperty("category")}): {f.GetProperty("problem")}").ToList();
var risk = new Dictionary<string, object>(input) { ["task"] = "risk",
["prior_audit"] = new { verdict = audit.GetProperty("verdict").GetString(), haircut = audit.GetProperty("haircut"), top_findings = top } };
var rbody = JsonSerializer.Serialize(risk);
var rkey = $"backtest-desk:risk:{Convert.ToHexString(SHA256.HashData(Encoding.UTF8.GetBytes(rbody)))[..16].ToLower()}:a1";
var riskJob = await Call(token, "run", risk, new() { ["Idempotency-Key"] = rkey });
// poll as in step 5, then ParseReply(...) and read "metrics", "limits", "sizing"
9. Use it in a research pipeline
The audit verdict is a gate. A shell step that fails the pipeline on unreliable or on any
critical finding is the smallest useful integration; a stricter shop also fails on a
discount above 50 percent, and generates the risk memo only for what passes.
#!/bin/sh
# gate.sh - fail the pipeline when the audit says the backtest cannot be believed.
set -eu
: "${SKILLSAFE_TOKEN:?set SKILLSAFE_TOKEN from the token page}"
INPUT=$(python3 build_input.py strategies/momentum.py results/equity.csv) # your metrics + lint -> the input object
KEY="backtest-desk:audit:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"
JOB=$(curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/run" -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
-H "Content-Type: application/json" -H "Idempotency-Key: $KEY" --data-binary "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')
while :; do
J=$(curl -sS "https://api.skillsafe.ai/v1/app-api/jobs/$JOB" -H "Authorization: Bearer $SKILLSAFE_TOKEN")
S=$(printf '%s' "$J" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
[ "$S" = succeeded ] && break
[ "$S" = failed ] && { echo "run failed"; exit 2; }
sleep 3
done
printf '%s' "$J" | python3 -c '
import sys, json
raw = json.load(sys.stdin)["data"]["output"]["output"]
r = json.loads(raw[raw.find("{"): raw.rfind("}") + 1])
crit = [f["id"] for f in r["findings"] if f["severity"] == "critical"]
print(r["verdict"], "-", r["verdict_reason"])
for f in r["findings"]: print(" ", f["id"], f["severity"], f["target"])
bad = r["verdict"] == "unreliable" or crit or r["haircut"]["discount_pct"] > 50
sys.exit(1 if bad else 0)
'
# gate.py - exit 1 when the audit says the backtest cannot be believed.
import sys
r = parse_reply(raw) # steps 5 and 7
crit = [f["id"] for f in r["findings"] if f["severity"] == "critical"]
print(r["verdict"], "-", r["verdict_reason"])
for f in r["findings"]:
print(f" {f['id']} {f['severity']:8s} {f['target']}")
if r["verdict"] == "unreliable" or crit or r["haircut"]["discount_pct"] > 50:
print("blocked: unreliable, a critical finding, or a haircut over 50%")
sys.exit(1)
# passed the gate: now write the risk memo (step 8) and attach it to the promotion ticket
// gate.mjs - exit 1 when the audit says the backtest cannot be believed.
const r = parseReply(raw); // steps 5 and 7
const crit = r.findings.filter(f => f.severity === "critical").map(f => f.id);
console.log(r.verdict, "-", r.verdict_reason);
for (const f of r.findings) console.log(" ", f.id, f.severity, f.target);
if (r.verdict === "unreliable" || crit.length || r.haircut.discount_pct > 50) {
console.error("blocked: unreliable, a critical finding, or a haircut over 50%");
process.exit(1);
}
// gate: os.Exit(1) when the audit says the backtest cannot be believed.
r, err := parseReply(raw)
if err != nil {
log.Fatal(err)
}
crit := 0
for _, f := range r.Findings {
if f.Severity == "critical" {
crit++
}
fmt.Printf(" %s %-8s %s\n", f.ID, f.Severity, f.Target)
}
if r.Verdict == "unreliable" || crit > 0 || r.Haircut.DiscountPct > 50 {
fmt.Println("blocked: unreliable, a critical finding, or a haircut over 50%")
os.Exit(1)
}
// gate: System.exit(1) when the audit says the backtest cannot be believed.
var r = parseReply(raw);
boolean crit = false;
for (var f : r.get("findings")) {
crit |= f.get("severity").asText().equals("critical");
System.out.printf(" %s %-8s %s%n", f.get("id").asText(), f.get("severity").asText(), f.get("target").asText());
}
if (r.get("verdict").asText().equals("unreliable") || crit || r.at("/haircut/discount_pct").asDouble() > 50) {
System.err.println("blocked: unreliable, a critical finding, or a haircut over 50%");
System.exit(1);
}
# gate.rb - exit 1 when the audit says the backtest cannot be believed.
r = parse_reply(raw)
crit = r["findings"].select { |f| f["severity"] == "critical" }
puts "#{r['verdict']} - #{r['verdict_reason']}"
r["findings"].each { |f| puts " #{f['id']} #{f['severity'].ljust(8)} #{f['target']}" }
if r["verdict"] == "unreliable" || crit.any? || r["haircut"]["discount_pct"] > 50
warn "blocked: unreliable, a critical finding, or a haircut over 50%"
exit 1
end
// gate.php - exit(1) when the audit says the backtest cannot be believed.
$r = parseReply($raw);
$crit = array_filter($r["findings"], fn($f) => $f["severity"] === "critical");
echo $r["verdict"], " - ", $r["verdict_reason"], "\n";
foreach ($r["findings"] as $f) printf(" %s %-8s %s\n", $f["id"], $f["severity"], $f["target"]);
if ($r["verdict"] === "unreliable" || $crit || $r["haircut"]["discount_pct"] > 50) {
fwrite(STDERR, "blocked: unreliable, a critical finding, or a haircut over 50%\n");
exit(1);
}
// gate: Environment.Exit(1) when the audit says the backtest cannot be believed.
var r = ParseReply(raw);
var crit = false;
foreach (var f in r.GetProperty("findings").EnumerateArray()) {
crit |= f.GetProperty("severity").GetString() == "critical";
Console.WriteLine($" {f.GetProperty("id")} {f.GetProperty("severity"),-8} {f.GetProperty("target")}");
}
if (r.GetProperty("verdict").GetString() == "unreliable" || crit || r.GetProperty("haircut").GetProperty("discount_pct").GetDouble() > 50) {
Console.Error.WriteLine("blocked: unreliable, a critical finding, or a haircut over 50%");
Environment.Exit(1);
}
Truncation and partial results
If the balance sits between min_credits and hold_credits, the run still executes with
a reduced output cap and the finished job carries truncated: true. The reply is then cut
mid-JSON. Do what the web app does: try to close the object (append "}]}-style tails until
one parses), keep every section that arrived, count them against the lane's section list
(title, verdict, exec_summary, biases, findings, haircut, validation_plan, questions_for_author,
coverage_check, summary for an audit; title, verdict, exec_summary, metrics, tail_risk,
drawdowns, limits, sizing, stress, monitoring, coverage_check, summary for a memo), report
"N of M sections", and top up before re-running with the same Idempotency-Key and the attempt counter
incremented.
Backtest Desk reads a backtest and a series; it never runs code, fetches prices, reaches a broker or recommends a trade. Nothing it returns is investment advice, and past performance - simulated or real - does not predict future results. Derived from @wshobson/backtesting-frameworks and @wshobson/risk-metrics-calculation (wshobson/agents, MIT).