← EDA Desk / API
Tokens

Drive EDA Desk from your own code

EDA Desk checks a tabular dataset before modeling. In the page, JavaScript ports of the three CSV/TSV tools of the agent skill @k-dense-ai/exploratory-data-analysis (k-dense-ai/scientific-agent-skills) run free: tabular_profile.py (kinds, missingness, distinct counts, numeric aggregates), missingness_leakage_audit.py (missingness by split and group; entities, groups, identical rows and time ranges in more than one split) and distribution_sensitivity.py (mean vs median, SD vs MAD, IQR fences, trimmed and winsorized means, log1p skewness). From code, run the same scripts with Python (python scripts/missingness_leakage_audit.py data.csv --root . --split-column split --entity-column subject_id) and build facts from their JSON.

The metered lanes read only aggregates: split judges whether the split supports an honest held-out score and returns a resplit plan with scikit-learn code; report drafts the EDA report the skill's template asks for. Neither ever sees a row, a cell value or an entity id.

Two lanes: the task field

task picks the lane. /estimate does not validate the body, so always send a JSON object with task set to one of the two lanes.

taskneedsreturns
splitfacts with a split role; context and question optionalstatus (resplit_required, usable_with_fixes, usable_as_is), a response to every flag, leak_findings, a split_plan (unit, method, hold-out, steps, rationale), preprocessing_rules, code, questions and assumptions.
reportfacts; context, question and split_review optionalstatus (proceed_to_modeling, proceed_with_caveats, do_not_model_yet), flag responses, sections (scope, structure, split, schema, missingness, distributions, transformations, limitations), key_findings, data_dictionary_requests, questions and assumptions.

Worked example: split

The clinic example on the page: visit rows split 75/25 at random, so the same subject appears in both splits. Column ids are tokens.

{
 "task": "split",
 "facts": "{\"page\":\"eda-desk\",\"tools\":\"tabular_profile.py, missingness_leakage_audit.py, distribution_sensitivity.py (exploratory-data-analysis skill, run in the browser on the user's file)\",\"identifiers\":\"tokenized (column_*, split_*, group_* pseudonyms); the user sees names, you see tokens\",\"file\":{\"format\":\"CSV\",\"rows_scanned\":176,\"row_limit_reached\":false,\"columns\":11,\"duplicate_rows\":0,\"missing_codes_declared\":1},\"roles\":{\"split\":\"column_4a334e7217932c40\",\"entity\":\"column_ce22b27fac9ca72c\",\"group\":\"column_86523fcfc53231f7\",\"time\":\"column_bfa48dcb85289118\",\"target\":\"column_73c6fa731c23780b\"},\"columns\":[{\"id\":\"column_ce22b27fac9ca72c\",\"role\":\"entity\",\"kind\":\"text\",\"missing_pct\":0,\"distinct\":56},{\"id\":\"column_bfa48dcb85289118\",\"role\":\"time\",\"kind\":\"text\",\"missing_pct\":0,\"distinct\":153},{\"id\":\"column_86523fcfc53231f7\",\"role\":\"group\",\"kind\":\"text\",\"missing_pct\":0,\"distinct\":3},{\"id\":\"column_4a334e7217932c40\",\"role\":\"split\",\"kind\":\"text\",\"missing_pct\":0,\"distinct\":2},{\"id\":\"column_73c6fa731c23780b\",\"role\":\"target\",\"kind\":\"integer\",\"missing_pct\":0,\"distinct\":2,\"mean\":0.2273,\"sd\":0.4203,\"min\":0,\"median\":0,\"max\":1,\"mad\":0,\"iqr\":0,\"trimmed_mean\":0.162,\"outside_iqr_fences\":\"40/176\",\"skew\":1.302,\"log1p_skew\":1.302},{\"id\":\"column_b4c12f0ab59f1f7f\",\"kind\":\"numeric\",\"missing_pct\":0,\"distinct\":90,\"mean\":6.355,\"sd\":8.974,\"min\":0.4,\"median\":3.8,\"max\":71.6,\"mad\":2.3,\"iqr\":5.75,\"trimmed_mean\":4.506,\"outside_iqr_fences\":\"13/176\",\"skew\":4.221,\"log1p_skew\":0.7836},{\"id\":\"column_2a39cbc6e4c57f3d\",\"kind\":\"mixed\",\"missing_pct\":0,\"distinct\":131,\"mean\":3.136,\"sd\":0.7404,\"min\":1.3,\"median\":3.14,\"max\":5.18,\"numeric_values\":175,\"text_values\":1,\"mad\":0.51,\"iqr\":1,\"trimmed_mean\":3.136,\"outside_iqr_fences\":\"1/175\",\"skew\":0.0182,\"log1p_skew\":-0.4604},{\"id\":\"column_974ca6e12d61c806\",\"kind\":\"numeric\",\"missing_pct\":8.5,\"distinct\":45,\"mean\":6.329,\"sd\":0.9731,\"min\":4.2,\"median\":6.2,\"max\":8.8,\"mad\":0.6,\"iqr\":1.3,\"trimmed_mean\":6.292,\"outside_iqr_fences\":\"0/161\",\"skew\":0.3328,\"log1p_skew\":0.01885},{\"id\":\"column_b315a92d8b82a14c\",\"kind\":\"integer\",\"missing_pct\":0,\"distinct\":34,\"mean\":56.16,\"sd\":13.26,\"min\":34,\"median\":59,\"max\":82,\"mad\":10,\"iqr\":23.25,\"trimmed_mean\":56.04,\"outside_iqr_fences\":\"0/176\",\"skew\":-0.00217,\"log1p_skew\":-0.3196},{\"id\":\"column_dbc0491d72d586d9\",\"kind\":\"text\",\"missing_pct\":0,\"distinct\":2},{\"id\":\"column_c50d01f329dff09d\",\"kind\":\"numeric\",\"missing_pct\":5.1,\"distinct\":98,\"mean\":28.28,\"sd\":3.981,\"min\":16.7,\"median\":28.2,\"max\":37,\"mad\":2.4,\"iqr\":5.25,\"trimmed_mean\":28.36,\"outside_iqr_fences\":\"2/167\",\"skew\":-0.1781,\"log1p_skew\":-0.5898}],\"columns_omitted\":0,\"tool_exits\":{\"profile\":0,\"leakage\":0,\"distribution\":0},\"flags\":[{\"id\":\"F1\",\"kind\":\"leak_entity\",\"severity\":\"high\",\"text\":\"31 entity value(s) of column_ce22b27fac9ca72c appear in more than one split.\"},{\"id\":\"F2\",\"kind\":\"leak_group\",\"severity\":\"high\",\"text\":\"3 group value(s) of column_86523fcfc53231f7 appear in more than one split.\"},{\"id\":\"F3\",\"kind\":\"leak_duplicate_rows\",\"severity\":\"high\",\"text\":\"2 row(s) identical apart from the split column appear in more than one split.\"},{\"id\":\"F4\",\"kind\":\"leak_time\",\"severity\":\"high\",\"text\":\"1 pair(s) of splits overlap in column_bfa48dcb85289118 (min-max intervals).\"},{\"id\":\"F5\",\"kind\":\"missing_split_gap\",\"severity\":\"medium\",\"text\":\"column_974ca6e12d61c806 is missing in 3.1%-24.4% of rows depending on the split (a gap of 21.4 points).\"},{\"id\":\"F6\",\"kind\":\"mixed_kind\",\"severity\":\"medium\",\"text\":\"column_2a39cbc6e4c57f3d mixes 175 numeric and 1 text value(s) - a detection-limit marker, an undeclared missing code or a unit suffix is the usual cause.\"},{\"id\":\"F7\",\"kind\":\"outliers\",\"severity\":\"low\",\"text\":\"column_b4c12f0ab59f1f7f: 13 of 176 sampled values sit outside the IQR fences; the raw mean is 6.355 against a 10% trimmed mean of 4.506.\"},{\"id\":\"F8\",\"kind\":\"skewed\",\"severity\":\"low\",\"text\":\"column_b4c12f0ab59f1f7f is skewed (moment skewness 4.221); log1p would bring it to 0.7836. A diagnostic only - any transform is fitted on training data.\"}],\"browser_status\":\"leakage_flagged\",\"leakage\":{\"status\":\"potential_leakage_detected\",\"entity_tokens_in_multiple_splits\":31,\"group_tokens_in_multiple_splits\":3,\"identical_row_hashes_in_multiple_splits\":2,\"overlapping_split_time_interval_pairs\":1,\"rows_with_missing_split\":0,\"unparseable_nonmissing_time_values\":0,\"tracking_truncated\":false,\"splits\":[{\"split\":\"split_2aefc8423ce0c051\",\"rows\":45,\"missing\":[\"column_974ca6e12d61c806:24.4%\",\"column_c50d01f329dff09d:8.9%\"]},{\"split\":\"split_373650b82b95498a\",\"rows\":131,\"missing\":[\"column_974ca6e12d61c806:3.1%\",\"column_c50d01f329dff09d:3.8%\"]}],\"groups\":3,\"group_gaps\":[]}}",
 "context": "Visit-level records from three outpatient sites. We want to predict 30-day readmission (readmit_30d) at a new visit. The train/test column was made with a random 75/25 split of rows."
}

The reply's status is resplit_required, with split_plan.unit set to the entity column's token and a GroupShuffleSplit in code. Load the example on the page to see a saved reply in full, free.

Worked example: report

The bioreactor example: runs split forward in time by weekly batch, with --reveal-identifiers so column names are sent (split and group values stay tokens).

{
 "task": "report",
 "facts": "{\"page\":\"eda-desk\",\"tools\":\"tabular_profile.py, missingness_leakage_audit.py, distribution_sensitivity.py (exploratory-data-analysis skill, run in the browser on the user's file)\",\"identifiers\":\"sanitized column names (--reveal-identifiers); split, group and entity values stay tokenized\",\"file\":{\"format\":\"TSV\",\"rows_scanned\":146,\"row_limit_reached\":false,\"columns\":10,\"duplicate_rows\":0,\"missing_codes_declared\":1},\"roles\":{\"split\":\"split\",\"entity\":\"run_id\",\"group\":\"batch\",\"time\":\"run_start\",\"target\":\"titer_g_l\"},\"columns\":[{\"id\":\"run_id\",\"role\":\"entity\",\"kind\":\"text\",\"missing_pct\":0,\"distinct\":146},{\"id\":\"batch\",\"role\":\"group\",\"kind\":\"text\",\"missing_pct\":0,\"distinct\":24},{\"id\":\"run_start\",\"role\":\"time\",\"kind\":\"text\",\"missing_pct\":0,\"distinct\":146},{\"id\":\"split\",\"role\":\"split\",\"kind\":\"text\",\"missing_pct\":0,\"distinct\":3},{\"id\":\"titer_g_l\",\"role\":\"target\",\"kind\":\"numeric\",\"missing_pct\":3.4,\"distinct\":103,\"mean\":1.964,\"sd\":0.7443,\"min\":0.69,\"median\":1.81,\"max\":5.38,\"mad\":0.35,\"iqr\":0.72,\"trimmed_mean\":1.852,\"outside_iqr_fences\":\"8/141\",\"skew\":1.835,\"log1p_skew\":0.9429},{\"id\":\"feed_rate_ml_h\",\"kind\":\"numeric\",\"missing_pct\":0,\"distinct\":104,\"mean\":12.39,\"sd\":8.507,\"min\":3.1,\"median\":10.4,\"max\":60.7,\"mad\":3.4,\"iqr\":6.925,\"trimmed_mean\":10.84,\"outside_iqr_fences\":\"12/146\",\"skew\":2.436,\"log1p_skew\":0.6074},{\"id\":\"operator\",\"kind\":\"text\",\"missing_pct\":0,\"distinct\":1},{\"id\":\"reactor\",\"kind\":\"text\",\"missing_pct\":0,\"distinct\":3},{\"id\":\"temp_c\",\"kind\":\"numeric\",\"missing_pct\":0,\"distinct\":84,\"mean\":36.85,\"sd\":0.352,\"min\":35.84,\"median\":36.8,\"max\":37.68,\"mad\":0.22,\"iqr\":0.4675,\"trimmed_mean\":36.85,\"outside_iqr_fences\":\"1/146\",\"skew\":-0.06966,\"log1p_skew\":-0.09774},{\"id\":\"ph\",\"kind\":\"numeric\",\"missing_pct\":0,\"distinct\":31,\"mean\":7.044,\"sd\":0.06326,\"min\":6.88,\"median\":7.045,\"max\":7.24,\"mad\":0.04,\"iqr\":0.08,\"trimmed_mean\":7.044,\"outside_iqr_fences\":\"1/146\",\"skew\":0.03359,\"log1p_skew\":0.006589}],\"columns_omitted\":0,\"tool_exits\":{\"profile\":0,\"leakage\":0,\"distribution\":0},\"flags\":[{\"id\":\"F1\",\"kind\":\"constant\",\"severity\":\"low\",\"text\":\"'operator' has a single distinct value.\"},{\"id\":\"F2\",\"kind\":\"outliers\",\"severity\":\"low\",\"text\":\"'feed_rate_ml_h': 12 of 146 sampled values sit outside the IQR fences; the raw mean is 12.39 against a 10% trimmed mean of 10.84.\"},{\"id\":\"F3\",\"kind\":\"skewed\",\"severity\":\"low\",\"text\":\"'feed_rate_ml_h' is skewed (moment skewness 2.436); log1p would bring it to 0.6074. A diagnostic only - any transform is fitted on training data.\"},{\"id\":\"F4\",\"kind\":\"outliers\",\"severity\":\"low\",\"text\":\"'titer_g_l': 8 of 141 sampled values sit outside the IQR fences; the raw mean is 1.964 against a 10% trimmed mean of 1.852.\"}],\"browser_status\":\"review_before_modeling\",\"leakage\":{\"status\":\"not_detected_in_scanned_rows\",\"entity_tokens_in_multiple_splits\":0,\"group_tokens_in_multiple_splits\":0,\"identical_row_hashes_in_multiple_splits\":0,\"overlapping_split_time_interval_pairs\":0,\"rows_with_missing_split\":0,\"unparseable_nonmissing_time_values\":0,\"tracking_truncated\":false,\"splits\":[{\"split\":\"split_2aefc8423ce0c051\",\"rows\":23,\"missing\":[]},{\"split\":\"split_373650b82b95498a\",\"rows\":100,\"missing\":[\"titer_g_l:3%\"]},{\"split\":\"split_c3544993fefff43f\",\"rows\":23,\"missing\":[\"titer_g_l:8.7%\"]}],\"groups\":24,\"group_gaps\":[\"titer_g_l:16.7 points\"]}}",
 "context": "One row per bioreactor run. Batches of media are used for one week each. Target is titer_g_l. We split forward in time by batch: the first 16 weekly batches train, the next 4 validate, the last 4 test."
}

Input fields

Every field is a string; facts is JSON text.

fieldtyperequiredmeaning
taskstringyes"split" or "report".
factsstringyesThe aggregates, as JSON text (below).
contextstringnoWhat the data is, the observational unit, how the split was made and what will be predicted. Sent as written, up to about 4,000 characters.
questionstringnoA question to answer in the reply.
split_reviewstringnoReport lane: an earlier split review as text (the page fills it from the review's status, findings and plan).
retry_notestringnoOnly on a reformat retry.

The facts string

file (format, rows_scanned, row_limit_reached, columns, duplicate_rows, missing_codes_declared); roles (split, entity, group, time, target: a column id or null); columns (per column: id, role, kind, missing_pct, distinct, mean, sd, min, median, max, mad, iqr, trimmed_mean, outside_iqr_fences as "count/sampled", skew and the log1p or signed_log1p skew, 4 significant digits); columns_omitted; leakage (the audit's counts, splits with each split's token, rows and missing columns, groups, group_gaps); flags (id, kind, severity, text); browser_status (leakage_flagged, not_assessed, review_before_modeling, no_flags); and identifiers. Column ids are the skill's tokens, "column_" + blake2s("eda-v1.1\0column\0" + name, digest_size=8), unless you use --reveal-identifiers; split and group values are always tokens.

The output

One JSON object, serialised as a string at data.output.output. Keys in both lanes: task, status, headline, flag_responses (ref, stance = confirmed | explained | dismissed | needs_data_owner, note), questions_for_data_owner, assumptions. Lane keys are listed in the table above.

Base URL and the envelope

Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses the same envelope, so one helper covers the whole API:

{"ok": true, "data": {"job_id": "job_...", "status": "queued"}}
{"ok": false, "error": {"code": "payment_required", "message": "..."}}

The token is minted for this app (the guest endpoint takes {"slug":"eda-desk"} in its body), so no slug header is needed afterwards. Send it as Authorization: Bearer ....

The input object IS the request body. There is no {"input": ...} wrapper. A wrapped body is answered with an unknown field 'input' warning, and the model never sees your text.

Error codes

statuscodewhat to do
400validation_errorA field is missing or the wrong type. Every field is a string: facts must be a JSON-encoded string, not an object.
401unauthorizedThe token is missing, malformed or expired. Get a new one from the token page.
402payment_requiredThe balance is below min_credits. Call /estimate first and top up.
403forbiddenThe token is valid but not for this app, or a guest token tried a metered run. A guest cannot run; sign in for a personal token.
404not_foundUnknown job id, or the app slug does not exist.
409conflictThe same Idempotency-Key was replayed with a different body. Change the key or send the original input.
429rate_limitedToo many requests. Back off and retry; do not tight-loop.
5xxinternalA server-side failure. Retry with the SAME Idempotency-Key so you are not billed twice.

1. A tiny client

One helper that sends the token, unwraps data and raises on ok: false. The token comes from the token page (Copy token or Copy shell export); step 2 covers the kinds of token and minting one from code.

# Every call is the same three things: the base URL, your bearer token,
# and a JSON body. Keep the token in a shell variable.
BASE="https://api.skillsafe.ai/v1/app-api"
SLUG="eda-desk"
TOKEN="$SKILLSAFE_TOKEN"   # from https://eda-desk.skillsafe.ai/tokens.html

call() {                  # call <path> [json-body]
  if [ -n "$2" ]; then
    curl -sS -X POST "$BASE/$1" \
      -H "Authorization: Bearer $TOKEN" \
      -H "Content-Type: application/json" \
      -d "$2"
  else
    curl -sS "$BASE/$1" -H "Authorization: Bearer $TOKEN"
  fi
}

2. Get a token

The easiest route is the token page: it shows the token this browser already holds, with Copy token and Copy shell export buttons, and a sign-in button for a personal token. A guest token, minted with POST /guest and {"slug":"eda-desk"}, can call /me and /estimate; the run is metered, so /run and /run-stream need a personal token.

# The token page is the shortest path. It shows the token this browser holds and
# hands you a ready-made shell export:
#
#   https://eda-desk.skillsafe.ai/tokens.html
#   export SKILLSAFE_TOKEN="..."
#
# To mint a guest token from the command line instead. A guest token is enough
# for /me and /estimate; a run needs a personal token from signing in.
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
  -H "Content-Type: application/json" -d '{"slug":"eda-desk"}'
# {"ok":true,"data":{"token":"...","subject_type":"guest"}}

3. Check the session and the balance

call me
# {"ok":true,"data":{"subject_type":"user","username":"you","credits":51234}}

4. Price the run (free)

/estimate returns the model binding and the credits a run would reserve. It creates no job and charges nothing. Expect model_alias gpt-terra. hold_credits is a reservation, not the price. min_credits is the least balance that can start a run. What you pay is charged_credits, reported on the finished job, usually far lower. The body is the input object itself, with no {"input": ...} wrapper. /estimate does no input validation, so check the shape yourself: an object whose every value is a string, task equal to split or report, and a facts that is a JSON string parsing to an object.

# body.json is the input object itself - no {"input": ...} wrapper. estimate does
# not validate it, so check the shape first:
python3 -c '
import json
b = json.load(open("body.json"))
assert isinstance(b, dict) and b.get("task") in ("split", "report")
assert all(isinstance(v, str) for v in b.values())
assert isinstance(json.loads(b["facts"]), dict)
'
INPUT=$(cat body.json)

call estimate "$INPUT"
# {"ok":true,"data":{"model":"...","model_alias":"gpt-terra",
#   "markup_bps":...,"hold_credits":...,"min_credits":...,"sponsor_enabled":false}}
# hold_credits is RESERVED, not the price; charged_credits after the run is the cost.

5. Run it, then poll

POST /run returns a job_id; poll GET /jobs/{id} until it is terminal. The reply is a string at data.output.output: JSON.parse it (step 7). Send an Idempotency-Key built from the lane, a hash of the input and the attempt number, eda-desk:<lane>:<hash>:a<attempt>, so a retried request returns the same job instead of billing a second run. Use one key per distinct input: edited facts, context, split_review or question are a new hash, and replaying an old key with a different body is a 409. Any stable digest of the body works. Leave retry_note out of the hash and bump the attempt instead.

# Always send an Idempotency-Key derived from the input. A retried request with
# the same key returns the SAME job instead of billing a second run.
LANE=$(printf '%s' "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["task"])')   # split or report
KEY="eda-desk:$LANE:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"

JOB=$(curl -sS -X POST "$BASE/run" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -d "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')

while :; do
  OUT=$(call "jobs/$JOB")
  STATUS=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
  [ "$STATUS" = "succeeded" ] && break
  [ "$STATUS" = "failed" ] && echo "$OUT" && exit 1
  sleep 2
done

# {"ok":true,"data":{"job_id":"job_...","status":"succeeded",
#   "output":{"output":"{\"task\":\"split\",\"status\":\"resplit_required\",\"headline\":\"...\", ...}"},
#   "charged_credits":...,"truncated":false}}
printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["output"]["output"])' > reply.json

6. Or stream it

POST /run-stream takes the same body and headers and answers with server-sent events: job (the job id), delta (chunks of the reply) and done (the status, charged_credits, truncated and, when present, the full output). A browser page may receive only tick heartbeats and then done, never a delta, so take the reply from done.output.output when it is there, fall back to the concatenated deltas, and fall back again to GET /jobs/{id}.

# Server-sent events. `delta` events carry chunks of the reply; `done` carries the
# status, charged_credits and the truncated flag. Ignore `tick` heartbeats.
curl -N -X POST "$BASE/run-stream" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -H "Accept: text/event-stream" \
  -d "$INPUT"

# event: job    {"job_id":"job_..."}
# event: delta  {"text":"{\"task\":\"split\",\"status\":\"resplit_required\",\"headline\":\"The"}
# event: done   {"status":"succeeded","charged_credits":...,"truncated":false}

7. Parse the reply

The reply is one JSON object serialised as a string. Parse it and check that task is the lane you asked for. Column ids in it are the tokens from your facts (or the sanitized names, if you built facts with --reveal-identifiers); map tokens back with the same BLAKE2s rule the skill's tools use. Read the code before running it.

# The reply is a JSON string inside data.output.output (saved as reply.json in step 5):
python3 -c 'import json;r=json.load(open("reply.json"));print(r["task"],r["status"],r["headline"])'
# A split review: the plan and the code
python3 -c '
import json
r = json.load(open("reply.json"))
if r["task"] == "split":
    print(r["split_plan"]["unit"], r["split_plan"]["method"])
    open("split_plan.py", "w").write(r["code"])   # read it before running it
else:
    for f in r["key_findings"]: print("-", f["finding"])
'
# Column ids in the reply are the tokens from facts; map them back with your own header:
# token = "column_" + blake2s(b"eda-v1.1\0column\0" + name.encode(), digest_size=8).hexdigest()

Costs

Invariants worth asserting