← SkillSafe / EDA Desk

Check a dataset's split for leakage before you model

Drop a CSV or TSV and name its split, subject, group and time columns. Your browser runs the exploratory-data-analysis skill's own tools on it - profile, missingness and split-leakage audit, outlier sensitivity - with the same output as the Python, free. Your rows never leave the page. A paid run reads only the aggregates and reviews the split or drafts the EDA report.

Both examples have saved model runs for both lanes - the whole page, free.

data.csv
Drop a .csv or .tsv here, or
Options (the tools' flags)
Drop or paste a table, or load an example.
Run the audits first to price the review.

Your recent runs

What this does, and what it does not

The exploratory-data-analysis agent skill ships standard-library Python tools for bounded, redacted EDA of CSV/TSV files. tabular_profile.py infers each column's kind and reports missingness, distinct counts and numeric aggregates. missingness_leakage_audit.py reports missingness overall and by group and split, and counts entities, groups, rows (identical apart from the split column) and time ranges that appear in more than one split. distribution_sensitivity.py compares mean and SD with median, IQR and MAD, counts values outside the 1.5 IQR fences, and reports trimmed and winsorized means and moment skewness before and after log1p. None of them prints a raw value: cells, entity ids and, by default, column names are replaced by BLAKE2s tokens.

This page runs JavaScript ports of all three. Checked against the Python (CPython 3.12) on fuzzed files - see the notice for the counts - stdout, stderr and exit code matched, including the cases where the Python refuses a file or stops with an uncaught exception. One exception is stated plainly: the skewness fields depend on the platform's pow and log1p, so CPython's own output differs between operating systems in the last digit or two; the page computes them correctly rounded. The page's flags are its own, labelled as such. A flag is not proof of bias or leakage, and "not detected" is not proof of independence. The paid lanes read only the aggregates, the flags and what you type - never a row or a cell value. Derived from the agent skill @k-dense-ai/exploratory-data-analysis (k-dense-ai/scientific-agent-skills, MIT; see the notice).

The page's free fixes are its own, not the skill's, and each runs only when you click it. When an entity, group or identical row sits in more than one split, Resplit by entity (or by group) moves each one into a single split, chosen by a hash of its value against the current row shares, audits the result and lets you download it as a CSV. It ignores time, so for forward-in-time prediction split by time instead. The next audit of a table with the same columns says what changed since the last one, for example "entities in more than one split 31 → 0". Semicolon exports with decimal commas, trailing delimiters and blank lines can be rewritten as a plain CSV, and text inside a numeric column such as NA or <0.5 is offered as a missing code.