Check a dataset's split for leakage before you model
Drop a CSV or TSV and name its split, subject, group and time columns. Your browser runs the exploratory-data-analysis skill's own tools on it - profile, missingness and split-leakage audit, outlier sensitivity - with the same output as the Python, free. Your rows never leave the page. A paid run reads only the aggregates and reviews the split or drafts the EDA report.
Both examples have saved model runs for both lanes - the whole page, free.
Your recent runs
What this does, and what it does not
The exploratory-data-analysis agent skill ships standard-library Python tools for bounded, redacted
EDA of CSV/TSV files. tabular_profile.py infers each column's kind and reports missingness,
distinct counts and numeric aggregates. missingness_leakage_audit.py reports missingness
overall and by group and split, and counts entities, groups, rows (identical apart from the split
column) and time ranges that appear in more than one split. distribution_sensitivity.py
compares mean and SD with median, IQR and MAD, counts values outside the 1.5 IQR fences, and reports
trimmed and winsorized means and moment skewness before and after log1p. None of them prints a raw
value: cells, entity ids and, by default, column names are replaced by BLAKE2s tokens.
This page runs JavaScript ports of all three. Checked against the Python (CPython 3.12) on fuzzed
files - see the notice for the counts - stdout, stderr and exit code matched,
including the cases where the Python refuses a file or stops with an uncaught exception. One exception
is stated plainly: the skewness fields depend on the platform's pow and log1p,
so CPython's own output differs between operating systems in the last digit or two; the page computes
them correctly rounded. The page's flags are its own, labelled as such. A flag is not proof of bias or
leakage, and "not detected" is not proof of independence. The paid lanes read only the aggregates,
the flags and what you type - never a row or a cell value. Derived from the agent skill
@k-dense-ai/exploratory-data-analysis
(k-dense-ai/scientific-agent-skills, MIT; see the notice).
The page's free fixes are its own, not the skill's, and each runs only when you click it. When an
entity, group or identical row sits in more than one split, Resplit by entity (or by
group) moves each one into a single split, chosen by a hash of its value against the current row
shares, audits the result and lets you download it as a CSV. It ignores time, so for forward-in-time
prediction split by time instead. The next audit of a table with the same columns says what changed
since the last one, for example "entities in more than one split 31 → 0". Semicolon exports with
decimal commas, trailing delimiters and blank lines can be rewritten as a plain CSV, and text inside a
numeric column such as NA or <0.5 is offered as a missing code.