Powered by RiskAICenter · Non-Profit · Advancing Agentic AI in Finance

Free Tool · No Login Required

Bulk-evaluate a whole file of inputs and outputs

Upload a CSV with one column of what you gave an AI system and one column of what it produced — a chat log, a Q&A export, a batch of model calls. Pick a window width K, and a K-row window slides across the whole file, recomputing input–output covariance at every position so you can see where (not just whether) the relationship held or drifted. Parsing and every calculation happen in your browser — the file is never uploaded anywhere.

1. Upload a CSV

Any CSV with a column of inputs and a column of outputs works. Row order matters — rows are read top to bottom, in file order.

No file selected yet.
Mean Covariance
–
Across all rolling windows
Std Dev
–
Of the rolling covariance series
Min / Max
–
Range across windows
Sign Changes
–
Times covariance crossed zero
In Your Range
–
Windows inside your acceptable range
Verdict
–
Within your acceptable range

Set your acceptable covariance range

Stability shouldn't be judged by sign changes alone — a covariance that flips from +0.0001 to -0.0001 has crossed zero but barely moved. Drag these to set the band of covariance values you consider acceptable; every rolling window is then checked against it directly.
–

Rolling covariance —

Each point is the input–output covariance computed over K consecutive rows, starting at that row. The dashed line marks zero; the shaded band is your acceptable range. Points outside it are highlighted. Click any point to see the underlying rows.

Interpretation

Window-by-window detail

One row per window position (start row → end row of that K-row window). Click a row to see the underlying records.
WindowRowsCovariance

Methodology: every input and output in the uploaded column is embedded into the same fixed-dimensional vector space and reduced to a scalar with a fixed projection (the same pipeline as the single-sample evaluator — see /tools/ai-evaluator for the full writeup). A window of K consecutive rows is then slid one row at a time across the whole file, and the ordinary sample covariance between the input and output scalars is computed inside each window, producing one covariance estimate per window position. Per Aldridge's "Regret Equals Covariance" framing, a rolling covariance that stays close to its own average is read as consistent, low-regret system behavior, while a covariance that trends or swings widely is read as behavioral drift. Rather than relying on a single fixed threshold (or on sign changes alone, which can flag a harmless flip near zero as a "change"), the stability verdict above is driven by the acceptable-covariance range you set with the sliders: a window only counts against stability if its covariance actually falls outside the range you defined. The default range (mean ± 1 standard deviation of the rolling series) is a starting point, not a recommendation — narrow it for a stricter read or widen it for a looser one. This is a reference implementation of the general public framing of that methodology, intended for teaching and quick triage rather than as a substitute for a full model-risk review.

Need this running continuously on a live system?

The free evaluator scores a file you upload by hand. The Pro and Team plans connect to your pipeline and re-run this rolling-window test automatically as new data arrives — with drift alerts.

See Pro & Team plans →