Skip to main content

BlindStats

Statistics on data you never decrypt

BlindStats computes descriptive statistics, hypothesis tests, correlations, regression, and drift reports from encrypted aggregate and count queries. Every number matches plaintext exactly. Zero rows decrypted.

Every statistic is an aggregate in disguise

Blind Insight answers aggregate and count queries over an encrypted index. That turns out to be enough. Sums and means come straight from numeric aggregates. Variance is two aggregate calls. Medians and quantiles are a bounded binary search over range counts. Contingency tables are a grid of filtered counts, and regression is solved locally from the same sufficient statistics. BlindStats only ever needs a query key—never a decryption key (how query and field keys work).

blind_stats
from blind_stats import BIStatsSession, BlindInsightClient

stats = BIStatsSession(
    client, org=org, dataset=dataset, schema=schema,
    field_domains={"risk_level": (0, 102)},
)

stats.mean("risk_level").statistic      # 48.31 — from an encrypted aggregate
stats.median("risk_level").statistic    # binary search over count queries
stats.chi2_independence("fraud_type", "is_active")
stats.describe("risk_level")             # 0 rows decrypted

What it computes

Descriptive

Count, sum, mean, min, max, variance, standard deviation, median, quantiles, IQR, histograms, frequencies, and a one-call describe().

Categorical tests

Crosstabs, chi-square, Fisher’s exact, odds ratios, relative risk, and proportion z-tests with confidence intervals.

Grouped and mean inference

Group-by counts, means, and variance; one- and two-sample t-tests; one-way and Welch ANOVA; Kolmogorov–Smirnov and Mann–Whitney.

Correlation and effect size

One-hot covariance and correlation, binned Pearson and Spearman, point-biserial, and eta-squared.

Regression

OLS and ridge regression summaries assembled from aggregate sufficient statistics, plus feature screening.

Monitoring and drift

Missingness, domain violations, outlier counts, population stability index, distribution divergence, chi-square drift, and period summaries.

The numbers.

Exact parity
Encrypted results match plaintext to the last digit
0 rows decrypted
Aggregate and count queries only
Query key only
No decryption key, ever
Small-cell suppression
min_cell_size is a setting, not a review step

Part of the Blind Insight platform, alongside BlindML for training models and Blind(L)LM for running LLMs and agents on encrypted data. See how the approach compares to FHE, TEEs and other methods, or browse the FAQ.

Run the notebook against your own schema.

The repo ships a notebook that checks every statistic against the same computation on plaintext, side by side. Point it at your schema and see the parity for yourself. You’ll need the Blind Proxy running; the getting started guide walks through it.