Skip to main content

BlindML

scikit-learn for Encrypted Data

BlindML trains Naive Bayes, Bayesian networks, decision trees, random forests, AdaBoost, logistic regression and histogram models on encrypted records. Zero plaintext exposure. Your machine learning workflows stay intact, your models get the data they need without violating compliance or security policies.

Free and open source on GitHub

Trained on aggregate counts, not plaintext records

Marginals are cross-tabulated count summaries over encrypted fields. The Blind Insight platform computes them by issuing aggregate queries against ciphertext and returning a count table. The model trains on that table—never on feature vectors derived from decrypted records.

blind_ml — training
from blind_ml import NaiveBayesModel

model = NaiveBayesModel().fit(marginals, n_pos=3201, n_neg=76402)
pred, risk = model.predict({"fraud_type": "card_fraud"})
# 0.0 F1 delta vs. plaintext — 600K records
blind grants create — scope training access
blind grants create --data '{
  "name": "blindml-training",
  "field_names": {"__query": true, "risk_level": true, "fraud_type": true, "is_fraud": true},
  "can_create_records": false
}'

Training needs only the query key, never a field’s decryption key. How grants and key sharing work →

Shipped models

Eight algorithms ship today, each training on aggregate counts from encrypted data. Zero records decrypted.

Naive Bayes

Conditional probabilities from class counts.

Gaussian Naive Bayes

Per-class means and variances on numeric fields.

Bayesian Network

Conditional probability tables over a dependency graph.

Decision Tree

Gini splits from class counts per partition.

Random Forest

Ensembles of count-based decision trees.

AdaBoost

Boosted decision stumps, reusing cached counts.

Logistic Regression

OLS from X′X, refined locally with IRLS.

Histogram Classifier

Per-value risk buckets, no independence assumption.

What a Blind Insight–optimized model looks like

Logistic regression needs X′X and X′y. With one-hot features, every cell of both is a count—so the whole design matrix comes back from encrypted aggregate queries, and the solve runs on your machine.

blind_ml — logistic regression from counts
from blind_ml import LogisticRegressionModel
from blind_ml.demo_helpers import get_encrypted_count

def count(q):  # one encrypted aggregate query → one integer
    return get_encrypted_count(client, org, dataset, schema, q)

# Every input to the model is a count
n_pos = count("risk_level:count(50~100)")
n_neg = count("risk_level:count(0~49)")
marginals = {(f, v): count(f"risk_level:count(0~100),{f}:{v}")
             for f, v in dummies}
pos_counts = {(f, v): count(f"risk_level:count(50~100),{f}:{v}")
              for f, v in dummies}
pairwise = {(fa, va, fb, vb):
            count(f"risk_level:count(0~100),{fa}:{va},{fb}:{vb}")
            for (fa, va), (fb, vb) in dummy_pairs}

# X′X and X′y are assembled from those counts and solved locally
model = LogisticRegressionModel(ridge_lambda=0.01).fit_from_counts(
    marginals, pos_counts, pairwise, dummies, n_pos, n_neg, features,
)
model.predict({"fraud_type": "mule_account", "is_active": "true"})
# P(high risk), trained on 0 decrypted records

Implementations, notebooks, and benchmarks: blind-insight/blind-ml on GitHub.

Supported models

Nearly every classical model reduces to counting. If an algorithm can learn from counts, it can train on encrypted data, with no record access required. Similarity search and KNN add encrypted fuzzy matching on top. Details for each are in the approach doc.

Algorithms that train on encrypted aggregate counts, exactly or from binned statistics. Source: blind-insight/blind-ml APPROACH.md.
Algorithm How it trains on encrypted data
Exact on encrypted counts
Decision Trees (entropy / ID3) Information gain instead of Gini, from the same class counts
Ridge Regression X′X from marginal and pairwise counts, plus an L2 penalty tuned on holdout
Gradient Boosted Trees Shallow trees fit to residual distributions built from aggregate counts
Association Rules Support and confidence as ratios of itemset counts
Statistical tests Chi-square, Fisher’s exact and more from contingency tables built with counts Live in BlindStats
Anomaly detection (statistical) Thresholds from means, extremes, and the distribution of counts per bin Live in BlindStats
Similarity search Hamming distance and n-gram fuzzy match on encrypted strings, combined with counts Live on the platform
From binned or summarized statistics
Gaussian Naive Bayes (continuous) Class means via avg; variance from binned count histograms when values can’t be enumerated
Linear Discriminant Analysis Class means via avg; shared covariance from pairwise cross-tab counts
PCA Eigendecomposition of X′X rebuilt from marginal and pairwise counts
K-Means (categorical) Mode-based clusters from per-cluster feature counts
Emerging
KNN on strings Fuzzy match already finds nearest neighbors on encrypted strings; count by class among matches to classify
Linear SVM Primal hinge loss approximated with IRLS-style iterations on aggregate statistics

See BlindStats for statistical tests, drift, and anomaly detection, and fuzzy and similarity matching on encrypted strings.

Build the next model with us

BlindML is open source under MIT. Fork it, open an issue with your approach, and send a PR—every contribution is reviewed by a Blind Insight maintainer. New to the platform? Start with uploading data and running the Blind Proxy.

New algorithm demos

Pick any model from the list above and build the notebook that proves it.

New domains

Insurance fraud, identity verification, clinical risk beyond breast cancer.

Performance

Query batching, smarter caching, and parallel execution patterns.

Docs

Clearer explanations, diagrams, and runnable tutorials.

The numbers.

0.0 F1 delta
vs. plaintext baseline, 600K records
HIPAA k=11
Suppression built in—F1 holds
8 shipped models
scikit-learn-style, validated vs. plaintext
No plaintext exposure
Training and inference on aggregates only

Training benchmarks against FHE alternatives →

Run statistics on the same encrypted data with BlindStats, put an LLM or agent on top of it with Blind(L)LM, and see why you probably don’t need fully homomorphic encryption to get there.

From demo to production in weeks.

Start in the sandbox — nothing to deploy, your schema on synthetic data, live in 72 hours. Validate against live pipelines, then ship across your stack.