Trained on aggregate counts, not plaintext records
Marginals are cross-tabulated count summaries over encrypted fields. The Blind Insight platform computes them by issuing aggregate queries against ciphertext and returning a count table. The model trains on that table—never on feature vectors derived from decrypted records.
from blind_ml import NaiveBayesModel
model = NaiveBayesModel().fit(marginals, n_pos=3201, n_neg=76402)
pred, risk = model.predict({"fraud_type": "card_fraud"})
# 0.0 F1 delta vs. plaintext — 600K records blind grants create --data '{
"name": "blindml-training",
"field_names": {"__query": true, "risk_level": true, "fraud_type": true, "is_fraud": true},
"can_create_records": false
}' Training needs only the query key, never a field’s decryption key. How grants and key sharing work →
Shipped models
Eight algorithms ship today, each training on aggregate counts from encrypted data. Zero records decrypted.
Naive Bayes
Conditional probabilities from class counts.
Gaussian Naive Bayes
Per-class means and variances on numeric fields.
Bayesian Network
Conditional probability tables over a dependency graph.
Decision Tree
Gini splits from class counts per partition.
Random Forest
Ensembles of count-based decision trees.
AdaBoost
Boosted decision stumps, reusing cached counts.
Logistic Regression
OLS from X′X, refined locally with IRLS.
Histogram Classifier
Per-value risk buckets, no independence assumption.
What a Blind Insight–optimized model looks like
Logistic regression needs X′X and X′y. With one-hot features, every cell of both is a count—so the whole design matrix comes back from encrypted aggregate queries, and the solve runs on your machine.
from blind_ml import LogisticRegressionModel
from blind_ml.demo_helpers import get_encrypted_count
def count(q): # one encrypted aggregate query → one integer
return get_encrypted_count(client, org, dataset, schema, q)
# Every input to the model is a count
n_pos = count("risk_level:count(50~100)")
n_neg = count("risk_level:count(0~49)")
marginals = {(f, v): count(f"risk_level:count(0~100),{f}:{v}")
for f, v in dummies}
pos_counts = {(f, v): count(f"risk_level:count(50~100),{f}:{v}")
for f, v in dummies}
pairwise = {(fa, va, fb, vb):
count(f"risk_level:count(0~100),{fa}:{va},{fb}:{vb}")
for (fa, va), (fb, vb) in dummy_pairs}
# X′X and X′y are assembled from those counts and solved locally
model = LogisticRegressionModel(ridge_lambda=0.01).fit_from_counts(
marginals, pos_counts, pairwise, dummies, n_pos, n_neg, features,
)
model.predict({"fraud_type": "mule_account", "is_active": "true"})
# P(high risk), trained on 0 decrypted records Implementations, notebooks, and benchmarks: blind-insight/blind-ml on GitHub.
Supported models
Nearly every classical model reduces to counting. If an algorithm can learn from counts, it can train on encrypted data, with no record access required. Similarity search and KNN add encrypted fuzzy matching on top. Details for each are in the approach doc.
| Algorithm | How it trains on encrypted data |
|---|---|
| Exact on encrypted counts | |
| Decision Trees (entropy / ID3) | Information gain instead of Gini, from the same class counts |
| Ridge Regression | X′X from marginal and pairwise counts, plus an L2 penalty tuned on holdout |
| Gradient Boosted Trees | Shallow trees fit to residual distributions built from aggregate counts |
| Association Rules | Support and confidence as ratios of itemset counts |
| Statistical tests | Chi-square, Fisher’s exact and more from contingency tables built with counts Live in BlindStats |
| Anomaly detection (statistical) | Thresholds from means, extremes, and the distribution of counts per bin Live in BlindStats |
| Similarity search | Hamming distance and n-gram fuzzy match on encrypted strings, combined with counts Live on the platform |
| From binned or summarized statistics | |
| Gaussian Naive Bayes (continuous) | Class means via avg; variance from binned count histograms when values can’t be enumerated |
| Linear Discriminant Analysis | Class means via avg; shared covariance from pairwise cross-tab counts |
| PCA | Eigendecomposition of X′X rebuilt from marginal and pairwise counts |
| K-Means (categorical) | Mode-based clusters from per-cluster feature counts |
| Emerging | |
| KNN on strings | Fuzzy match already finds nearest neighbors on encrypted strings; count by class among matches to classify |
| Linear SVM | Primal hinge loss approximated with IRLS-style iterations on aggregate statistics |
See BlindStats for statistical tests, drift, and anomaly detection, and fuzzy and similarity matching on encrypted strings.
Build the next model with us
BlindML is open source under MIT. Fork it, open an issue with your approach, and send a PR—every contribution is reviewed by a Blind Insight maintainer. New to the platform? Start with uploading data and running the Blind Proxy.
New algorithm demos
Pick any model from the list above and build the notebook that proves it.
New domains
Insurance fraud, identity verification, clinical risk beyond breast cancer.
Performance
Query batching, smarter caching, and parallel execution patterns.
Docs
Clearer explanations, diagrams, and runnable tutorials.
The numbers.
- 0.0 F1 delta
- vs. plaintext baseline, 600K records
- HIPAA k=11
- Suppression built in—F1 holds
- 8 shipped models
- scikit-learn-style, validated vs. plaintext
- No plaintext exposure
- Training and inference on aggregates only
Training benchmarks against FHE alternatives →
Run statistics on the same encrypted data with BlindStats, put an LLM or agent on top of it with Blind(L)LM, and see why you probably don’t need fully homomorphic encryption to get there.
From demo to production in weeks.
Start in the sandbox — nothing to deploy, your schema on synthetic data, live in 72 hours. Validate against live pipelines, then ship across your stack.