Season 2, same 18 PDFs, 149 private details, 12 engines ranked by mistakes (leaked private details weigh most, and over-redactions count too):
Sending PDFs with private identity information to a cloud Ai like Claude or ChatGPT is a bad idea. But could a local Ai on your own computer redact them instead? This bake-off runs Cover’s engine and the leading local Ai engines on the same 18 bank, tax and medical PDFs. So far the local Ai engines fall behind Cover in three ways: they over-redact, most are slow, and none offer a fast way to fix their mistakes.
PDF selection focuses on money, home and insurance documents. Real-world samples for anyone asking Ai about their own money, or accountants and tax pros who redact a client's documents before asking Ai for a second opinion. Watch Cover's engine improve with each release, even outpacing Google's tech.
| # | engine | PII accuracy | Over-redaction | Speed |
|---|---|---|---|---|
| 1 | 🥇 Cover | Excellent | Very good | Excellent |
| 2 | 🥈 Gemma 4Google | Very good | Very poor | Good |
| 3 | 🥉 Phi-4Microsoft | Good | Very poor | Poor |
| 4 | Qwen 3Alibaba | Good | Very poor | Very poor |
| 5 | Qwen 2.5Alibaba | Poor | Very poor | Good |
| 6 | DeepSeek R1DeepSeek | Poor | Very poor | Very poor |
| 7 | Llama 3.1Meta | Poor | Very poor | Good |
| 8 | Apple IntelligenceApple | Very poor | Very poor | Excellent |
| 9 | Gemma 3Google | Poor | Very poor | Good |
| 10 | Privacy FilterOpenAI | Poor | Very poor | Very good |
| 11 | RampartAmerica.gov | Very poor | Very poor | Excellent |
| 12 | NemotronOpenMed | Very poor | Very poor | Excellent |
Check out the actual pages below and how each engine redacted them, side by side with the original.
Catches 94% of the private details to Cover's 100%, but still leaves 13 visible and over-redacts about eleven times as much.The strongest of the opposing detectors, catching 94% of the PII against Cover's 100%, with 13 items left visible against none for Cover. It costs the page more: it over-redacts about eleven times as much clean text (157 against 14), and it takes roughly twenty-six seconds per PDF against about three. Cover also needs no multi-gigabyte model download or GPU, and lets you click a box to fix a miss, where this model offers no correction at all.












































































































































































































































































































Cover v1.0.2 vs Google Gemma 4 gemma4:12b-mlx · 18 real-world sample documents.
| metric · all 18 docs | Cover | Gemma 4 |
|---|---|---|
| Recall: PII correctly redacted | 100% | 93% |
| Leaks: confirmed PII left visible | 0 | 16 |
| Over-redaction: non-PII redacted by mistake | 12 | 157 |
| Speed: median time per PDF | ~3.2 sec/PDF | ~26 sec/PDF |
| document | Cover leaks | Gemma 4 leaks | Cover over | Gemma 4 over |
|---|---|---|---|---|
| ubs-portfolio-summary-sample | 0 | 1 | 0 | 86 |
| transunion-cameo-credit-report-sample | 0 | 3 | 0 | 22 |
| pershing-2024-sample-tax-yearend | 0 | 0 | 0 | 17 |
| 2025_Tax_Return_Okafor-Brennan | 0 | 6 | 0 | 15 |
| commercebank-sample-statement | 0 | 0 | 3 | 6 |
| Fidelity-sample-statement | 0 | 1 | 2 | 4 |
| WellsFargo-sample-statement | 0 | 0 | 2 | 3 |
| Lease_Agreement_signed | 0 | 3 | 0 | 2 |
| allstate-auto-declarations-sample | 0 | 0 | 0 | 1 |
| medicare-sample-part-b-msn | 0 | 1 | 4 | 1 |
| att-uverse-sample-bill | 0 | 1 | 0 | 0 |
| capitalone-sample-estatement | 0 | 0 | 0 | 0 |
| cfpb-draft-periodic-mortgage-statement | 0 | 0 | 0 | 0 |
| cfpb-ficus-bank-loan-estimate-sample | 0 | 0 | 0 | 0 |
| cfpb-sample-credit-card-statement-handout | 0 | 0 | 0 | 0 |
| cigna-eob-example | 0 | 0 | 0 | 0 |
| fidelity-large-print-sample-statement | 0 | 0 | 0 | 0 |
| pge-sample-bill-sparky-joule | 0 | 0 | 1 | 0 |
| TOTAL | 0 | 16 | 12 | 157 |
Test set: 18 public sample documents (149 verified PII items), held fixed across runs so results stay comparable. Every engine redacted the same documents; each output is scored against a Claude-authored gold standard (confirmed + ambiguous tiers), and every counted over-redaction is kept auditable. Pages and the leaderboard use the same weighting: leaked private details weigh most, and over-redactions count too.
Opening checkout with your email filled in.