De-identification requires more than recognizing an entity: a system must localize its complete surface form and replace it without corrupting the relationships that make a corpus useful. We study these requirements as separate, measurable objectives: typed span extraction, sensitive-character coverage, and coherent transformation.
We evaluate Nyne Blackvault Ensemble, an ensemble, and Nyne Blackvault, a single model. On all 5,198 human-annotated PrivacyBench messages, Blackvault Ensemble obtains 94.17 exact F1 (95% document-bootstrap CI 93.70 to 94.64), above micro1's published 93.40. Blackvault obtains 93.86 exact F1 and also achieves 92.11 and 93.41 strict typed-span F1 on 500-record AI4Privacy and Nemotron-PII cohorts. The evaluated Blackvault transformation configuration records 97.09 combined detection-and-coherence accuracy.
We report benchmark comparisons with explicit metric definitions and uncertainty estimates. Exact localization, broad native-label evaluation, and coherent replacement are measured separately, so each result answers a defined privacy-engineering question.
Protect the identity.
Keep the story connected.
Enterprise information is relational. A person appears in a message, an email address, an approval, and a company record. Removing every identifier can erase the distinctions a useful dataset needs. Replacing each mention independently can produce fluent text while turning one person into several unrelated fictional people.
We study two coupled requirements: complete, correctly typed localization and coherent replacement across related mentions. The first asks whether the right characters were found. The second asks whether the transformed corpus still expresses the intended relationships. They require different measurements, and neither is established by a single detection score.
- 01
Define the input
Keep the document inventory, original character coordinates, and declared entity policy.
- 02
Locate identifiers
Predict typed spans. Blackvault Ensemble combines detectors; Blackvault is the single-model alternative.
- 03
Connect related mentions
Use available corpus context to associate names, contact identifiers, and organization variants.
- 04
Transform coherently
Assign related fictional values and render them in context, retaining relevant relationships.
- 05
Measure the outcome
Score boundaries, coverage, unnecessary masking, and replacement coherence as separate endpoints.
This view describes the evaluated text workflow. The transformation track uses a recall-oriented detection input; its operating configuration is distinct from the exact-F1 detection headline.
One order. Three records. The same identities.
Source records
Mara Chen sent PO-2048 to Owen Vale at Cedar Labs.
Owen confirmed PO-2048 with m.chen@example.com.
Cedar Labs will deliver PO-2048 on Tuesday.
Change the identities. Preserve the allowed relationships.
Iris Liu sent PO-2048 to Theo Reed at Orchard Works.
Theo confirmed PO-2048 with i.liu@example.net.
Orchard Works will deliver PO-2048 on Tuesday.
Mara Chen and the related email become Iris Liu and i.liu@example.net. Owen remains Theo, and Cedar Labs remains Orchard Works. The order and delivery context stay intact.
What this report contributes
We compare exact and overlap detection against the complete available published field, test broader native label inventories on AI4Privacy and Nemotron-PII, and separate missed detections from incoherent replacements. Explicit metric definitions, uncertainty estimates, and observed output examples explain what each result measures.
What does a correct detection mean?
PII extraction is structured prediction under an annotation contract. A detector can identify the right person but assign the wrong name role, cover only part of an identifier, or merge several entities into one span. These errors have different consequences for privacy and downstream utility. A single overlap score obscures those distinctions.
1.1 Typed spans and label-conditioned extraction
Let xi be a document and ℒ its allowed entity labels. Gold and predicted entities are triples (a, b, ℓ): a half-open interval of original character offsets and an entity type. This fixes the unit of comparison before tokenization, normalization, or replacement.
Prior work on label-conditioned extraction, including GLiNER [1], formulates entity detection through contextual spans and entity descriptions. This perspective makes label semantics central to measurement: a complete personal name, a given name, and an organization are different annotation targets. Our evaluation retains each cohort's declared label inventory and states the matching rule explicitly.
1.2 Official overlap and exact-match estimands
PrivacyBench's official scorer [6] uses two directional overlap counts. Let G and N denote total gold and predicted spans. D counts gold spans touched by a same-label prediction; M counts predictions touching a same-label gold span. Write g ∼ p when their labels agree and their character intervals have a nonempty intersection:
Overlap matching is not one-to-one: D and M can differ. Exact scoring instead requires the complete typed triple. The exact metric therefore requires both the correct label and complete character boundaries:
Overlap is not complete coverage.
Consider two sensitive spans and one prediction with the same entity label. The prediction touches both spans, but covers only one character in each.
[0, 4) ∪ [8, 12)[3, 9)2 of 8- Overlap F1
- 100%
- Both gold spans are touched; the prediction touches gold.
- Exact F1
- 0%
- Neither gold span has an exact boundary match.
- Sensitive-character recall
- 25%
- Only 2 of 8 sensitive characters are covered.
1.3 From entity recognition to masking coverage
For privacy diagnostics, let UiG and UiP be unions of gold and predicted character intervals, ignoring type. Sensitive-character recall measures what a redaction would cover; excess masking measures the fraction of annotated non-sensitive characters it would remove:
These diagnostics separate missed sensitive content from loss of useful context. They inform acceptance checks independently of leaderboard F1. They are coverage proxies; coherent replacement is a distinct endpoint, defined in Section 6.
1.4 From span recall to complete identity coverage
An identity can recur across many records. Masking nine mentions while leaving a tenth intact can still reveal the person behind the corpus. The Text Anonymization Benchmark makes this distinction explicit by evaluating coverage at the entity level [12]. Let ℰ be the nonempty set of annotated entities, and let Me be the nonempty set of mentions of entity e. The binary indicator zj equals one if and only if mention j is completely covered. Complete-entity coverage is:
The product makes coverage conjunctive: every mention must be covered. Let pj = Pr(zj = 0), the probability of failing to cover mention j completely. The chance of at least one coverage failure obeys:
These bounds require no independence assumption. Repeated aliases and shared formatting can correlate errors, so multiplying independent success probabilities is generally unjustified. This is a statement about annotated coverage, not a re-identification probability; F1 cannot be substituted for the individual miss probabilities. It explains why Nyne treats identity continuity as a corpus-level problem.
1.5 Preserve context under a coverage requirement
Removing an entire document can maximize masking coverage while destroying its value. A more useful objective is to minimize excess masking subject to a declared missed-character budget ε. For a masking policy T, Equation 4's coverage and excess-masking quantities define:
This separates the privacy requirement from the utility cost instead of hiding both inside one weighted average. Constrained treatment of asymmetric errors has a formal precedent in Neyman-Pearson statistical learning [13]. Here, excess masking is a proxy for context loss; downstream task utility remains a separate question. For a nonempty feasible set with an attained minimum, the optimizer states the preferred policy. These equations describe analytical requirements, not additional measured results or a claim that the optimization has been solved.
1.6 From metric definitions to observed behavior
What Nyne gets right
that the baseline misses.
Two recorded contrasts between Nyne Blackvault and untuned NVIDIA GLiNER-PII, drawn from the public synthetic finance probe.
<fpml:name>Gallet-Bonneau</fpml:name>
Annotated target: Family nameOriginal offsets [962, 976)
Gallet-Bonneau
last name
The prediction matches the entire annotated surname and its type. All 14 characters are covered.
- No prediction overlaps this surname.
A redaction based on these predictions would leave the complete surname visible in this field.
For this annotated surname, Blackvault contributes one true positive and the baseline contributes one false negative under strict typed matching. Name merging does not change this target's result.
WHY IT MATTERS TO NYNEPrivacy-sensitive names occur inside schemas and machine-readable records, not only conversational prose. This case shows why we inspect complete span coverage: recognizing nearby entities is insufficient if one intact identifier remains.
Whole-record scores and scoring convention
These scores use all 3 annotated entities in the complete 1,204-character input, not only the excerpt. Both systems used the same mapped prompt list and operating point.
| System | Reading | TP / FP / FN | F1 |
|---|---|---|---|
| Nyne Blackvault | Raw strict | 2 / 5 / 1 | 40.00 |
| Nyne Blackvault | Name-merged strict | 2 / 5 / 1 | 40.00 |
| Untuned NVIDIA GLiNER-PII | Raw strict | 1 / 7 / 2 | 18.18 |
| Untuned NVIDIA GLiNER-PII | Name-merged strict | 2 / 4 / 1 | 44.44 |
Blackvault identifies this surname, but a different annotated name elsewhere does not receive a complete exact match. Under the declared name-merged full-record reading, both systems match two of three annotations. Blackvault has five unmatched predictions versus four for the baseline, so Blackvault's full-record F1 is lower despite its success on the displayed surname.
Name merging accommodates the source's full-name annotation convention by joining adjacent predicted given names
and surnames using the same declared rule for both systems. It does not change gold annotations. The baseline is
pinned to NVIDIA revision bd23e8ef.
Explore a harder iteration: four predictions, their scores, and why the difference matters
One sentence. Four detection outcomes.
Follow a difficult boundary repair through overcorrection, a name-role error, and exact localization.
Invented text and predictions, not recorded model outputsForward ChenFamily name's note to MaraGiven name at m.chen+finance@example.comEmail; MaraGiven name ChenFamily name owns invoice #2048.
For this example, all name mentions refer to the same fictional person. Names and the complete email are sensitive; the possessive ending, surrounding prose, and invoice number are ordinary context under this stated annotation policy.
Every entity is touched. The email is still incomplete.
Forward Chen's note to Mara at m.chen+finance@example.com; Mara Chen owns invoice #2048.
- Overlap F1 ↑
- 100% M = 5, D = 5
- Exact F1 ↑
- 80% 2 × 4 / (5 + 5)
- Character recall ↑
- 52.38% 22 / 42 sensitive characters
- Excess masking ↓
- 0% 0 / 46 ordinary characters
This hypothesis emits 5 spans against five annotations (N = 5, G = 5). The arrows show the preferred direction. Scores follow Equations 2-4.
All four name mentions match exactly. The email prediction stops at m.chen, leaving +finance@example.com outside the mask. Because that fragment overlaps the annotated email with the right type, overlap F1 still awards full credit.
Forward [PII]'s note to [PII] at [PII]+finance@example.com; [PII] [PII] owns invoice #2048.
Inspect the exact spans and count calculation
Offsets use the original 88-character sentence and half-open intervals [start, end). Overlapping predictions count separately for entity metrics; character coverage takes their union.
| Reference span | Prediction | Exact match |
|---|---|---|
ChenFamily name · [8, 12) | ChenFamily name · [8, 12) | Yes |
MaraGiven name · [23, 27) | MaraGiven name · [23, 27) | Yes |
m.chen+finance@example.comEmail · [31, 57) | m.chenEmail · [31, 37) | No |
MaraGiven name · [59, 63) | MaraGiven name · [59, 63) | Yes |
ChenFamily name · [64, 68) | ChenFamily name · [64, 68) | Yes |
Gold-side overlap hits D = 5; prediction-side overlap hits M = 5; exact matches X = 4. Overlap precision = 5/5 and recall = 5/5. Exact precision = 4/5; exact recall = 4/5.
Protect the identifier. Preserve the relationship.
A useful transformed record should still tell us who owns the invoice. After correct detection, a consistent fictional replacement could read:
Forward Vale's note to Iris at i.vale+finance@example.net; Iris Vale owns invoice #2048.
Names and email now belong to one synthetic identity. Replacing each mention independently could sever that connection. This is an illustrative coherence target, not a judged synthesis result.
A better number must represent a better transformation.
We evaluate exact boundaries, sensitive-character coverage, excess masking, and coherent replacement separately because each reveals a different failure. An iteration should address a diagnosed error and then satisfy declared acceptance criteria.
From candidate spans
to coherent replacement.
Consider “Mara Chen signed for Cedar Labs.” Detecting “Mara” is not equivalent to detecting “Mara Chen,” and replacing either string does not establish that a later email belongs to the same person. The system has to carry a decision through three representations: scores over candidate spans, a set of accepted mentions, and identities that remain consistent after replacement. The formulation below follows that chain.
2.1 Make the complete identifier win
Nyne Blackvault uses a fine-tuned GLiNER-PII detector. Following GLiNER's architecture, contextual word and label representations are projected into a shared space [1][10]. Let hi and vℓ denote those contextual representations, with learned functions producing the span vector sij and label vector eℓ. Stack n span vectors and k label vectors in a shared dimension d:
“Mara Chen” and “Mara” occupy different rows of S but share the PERSON label vector. Their scores depend on much of the same context and many of the same parameters. The sigmoid acts independently on each pair; a high PERSON score for the fragment does not suppress the full name. Their competition is resolved later.
To see how learning changes that competition, fix the input and candidate inventory and flatten Z in column order into z. Let Gθ = ∂z/∂θT be the score Jacobian, and let r = ∇zJ be the residual of a scoring loss J(z). A local gradient-descent update gives, for sufficiently small η > 0:
Here the score map and loss are twice continuously differentiable locally, with bounded second derivatives; all derivatives on the right are evaluated before the update. Each entry Kab is the inner product of two candidates' parameter gradients. Its off-diagonal terms describe how changing one candidate affects another. This finite-network tangent-kernel view [20] analyzes the coupling; it does not require materializing a kernel in the detector. The displayed step is a local reference update rather than a specification of the optimizer [16].
The boundary decision depends on a contrast within that coupled system. Let a be the correctly typed full name and b an overlapping fragment. Define vab as the vector with +1 at a, −1 at b, and zeros elsewhere. The ranking margin Δab and eligibility margin γa are:
For 0 < τa < 1, γa says whether the full span clears its score cutoff; Δab says whether it outranks this fragment. Although the first-order change in the loss in Equation 9 is nonpositive, the change in this particular ranking margin can have either sign. A better aggregate objective can therefore coexist with a worse boundary decision. Exact-span evaluation tests the decision that survives this competition, including competing candidates beyond this pair.
2.2 Choose which candidates survive
The scores now have to become an admissible set of spans. Map word endpoints back to original character intervals, index typed candidates by a, and put conflicting pairs in ℭ. Under a flat annotation policy, overlap makes a pair incompatible. If ua records whether candidate a is selected, the admissible set is:
A greedy decoder visits eligible candidates in score order and retains compatible spans. The full name can lose to a fragment even when both clear the cutoff; conversely, raising the cutoff can remove both. This is the connection between the margins above and the sensitive characters left in the output. The constraints specify feasible selections, without asserting that greedy selection solves a global maximum-score problem.
A high score can still select the wrong boundary.
Change one span vector, then adjust the cutoff to see which characters a simple decoder masks.
Mara Chen signed for Cedar Labs.
1. Score each span against each label
Label vectors are PERSON = (1, 0) and ORG = (0, 1). Their dot products with each span
vector give the logits below. Apply σ(z) = 1 / (1 + e⁻ᶻ) independently to each cell.
| Candidate / vector | PERSON | ORG |
|---|---|---|
Mara Chen(2.4, -1.2) Kept: PERSON | 91.68% logit 2.4 Selected | 23.15% logit -1.2 |
Mara(1.8, -0.8) Overlap rejected | 85.81% logit 1.8 | 31.00% logit -0.8 |
Cedar Labs(-1.4, 2.2) Kept: ORG | 19.78% logit -1.4 | 90.02% logit 2.2 Selected |
Sigmoid scores are independent, not a softmax distribution. They need not sum to 100%.
2. Keep high scores, reject overlapping spans
This toy decoder sorts eligible candidate-label pairs by score, highest first, and keeps a pair only if its span does not overlap one already kept.
- Mara Chen PERSON · 91.68% Keep
- Cedar Labs ORG · 90.02% Keep
- Mara PERSON · 85.81% Overlap rejected
3. Apply the selected masks
[PERSON] signed for [ORG].
2 spans selected. 19 of 19 sensitive characters covered.
Both full sensitive spans are covered. Increasing the threshold can discard a correct entity.
Some decisions depend on the candidate's meaning in context as well as its neural score. The evaluated Blackvault Ensemble includes structured candidate review using TypeSafe's Jev. Jev assesses proposed spans and returns constrained type or non-entity decisions; the surrounding system uses these outputs to retain, reject, or relabel candidates and handle uncertain outcomes. Blackvault performs detection without this review stage. Jev supplies contextual review here, while identity linking and replacement are downstream tasks. [15]
This review can be expressed as a decision over a finite action set 𝒜. Let qk(x, a) be the posterior probability of true class k for candidate a in context x, and let C(α, k) assign the cost of an action α under that class. Then the minimum-risk action is [17]:
Retaining an entity, rejecting non-entity text, changing a type, and deferring a decision have different consequences. The action set and cost matrix express those differences. This is a decision-theoretic interpretation of the interface; the evaluated review policy is not claimed to solve this optimization. Substituting estimated probabilities requires calibration in the decision domain [19]. The span scores in Equation 8 establish neither that calibration nor a probability of disclosure.
2.3 Keep one identity consistent across records
The accepted spans identify where a transformation may act. They do not yet establish which mentions must share a replacement. Suppose “Mara Chen,” “M. Chen,” and a related email have been resolved to one identity. Represent the assignments of m = 1Tu accepted mentions to c identity classes by a binary matrix B, with exactly one nonzero entry in each row. Its induced relation H records which mention pairs share an identity:
This representation enforces transitivity: if two mentions share a class with a third, they share a class with each other. Independent pairwise decisions need not satisfy that property [18]. The task is to infer the right partition and then preserve it during replacement; a coherent but incorrect partition would consistently repeat the original resolution error.
Let R assign the c original identity classes to s fictional identities, where s ≥ c. Giving every class one surrogate and preventing collisions between classes yields an exact preservation condition:
The equality follows because the rows of R are distinct one-hot vectors, so RRT = Ic. It states precisely what “the same fictional person across records” means. A surname and an email can render different attributes of that identity while preserving the relation. If two original classes collide, RRT gains off-diagonal entries and unrelated mentions merge. If one original class receives several surrogates, a single assignment matrix R no longer describes the transformation and its links can split.
Nyne's evaluated transformation stage uses deterministic identity resolution, surrogate assignment, and contextual rendering. Equation 14 expresses the invariant that this process seeks to preserve. The pairwise measures in Section 6 define how to count departures from it as merges and splits [14].
2.4 Apply the decisions to an entire corpus
The same construction scales from one sentence to a corpus 𝒳 = (x1, …, xD). Concatenate only its coordinate inventories: let L be the total number of original character positions and N the total number of typed candidates. Each document has its own contextual span and label matrices; stack their scores and selection vectors in the same document order. A sparse incidence matrix A records which characters each candidate covers. Each column lies entirely inside one document's block. Selection then projects from candidate space back into character space:
The indicator is elementwise: a character is covered if at least one selected candidate contains it. This is the coverage vector underlying Equation 4, now expressed directly as a function of the decisions u. Meanwhile, BR assigns every selected mention a surrogate identity. Together they specify where the renderer acts and which fictional identity supplies each replacement.
For document d, sort its md selected spans by original offsets (adj, bdj, ℓdj). Let i(d, j) identify the corresponding row of B, and let ψ render the assigned surrogate in the entity type and surface form of that mention. Writing ⊕ for string concatenation and setting bd0 = 0, the transformed document is:
Each untouched gap is copied from the original document; each selected span is rendered from its shared identity assignment. Replacement strings may have different lengths because every slice still refers to the source coordinates. When a document has no selected spans, the empty concatenation leaves the whole document intact. Applying this operator to every document produces the transformed corpus without treating repeated mentions as independent replacement problems. Coverage specifies the selected characters; rendering must also change their sensitive values appropriately.
2.5 What success requires across records
Preserving the accepted-mention partition cannot recover an identifier omitted before that partition was built. Return to all annotated mentions of an identity e. Let Ae mean that every such mention is correctly localized and typed, Be that their identity assignments are correct, and Ce that their replacements change the sensitive values coherently. Complete success requires their intersection:
This factorization retains the dependence between stages. A boundary error changes what the resolver receives; a resolution error changes which surrogate the renderer uses. Multiplying their marginal accuracies would erase those interactions. If an earlier conditioning event has probability zero, complete success is zero. Otherwise the conditional factors describe exactly where the remaining risk enters.
The chain connects representation learning to the output requirements: complete spans must survive selection, accepted mentions must form the right identity classes, and replacement must preserve those classes while changing sensitive values. These identity-level events define the analytical target. The following results report the measured detection and transformation endpoints, with their respective populations and denominators.
The benchmark, the metric,
and the comparison field.
Evaluation begins with a defined population, annotation policy, and endpoint. Selection can overfit a noisy metric, as analyzed by Cawley and Talbot [2] and Varma and Simon [3]. We therefore interpret these measurements at their stated scope, distinguish exact localization from overlap, and examine errors beyond the headline score.
Nyne Blackvault Ensemble
An ensemble evaluated for high-fidelity PII localization on PrivacyBench.
Nyne Blackvault
A single detector evaluated across PrivacyBench, AI4Privacy, and Nemotron-PII.
3.1 Benchmark populations
We evaluate all five PrivacyBench entity types, then broaden coverage to the 20 native AI4Privacy types and 55 Nemotron-PII types. Each declared cohort retains negative records and every annotated type. The external evaluations use 500 English validation records from AI4Privacy and 500 test records from Nemotron-PII. [8][9]
| Benchmark | Population | Entity types |
|---|---|---|
| PrivacyBench | 5,198 messages | 5 entity types |
| AI4Privacy | 500 English validation records | 20 native label types |
| Nemotron-PII | 500 test records | 55 native label types |
3.2 Interpreting a reported score
The primary detection results pool counts across records before computing F1. We report 95% bootstrap intervals for Nyne measurements, following the resampling principle of Efron [4]. These quantify sampling uncertainty in the evaluated outputs. Published competitor results are reproduced from their sources; they are not paired reruns on our records. Detailed comparison scope is collected in the source notes.
#1 in exact F1.
The entire field, compared.
Across all nine systems shown, Blackvault Ensemble ranks first in exact-match F1. On 5,198 messages containing 8,805 human-annotated spans, it obtains 94.17 exact F1, a 0.77-point improvement over micro1's published 93.40. Exact precision is 94.44 and recall is 93.90. Blackvault obtains 93.86 exact F1. Blackvault Ensemble's exact-match interval lies above the published reference; Blackvault's higher point estimate has an interval that includes it. These comparisons use reported reference values rather than paired predictions. [7][P1]
PrivacyBench detection
Blackvault Ensemble leads the displayed field and is above micro1’s published exact-match F1.
F1 / score (%) · scale 0 to 100
- 1Nyne Blackvault Ensemble94.17
95% CI 93.70 to 94.64
- 2Nyne Blackvault93.86
95% CI 93.37 to 94.32
- 3micro1 flow-transform 1.093.40
- 4Tonic Textual92.30
- 5Claude Opus 4.889.50
- 6Claude Sonnet 4.688.50
- 7Microsoft Presidio86.20
- 8Claude Haiku 4.585.50
- 9GLiNER2 / NeMo Anonymizer84.40
All available field rows · ranked by point estimate · equal displayed scores share a rank
Blackvault Ensemble's overlap F1 is 96.14 [95.79, 96.48], numerically above micro1's 96.00; its interval includes that reference. Blackvault ties micro1 at the displayed overlap F1 of 96.00. The difference between the two readings is scientifically useful: exact matching demands type and boundary fidelity, while overlap asks whether entities were touched. The stronger exact result therefore supports a localization claim. It does not isolate which component caused the gain without a controlled ablation.
Equal-export macro F1 changes Blackvault Ensemble's exact score from 94.17 to 94.10 and Blackvault's from 93.86 to 93.83. The small shift shows that the pooled result is not driven solely by weighting larger exports more heavily. Precision, recall, and both F1 readings are retained in the complete table below.
Table 2. Full precision, recall and F1; aggregation sensitivity
| System | Exact P | Exact R | Exact F1 | 95% CI | Overlap P | Overlap R | Overlap F1 | 95% CI |
|---|---|---|---|---|---|---|---|---|
| Nyne Blackvault Ensemble | 94.44 | 93.90 | 94.17 | 93.70 to 94.64 | 96.42 | 95.87 | 96.14 | 95.79 to 96.48 |
| Nyne Blackvault | 94.60 | 93.13 | 93.86 | 93.37 to 94.32 | 96.74 | 95.28 | 96.00 | 95.64 to 96.34 |
| micro1 flow-transform 1.0 | 94.34 | 92.48 | 93.40 | Not reported | 96.95 | 95.07 | 96.00 | Not reported |
| Tonic Textual | 94.30 | 90.50 | 92.30 | Not reported | 96.50 | 92.60 | 94.50 | Not reported |
| Claude Opus 4.8 | 93.10 | 86.10 | 89.50 | Not reported | 95.40 | 88.20 | 91.70 | Not reported |
| Claude Sonnet 4.6 | 94.70 | 83.00 | 88.50 | Not reported | 97.00 | 85.00 | 90.60 | Not reported |
| Microsoft Presidio | 85.40 | 87.10 | 86.20 | Not reported | 87.30 | 89.00 | 88.10 | Not reported |
| Claude Haiku 4.5 | 95.00 | 77.70 | 85.50 | Not reported | 97.60 | 79.90 | 87.80 | Not reported |
| GLiNER2 / NeMo Anonymizer | 82.80 | 86.10 | 84.40 | Not reported | 85.20 | 88.60 | 86.90 | Not reported |
Blackvault Ensemble pooled / macro exact F1: 94.17 / 94.10; overlap: 96.14 / 96.08. Blackvault pooled / macro exact F1: 93.86 / 93.83; overlap: 96.00 / 95.96. micro1 does not specify its aggregation; pooled is assumed for the reference comparison.
Scores are rounded directly from the underlying measurements. The final field row preserves the GLiNER2 label used in micro1's detection table.
Leading scores.
Two broader label inventories.
PrivacyBench's five labels cannot by themselves establish broad PII coverage. We retain all 20 AI4Privacy and 55 Nemotron-PII native types on fixed 500-record cohorts, including 1,237 and 4,019 annotated spans respectively. Blackvault is evaluated on both.
Results use strict typed exact-span F1: a detection must recover both boundaries and the correct entity type. The complete field includes micro1, NeMo Anonymizer, NVIDIA's published GLiNER-PII model-card result, and a GLiNER-PII control measured on our cohorts. Published figures retain the metric stated by their source.
External benchmark performance
Strict typed exact-span F1 (%) · All available reference systems
AI4Privacy
500 RECORDSF1 / score (%) · scale 0 to 100
- 1Blackvault92.11
Nyne measurement
95% CI 90.39 to 93.71
- 2micro1 flow-transform 1.083.74
Published reference
- 3NeMo Anonymizer73.08
Published reference
- 4NVIDIA GLiNER-PII (same cohort)64.36
Same-cohort baseline
95% CI 61.41 to 67.17
- 5NVIDIA gliner-PII (card)64.00
Published reference
All available field rows · ranked by point estimate · equal displayed scores share a rank
Nemotron-PII
500 RECORDSF1 / score (%) · scale 0 to 100
- 1Blackvault93.41
Nyne measurement
95% CI 92.42 to 94.37
- 2NVIDIA GLiNER-PII (same cohort)89.66
Same-cohort baseline
95% CI 88.48 to 90.77
- 3micro1 flow-transform 1.088.33
Published reference
- 4NeMo Anonymizer87.84
Published reference
- 5NVIDIA gliner-PII (card)87.00
Published reference
All available field rows · ranked by point estimate · equal displayed scores share a rank
5.1 Broad PII coverage
Under the strict reading, Blackvault reaches 92.11 [90.39, 93.71] on AI4Privacy and 93.41 [92.42, 94.37] on Nemotron-PII, the highest strict scores in each displayed comparison.
micro1 reports strict-span F1 of 83.74 on AI4Privacy and 88.33 on Nemotron; its NeMo references are 73.08 and 87.84. We preserve that stated metric. The numerical comparison is informative, while the directly measured NVIDIA baseline provides the stronger control on population and scoring. Published rows are not paired reruns. [P2]
Table 3. Full external benchmark results
| System | Strict F1 | 95% CI |
|---|---|---|
| NVIDIA GLiNER-PII (same cohort) | 64.36 | 61.41 to 67.17 |
| Nyne Blackvault | 92.11 | 90.39 to 93.71 |
| micro1 flow-transform 1.0 | 83.74 | Not reported |
| NeMo Anonymizer | 73.08 | Not reported |
| NVIDIA gliner-PII model card | 64.00 | Not reported |
| System | Strict F1 | 95% CI |
|---|---|---|
| NVIDIA GLiNER-PII (same cohort) | 89.66 | 88.48 to 90.77 |
| Nyne Blackvault | 93.41 | 92.42 to 94.37 |
| micro1 flow-transform 1.0 | 88.33 | Not reported |
| NeMo Anonymizer | 87.84 | Not reported |
| NVIDIA gliner-PII model card | 87.00 | Not reported |
Rows labelled published reference retain the source's stated strict metric. The measured NVIDIA baseline uses the same cohort and declared scorer as Nyne.
Detection and replacement
are coupled objectives.
Replacing each detected string independently can preserve local readability while breaking corpus-level identity relationships. Coherent transformation must keep a person's related names, handles, and email addresses consistent. We evaluate replacement quality separately from detection to distinguish missed entities from failures to transform a detected entity correctly.
6.1 A decomposable endpoint
Tonic's synthesis scorer selects the same-label prediction with the largest intersection for each gold span. Let G be annotated spans, D those detected, and C those changed and judged coherent. Unchanged replacements receive no credit; missing verdicts remain in the denominator. The three reported quantities are:
This factorization distinguishes a missed detection from an incoherent replacement. Equivalently, the combined error can be partitioned into two disjoint contributions:
For the evaluated Blackvault transformation configuration, recall is 98.07% and conditional accuracy is 99.00%. Using the rounded rates, about 1.93 percentage points of annotated spans are missed and 0.98 points are detected without coherent transformation. The residual is about 2.91%, consistent with the measured 97.09% combined result. This endpoint has no precision term because the generated annotation set is not exhaustive.
6.2 Identity links: two distinct failure modes
Local fluency cannot establish relational consistency. On a fixed set of aligned mentions, let A be the set of unordered mention pairs that share an original identity, and let BT contain pairs assigned the same fictional identity. Pairwise link precision and recall distinguish accidental identity merges from identity splits:
These errors count mention pairs: a merge invents a relationship between different people; a split breaks the connection between two mentions of one person. Pairwise evaluation is established in coreference research [14]. Its pair counts give larger identity clusters more weight; singleton entities and missing mentions need separate coverage accounting. For nonempty pair sets, the equations state a corpus-level consistency objective. The judged transformation scores below measure coherent replacements, not this separate link metric.
The highest combined score in the reported field
Blackvault + transformation
adaptive after two judged looks
#1 combined score in this comparison
F1 / score (%) · scale 0 to 100
- 1Blackvault + synthesis v1.697.09
adaptive after two judged looks
95% CI 96.68 to 97.49
- 2micro1 flow-transform 1.095.46
published; development exposure undisclosed
- 3Tonic Textual + Opus 4.892.00
published
All available field rows · ranked by point estimate · equal displayed scores share a rank
6.3 Conditional accuracy and end-to-end success
The evaluated Blackvault transformation configuration records 99.00 [98.76, 99.23] synthesis accuracy and 97.09 [96.68, 97.49] combined (adaptive after two judged looks), alongside micro1's published 98.28 and 95.46. The useful distinction is between conditional replacement quality and total annotated-span success: a coherent replacement cannot recover a span that was never detected.
We use Tonic's synthesis evaluator to assess whether replacements change the source value and remain coherent, with Claude Opus 4.7 providing the stored judgments. Confidence intervals resample document counts while holding those judgments fixed. They describe the reported configuration and do not capture repeated-judge variability. The comparator values are published references, with potentially different evaluation populations and judge configurations.
Table 4. Transformation results
| System | NER recall | Synthesis | Combined |
|---|---|---|---|
| Nyne Blackvault + synthesis v1.6adaptive after two judged looks | 98.07 | 99.00 | 97.09 |
| micro1 flow-transform 1.0published; development exposure undisclosed | 97.13 | 98.28 | 95.46 |
| Tonic Textual + Claude Opus 4.8published | 95.00 | 97.00 | 92.00 |
This task uses 8,253 generated-gold spans, a different denominator from the 8,805 human-annotated detection spans. Synthesis scores do not establish enterprise-wide Transformation Quality Index (TQI), re-identification resistance or guarantees of anonymity.
What the evidence supports.
Three findings emerge. First, exact localization provides the strongest PrivacyBench comparison: Blackvault Ensemble's point estimate and interval exceed the published reference. Second, Blackvault supports broader native-label inventories, with competitive strict scores on AI4Privacy and Nemotron-PII. Third, the transformation evaluation measures detection and coherent replacement together.
Our endpoint is empirical extraction and transformation quality. Differential privacy, for example, constrains output distributions on neighboring datasets, a different mathematical object [5]. The measurements here concern detection, coverage, and coherence; they do not establish such a mechanism-level guarantee.
Source and measurement notes
Three short notes define the scope of the reference comparisons and uncertainty estimates.
P1Published reference comparisons
PrivacyBench field rows and external micro1/NVIDIA model-card references are published values, not paired reruns. Exact samples, aggregation and complete scoring details are not fully specified. Transformation populations and judge configurations may also differ. Nyne intervals describe its measured scores, not paired differences against these references. Both Nyne overlap F1 intervals include micro1's published value.
P2External scoring readings
Strict typed exact-span F1 is the primary measured endpoint. micro1 and NeMo published values retain their source's stated strict-span metric; their exact samples and implementations are not verified here. The Nyne-executed NVIDIA control uses the same 500-record cohorts and native label inventories as Nyne.
P3Conditional uncertainty
Intervals condition on observed outputs and evaluation samples. Blackvault Ensemble resamples documents; transformation resamples documents within source groups; external results resample records. Blackvault's PrivacyBench interval uses 10,000 document resamples across the full evaluated population. Transformation intervals also condition on fixed judge verdicts.
Methods and prior work.
The references identify the prior work, benchmark definitions, datasets, and published comparisons used in this report.
- [1]GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer
NAACL-HLT, pp. 5364-5376.
- [2]On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation
Journal of Machine Learning Research 11, pp. 2079-2107.
- [3]Bias in error estimation when using cross-validation for model selection
BMC Bioinformatics 7, article 91. doi:10.1186/1471-2105-7-91.
- [4]Bootstrap Methods: Another Look at the Jackknife
The Annals of Statistics 7(1), pp. 1-26. doi:10.1214/aos/1176344552.
- [5]Differential Privacy
ICALP II, LNCS 4052, pp. 1-12. doi:10.1007/11787006_1.
- [6]PrivacyBench: dataset and evaluation implementation
Detection and synthesis scorers. Replay checked at revisions bc120874 and 7f5f131f.
- [7]PII transformation for enterprise datasets
Published reference values, particularly Section 4; accessed September 27, 2026.
- [8]Open PII Masking 500k
Dataset card; English validation cohort in this report.
- [9]Nemotron-PII
Dataset card; test cohort in this report.
- [10]GLiNER-PII model card
Published strict-F1 reference scores; pinned revision bd23e8ef.
- [11]Synthetic PII Finance Multilingual
Dataset card; source of the observed public examples.
- [12]The Text Anonymization Benchmark (TAB): A Dedicated Corpus and Evaluation Framework for Text Anonymization
Computational Linguistics 48(4), pp. 1053-1101. Entity-level coverage evaluation.
- [13]A Neyman-Pearson Approach to Statistical Learning
IEEE Transactions on Information Theory 51(11), pp. 3806-3819. Constrained treatment of error types.
- [14]Which Coreference Evaluation Metric Do You Trust? A Proposal for a Link-based Entity Aware Metric
ACL, pp. 632-642. Analysis of link-based coreference evaluation.
- [15]Jev and the typed probabilistic decision interface
Official documentation. Structured decisions and probability outputs; accessed September 27, 2026.
- [16]Deep Learning: Optimization for Training Deep Models
MIT Press, Chapter 8. Gradients, curvature, and numerical optimization.
- [17]The Foundations of Cost-Sensitive Learning
IJCAI. Decision costs and optimal classification thresholds.
- [18]Enforcing Transitivity in Coreference Resolution
ACL-HLT Short Papers, pp. 45–48. Consistency constraints for identity assignments.
- [19]On Calibration of Modern Neural Networks
ICML, Proceedings of Machine Learning Research 70, pp. 1321–1330.
- [20]Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Advances in Neural Information Processing Systems 31. Parameter-gradient coupling of model outputs.
Research snapshot: September 27, 2026. Scores are percentages; differences are percentage points. Results are rounded directly from the corresponding measurements.