Skip to research
Nyne.aiResearch
NYNE RESEARCH / PII BENCHMARK LEADERSHIP

Benchmark-Leading PII Detection
for Enterprise AI

Nyne Blackvault ranks #1 on PrivacyBench. Explore the full field below.

THE FULL COMPARISON · 8 SYSTEMS

PrivacyBench detection

Exact-match F1 (%) Higher is better

#1 NYNE93.86%

+0.46 points vs micro1's published 93.40

PrivacyBench detection: Nyne Blackvault ranks first in exact-match f11. Nyne Blackvault: 93.86%. Measured by Nyne. 95% CI 93.37 to 94.32. 2. micro1 flow-transform 1.0: 93.40%. Published reference. 3. Tonic Textual: 92.30%. Published reference. 4. Claude Opus 4.8: 89.50%. Published reference. 5. Claude Sonnet 4.6: 88.50%. Published reference. 6. Microsoft Presidio: 86.20%. Published reference. 7. Claude Haiku 4.5: 85.50%. Published reference. 8. GLiNER2 / NeMo Anonymizer: 84.40%. Published reference.0255075100Nyne Blackvault: 93.86%; Measured by Nyne; 95% CI 93.37 to 94.3293.86#1NyneBlackvaultmicro1 flow-transform 1.0: 93.40%; Published reference93.40#2micro1flow-transform 1.0Tonic Textual: 92.30%; Published reference92.30#3TonicTextualClaude Opus 4.8: 89.50%; Published reference89.50#4ClaudeOpus 4.8Claude Sonnet 4.6: 88.50%; Published reference88.50#5ClaudeSonnet 4.6Microsoft Presidio: 86.20%; Published reference86.20#6MicrosoftPresidioClaude Haiku 4.5: 85.50%; Published reference85.50#7ClaudeHaiku 4.5GLiNER2 / NeMo Anonymizer: 84.40%; Published reference84.40#8GLiNER2 /NeMo Anonymizer
  1. 1Nyne Blackvault93.86
  2. 2micro1 flow-transform 1.093.40
  3. 3Tonic Textual92.30
  4. 4Claude Opus 4.889.50
  5. 5Claude Sonnet 4.688.50
  6. 6Microsoft Presidio86.20
  7. 7Claude Haiku 4.585.50
  8. 8GLiNER2 / NeMo Anonymizer84.40
Scale 0-100 Whiskers: 95% CI Methods & full results

Nyne measurements and seven published references reported by micro1. Published comparisons are unpaired; ranks compare point estimates. Comparison scope

THE NYNE ADVANTAGE

Identity expertise.
At global scale.

Privacy research benefits from understanding how people, organizations, and identifiers connect. Our identity graph gives Nyne a distinctive foundation for that work.

People in the graph
2.2B
Businesses in the graph
380M

Working on identity resolution at this scale makes name variation, person-versus-company distinctions, and relationships across records central engineering questions. That perspective shapes our research objective: protect complete identifiers while preserving the relationships that make data useful.

The graph also helps us find relevant organizations and data owners, connecting research questions to domain expertise and potential licensed data partnerships.

Explore the Nyne identity graph Company graph coverage, not the size of a benchmark or a disclosed training corpus.
ABSTRACT

De-identification requires more than recognizing an entity: a system must localize its complete surface form and replace it without corrupting the relationships that make a corpus useful. We study these requirements as separate, measurable objectives: typed span extraction, sensitive-character coverage, and coherent transformation.

We evaluate Nyne Blackvault Ensemble, an ensemble, and Nyne Blackvault, a single model. On all 5,198 human-annotated PrivacyBench messages, Blackvault Ensemble obtains 94.17 exact F1 (95% document-bootstrap CI 93.70 to 94.64), above micro1's published 93.40. Blackvault obtains 93.86 exact F1 and also achieves 92.11 and 93.41 strict typed-span F1 on 500-record AI4Privacy and Nemotron-PII cohorts. The evaluated Blackvault transformation configuration records 97.09 combined detection-and-coherence accuracy.

We report benchmark comparisons with explicit metric definitions and uncertainty estimates. Exact localization, broad native-label evaluation, and coherent replacement are measured separately, so each result answers a defined privacy-engineering question.

Protect the identity.
Keep the story connected.

Enterprise information is relational. A person appears in a message, an email address, an approval, and a company record. Removing every identifier can erase the distinctions a useful dataset needs. Replacing each mention independently can produce fluent text while turning one person into several unrelated fictional people.

We study two coupled requirements: complete, correctly typed localization and coherent replacement across related mentions. The first asks whether the right characters were found. The second asks whether the transformed corpus still expresses the intended relationships. They require different measurements, and neither is established by a single detection score.

FUNCTIONAL VIEWFrom input text to measured transformation
  1. 01

    Define the input

    Keep the document inventory, original character coordinates, and declared entity policy.

  2. 02

    Locate identifiers

    Predict typed spans. Blackvault Ensemble combines detectors; Blackvault is the single-model alternative.

  3. 03

    Connect related mentions

    Use available corpus context to associate names, contact identifiers, and organization variants.

  4. 04

    Transform coherently

    Assign related fictional values and render them in context, retaining relevant relationships.

  5. 05

    Measure the outcome

    Score boundaries, coverage, unnecessary masking, and replacement coherence as separate endpoints.

This view describes the evaluated text workflow. The transformation track uses a recall-oriented detection input; its operating configuration is distinct from the exact-F1 detection headline.

AN AUTHORED THREE-RECORD EXAMPLE

One order. Three records. The same identities.

Source records

Purchase request

Mara Chen sent PO-2048 to Owen Vale at Cedar Labs.

Follow-up message

Owen confirmed PO-2048 with m.chen@example.com.

Delivery record

Cedar Labs will deliver PO-2048 on Tuesday.

Change the identities. Preserve the allowed relationships.

Purchase request

Iris Liu sent PO-2048 to Theo Reed at Orchard Works.

Follow-up message

Theo confirmed PO-2048 with i.liu@example.net.

Delivery record

Orchard Works will deliver PO-2048 on Tuesday.

Mara Chen and the related email become Iris Liu and i.liu@example.net. Owen remains Theo, and Cedar Labs remains Orchard Works. The order and delivery context stay intact.

All three variants are authored illustrations, not recorded model outputs. They explain the objective; the observed comparisons and measured transformation results follow below. Preserving relationships is a utility objective, not by itself a guarantee against re-identification.

What this report contributes

We compare exact and overlap detection against the complete available published field, test broader native label inventories on AI4Privacy and Nemotron-PII, and separate missed detections from incoherent replacements. Explicit metric definitions, uncertainty estimates, and observed output examples explain what each result measures.

What does a correct detection mean?

PII extraction is structured prediction under an annotation contract. A detector can identify the right person but assign the wrong name role, cover only part of an identifier, or merge several entities into one span. These errors have different consequences for privacy and downstream utility. A single overlap score obscures those distinctions.

1.1 Typed spans and label-conditioned extraction

Let xi be a document and ℒ its allowed entity labels. Gold and predicted entities are triples (a, b, ℓ): a half-open interval of original character offsets and an entity type. This fixes the unit of comparison before tokenization, normalization, or replacement.

Gi,Pi⊆{(a,b,ℓ):0≤a<b≤∣xi∣, ℓ∈L},Pi=fθ(xi,L).\begin{aligned}\mathcal G_i,\mathcal P_i&\subseteq\{(a,b,\ell):0\leq a<b\leq |x_i|,\ \ell\in\mathcal L\},\\\mathcal P_i&=f_\theta(x_i,\mathcal L).\end{aligned}

Prior work on label-conditioned extraction, including GLiNER [1], formulates entity detection through contextual spans and entity descriptions. This perspective makes label semantics central to measurement: a complete personal name, a given name, and an organization are different annotation targets. Our evaluation retains each cohort's declared label inventory and states the matching rule explicitly.

1.2 Official overlap and exact-match estimands

PrivacyBench's official scorer [6] uses two directional overlap counts. Let G and N denote total gold and predicted spans. D counts gold spans touched by a same-label prediction; M counts predictions touching a same-label gold span. Write g ∼ p when their labels agree and their character intervals have a nonempty intersection:

D=∑i∑g∈Gi1{∃p∈Pi:g∼p},M=∑i∑p∈Pi1{∃g∈Gi:g∼p},Pov=MN,Rov=DG,F1,ov=2PovRovPov+Rov.\begin{aligned}D&=\sum_i\sum_{g\in\mathcal G_i}\mathbf1\{\exists p\in\mathcal P_i:g\sim p\},\\M&=\sum_i\sum_{p\in\mathcal P_i}\mathbf1\{\exists g\in\mathcal G_i:g\sim p\},\\P_{\rm ov}&=\frac M N,\quad R_{\rm ov}=\frac D G,\quad F_{1,\rm ov}=\frac{2P_{\rm ov}R_{\rm ov}}{P_{\rm ov}+R_{\rm ov}}.\end{aligned}

Overlap matching is not one-to-one: D and M can differ. Exact scoring instead requires the complete typed triple. The exact metric therefore requires both the correct label and complete character boundaries:

X=∑i∑g∈Gi1{g∈Pi},Pex=XN,Rex=XG,F1,ex=2XN+G.\begin{aligned}X&=\sum_i\sum_{g\in\mathcal G_i}\mathbf1\{g\in\mathcal P_i\},\\P_{\rm ex}&=\frac X N,\quad R_{\rm ex}=\frac X G,\quad F_{1,\rm ex}=\frac{2X}{N+G}.\end{aligned}
Figure 01 / An illustrative construction

Overlap is not complete coverage.

Consider two sensitive spans and one prediction with the same entity label. The prediction touches both spans, but covers only one character in each.

Overlap F1
100%
Both gold spans are touched; the prediction touches gold.
Exact F1
0%
Neither gold span has an exact boundary match.
Sensitive-character recall
25%
Only 2 of 8 sensitive characters are covered.
Invented intervals. No benchmark text. Under Tonic’s independent overlap matching, gold-side recall and prediction-side precision are both 100% in this construction. That does not imply that every sensitive character is covered. These are illustrative values, not measured model results.

1.3 From entity recognition to masking coverage

For privacy diagnostics, let UiG and UiP be unions of gold and predicted character intervals, ignoring type. Sensitive-character recall measures what a redaction would cover; excess masking measures the fraction of annotated non-sensitive characters it would remove:

Rchar=∑i∣UiG∩UiP∣∑i∣UiG∣,Emask=∑i∣UiP∖UiG∣∑i(∣xi∣−∣UiG∣).\begin{aligned}R_{\rm char}&=\frac{\sum_i|U_i^G\cap U_i^P|}{\sum_i|U_i^G|},\\E_{\rm mask}&=\frac{\sum_i|U_i^P\setminus U_i^G|}{\sum_i(|x_i|-|U_i^G|)}.\end{aligned}

These diagnostics separate missed sensitive content from loss of useful context. They inform acceptance checks independently of leaderboard F1. They are coverage proxies; coherent replacement is a distinct endpoint, defined in Section 6.

1.4 From span recall to complete identity coverage

An identity can recur across many records. Masking nine mentions while leaving a tenth intact can still reveal the person behind the corpus. The Text Anonymization Benchmark makes this distinction explicit by evaluating coverage at the entity level [12]. Let ℰ be the nonempty set of annotated entities, and let Me be the nonempty set of mentions of entity e. The binary indicator zj equals one if and only if mention j is completely covered. Complete-entity coverage is:

Ce=∏j∈Mezj,Rentity=1∣E∣∑e∈ECe.\begin{aligned}C_e&=\prod_{j\in M_e}z_j,\\R_{\rm entity}&=\frac{1}{|\mathcal E|}\sum_{e\in\mathcal E}C_e.\end{aligned}

The product makes coverage conjunctive: every mention must be covered. Let pj = Pr(zj = 0), the probability of failing to cover mention j completely. The chance of at least one coverage failure obeys:

max⁡j∈Mepj≤Pr⁡(Ce=0)≤min⁡ ⁣(1,∑j∈Mepj).\begin{aligned}\max_{j\in M_e}p_j&\leq\Pr(C_e=0)\\&\leq\min\!\left(1,\sum_{j\in M_e}p_j\right).\end{aligned}

These bounds require no independence assumption. Repeated aliases and shared formatting can correlate errors, so multiplying independent success probabilities is generally unjustified. This is a statement about annotated coverage, not a re-identification probability; F1 cannot be substituted for the individual miss probabilities. It explains why Nyne treats identity continuity as a corpus-level problem.

1.5 Preserve context under a coverage requirement

Removing an entire document can maximize masking coverage while destroying its value. A more useful objective is to minimize excess masking subject to a declared missed-character budget ε. For a masking policy T, Equation 4's coverage and excess-masking quantities define:

Tε={T:1−Rchar(T)≤ε},T⋆∈arg min⁡T∈Tε Emask(T).\begin{aligned}\mathcal T_\varepsilon&=\{T:1-R_{\rm char}(T)\leq\varepsilon\},\\T^\star&\in\underset{T\in\mathcal T_\varepsilon}{\operatorname{arg\,min}}\ E_{\rm mask}(T).\end{aligned}

This separates the privacy requirement from the utility cost instead of hiding both inside one weighted average. Constrained treatment of asymmetric errors has a formal precedent in Neyman-Pearson statistical learning [13]. Here, excess masking is a proxy for context loss; downstream task utility remains a separate question. For a nonempty feasible set with an attained minimum, the optimizer states the preferred policy. These equations describe analytical requirements, not additional measured results or a claim that the optimization has been solved.

1.6 From metric definitions to observed behavior

OBSERVED MODEL OUTPUTS / SAME INPUT, SAME PROMPTS

What Nyne gets right
that the baseline misses.

Two recorded contrasts between Nyne Blackvault and untuned NVIDIA GLiNER-PII, drawn from the public synthetic finance probe.

EXCERPT FROM THE EVALUATED DOCUMENTFpML
<fpml:name>Gallet-Bonneau</fpml:name>

Annotated target: Family nameOriginal offsets [962, 976)

NYNE BLACKVAULT
Exact surname recovered ACTUAL EMITTED SPANS
  • Gallet-Bonneaulast name

The prediction matches the entire annotated surname and its type. All 14 characters are covered.

NVIDIA GLINER-PII · UNTUNED
Surname left undetected ACTUAL EMITTED SPANS
  • No prediction overlaps this surname.

A redaction based on these predictions would leave the complete surname visible in this field.

HOW THIS CASE IS SCORED

For this annotated surname, Blackvault contributes one true positive and the baseline contributes one false negative under strict typed matching. Name merging does not change this target's result.

WHY IT MATTERS TO NYNE

Privacy-sensitive names occur inside schemas and machine-readable records, not only conversational prose. This case shows why we inspect complete span coverage: recognizing nearby entities is insufficient if one intact identifier remains.

Whole-record scores and scoring convention

These scores use all 3 annotated entities in the complete 1,204-character input, not only the excerpt. Both systems used the same mapped prompt list and operating point.

Full-record F1 (%) and counts, by scoring convention
SystemReadingTP / FP / FNF1
Nyne BlackvaultRaw strict2 / 5 / 140.00
Nyne BlackvaultName-merged strict2 / 5 / 140.00
Untuned NVIDIA GLiNER-PIIRaw strict1 / 7 / 218.18
Untuned NVIDIA GLiNER-PIIName-merged strict2 / 4 / 144.44

Blackvault identifies this surname, but a different annotated name elsewhere does not receive a complete exact match. Under the declared name-merged full-record reading, both systems match two of three annotations. Blackvault has five unmatched predictions versus four for the baseline, so Blackvault's full-record F1 is lower despite its success on the displayed surname.

Name merging accommodates the source's full-name annotation convention by joining adjacent predicted given names and surnames using the same declared rule for both systems. It does not change gold annotations. The baseline is pinned to NVIDIA revision bd23e8ef.

Explore a harder iteration: four predictions, their scores, and why the difference matters
WORKED EXAMPLE / EXPLORE THE SCORING

One sentence. Four detection outcomes.

Follow a difficult boundary repair through overcorrection, a name-role error, and exact localization.

Invented text and predictions, not recorded model outputs
REFERENCE ANNOTATION5 SPANS · NAMES + EMAIL

Forward ChenFamily name's note to MaraGiven name at m.chen+finance@example.comEmail; MaraGiven name ChenFamily name owns invoice #2048.

For this example, all name mentions refer to the same fictional person. Names and the complete email are sensitive; the possessive ending, surrounding prose, and invoice number are ordinary context under this stated annotation policy.

HYPOTHESIS 1 / 4

Every entity is touched. The email is still incomplete.

WHAT THE PREDICTED SPANS WOULD COVER

Forward Chen's note to Mara at m.chen+finance@example.com; Mara Chen owns invoice #2048.

Covered sensitive textUncovered sensitive textContext removed
Overlap F1 ↑
100%
M = 5, D = 5
Exact F1 ↑
80%
2 × 4 / (5 + 5)
Character recall ↑
52.38%
22 / 42 sensitive characters
Excess masking ↓
0%
0 / 46 ordinary characters

This hypothesis emits 5 spans against five annotations (N = 5, G = 5). The arrows show the preferred direction. Scores follow Equations 2-4.

All four name mentions match exactly. The email prediction stops at m.chen, leaving +finance@example.com outside the mask. Because that fragment overlaps the annotated email with the right type, overlap F1 still awards full credit.

IF THESE SPANS WERE REDACTED

Forward [PII]'s note to [PII] at [PII]+finance@example.com; [PII] [PII] owns invoice #2048.

WHAT THE NEXT ITERATION MUST SOLVEThe next hypothesis expands the email boundary. But how far should it expand?

Inspect the exact spans and count calculation

Offsets use the original 88-character sentence and half-open intervals [start, end). Overlapping predictions count separately for entity metrics; character coverage takes their union.

Authored gold and predicted spans
Reference spanPredictionExact match
ChenFamily name · [8, 12)ChenFamily name · [8, 12)Yes
MaraGiven name · [23, 27)MaraGiven name · [23, 27)Yes
m.chen+finance@example.comEmail · [31, 57)m.chenEmail · [31, 37)No
MaraGiven name · [59, 63)MaraGiven name · [59, 63)Yes
ChenFamily name · [64, 68)ChenFamily name · [64, 68)Yes

Gold-side overlap hits D = 5; prediction-side overlap hits M = 5; exact matches X = 4. Overlap precision = 5/5 and recall = 5/5. Exact precision = 4/5; exact recall = 4/5.

WHY IT MATTERS

Protect the identifier. Preserve the relationship.

A useful transformed record should still tell us who owns the invoice. After correct detection, a consistent fictional replacement could read:

Forward Vale's note to Iris at i.vale+finance@example.net; Iris Vale owns invoice #2048.

Names and email now belong to one synthetic identity. Replacing each mention independently could sever that connection. This is an illustrative coherence target, not a judged synthesis result.

WHY NYNE MEASURES IT

A better number must represent a better transformation.

We evaluate exact boundaries, sensitive-character coverage, excess masking, and coherent replacement separately because each reveals a different failure. An iteration should address a diagnosed error and then satisfy declared acceptance criteria.

From candidate spans
to coherent replacement.

Consider “Mara Chen signed for Cedar Labs.” Detecting “Mara” is not equivalent to detecting “Mara Chen,” and replacing either string does not establish that a later email belongs to the same person. The system has to carry a decision through three representations: scores over candidate spans, a set of accepted mentions, and identities that remain consistent after replacement. The formulation below follows that chain.

2.1 Make the complete identifier win

Nyne Blackvault uses a fine-tuned GLiNER-PII detector. Following GLiNER's architecture, contextual word and label representations are projected into a shared space [1][10]. Let hi and vℓ denote those contextual representations, with learned functions producing the span vector sij and label vector eℓ. Stack n span vectors and k label vectors in a shared dimension d:

sij=gθ([ϕθ(hi);ψθ(hj)]),eℓ=qθ(vℓ),S∈Rn×d,E∈Rk×d,Z=SET,P=σ(Z),σ(z)=11+e−z.\begin{aligned}s_{ij}&=g_\theta([\phi_\theta(h_i);\psi_\theta(h_j)]),\quad e_\ell=q_\theta(v_\ell),\\S&\in\mathbb R^{n\times d},\quad E\in\mathbb R^{k\times d},\\Z&=SE^{\mathsf T},\qquad P=\sigma(Z),\quad \sigma(z)=\frac{1}{1+e^{-z}}.\end{aligned}

“Mara Chen” and “Mara” occupy different rows of S but share the PERSON label vector. Their scores depend on much of the same context and many of the same parameters. The sigmoid acts independently on each pair; a high PERSON score for the fragment does not suppress the full name. Their competition is resolved later.

To see how learning changes that competition, fix the input and candidate inventory and flatten Z in column order into z. Let Gθ = ∂z/∂θT be the score Jacobian, and let r = ∇zJ be the residual of a scoring loss J(z). A local gradient-descent update gives, for sufficiently small η > 0:

θ+=θ−ηGθTr,Kθ=GθGθT⪰0,z+=z−ηKθr+O(η2),J(z+)−J(z)=−ηrTKθr+O(η2).\begin{aligned}\theta^+&=\theta-\eta G_\theta^{\mathsf T}r,\qquad K_\theta=G_\theta G_\theta^{\mathsf T}\succeq0,\\z^+&=z-\eta K_\theta r+O(\eta^2),\\J(z^+)-J(z)&=-\eta r^{\mathsf T}K_\theta r+O(\eta^2).\end{aligned}

Here the score map and loss are twice continuously differentiable locally, with bounded second derivatives; all derivatives on the right are evaluated before the update. Each entry Kab is the inner product of two candidates' parameter gradients. Its off-diagonal terms describe how changing one candidate affects another. This finite-network tangent-kernel view [20] analyzes the coupling; it does not require materializing a kernel in the detector. The displayed step is a local reference update rather than a specification of the optimizer [16].

The boundary decision depends on a contrast within that coupled system. Let a be the correctly typed full name and b an overlapping fragment. Define vab as the vector with +1 at a, −1 at b, and zeros elsewhere. The ranking margin Δab and eligibility margin γa are:

Δab=vabTz=za−zb,Δab+=Δab−ηvabTKθr+O(η2),γa=za−λa,λa=log⁡τa1−τa.\begin{aligned}\Delta_{ab}&=v_{ab}^{\mathsf T}z=z_a-z_b,\\\Delta_{ab}^+&=\Delta_{ab}-\eta v_{ab}^{\mathsf T}K_\theta r+O(\eta^2),\\\gamma_a&=z_a-\lambda_a,\qquad\lambda_a=\log\frac{\tau_a}{1-\tau_a}.\end{aligned}

For 0 < τa < 1, γa says whether the full span clears its score cutoff; Δab says whether it outranks this fragment. Although the first-order change in the loss in Equation 9 is nonpositive, the change in this particular ranking margin can have either sign. A better aggregate objective can therefore coexist with a worse boundary decision. Exact-span evaluation tests the decision that survives this competition, including competing candidates beyond this pair.

2.2 Choose which candidates survive

The scores now have to become an admissible set of spans. Map word endpoints back to original character intervals, index typed candidates by a, and put conflicting pairs in ℭ. Under a flat annotation policy, overlap makes a pair incompatible. If ua records whether candidate a is selected, the admissible set is:

U(z)={u∈{0,1}nk:ua=0if za<λa,ua+ub≤1∀(a,b)∈C}.\mathcal U(z)=\left\{u\in\{0,1\}^{nk}:\begin{array}{ll}u_a=0&\text{if }z_a<\lambda_a,\\u_a+u_b\leq1&\forall(a,b)\in\mathcal C\end{array}\right\}.

A greedy decoder visits eligible candidates in score order and retains compatible spans. The full name can lose to a fragment even when both clear the cutoff; conversely, raising the cutoff can remove both. This is the connection between the margins above and the sensitive characters left in the output. The constraints specify feasible selections, without asserting that greedy selection solves a global maximum-score problem.

Interactive illustration / Authored inputs

A high score can still select the wrong boundary.

Change one span vector, then adjust the cutoff to see which characters a simple decoder masks.

Invented text

Mara Chen signed for Cedar Labs.

1. Score each span against each label

Label vectors are PERSON = (1, 0) and ORG = (0, 1). Their dot products with each span vector give the logits below. Apply σ(z) = 1 / (1 + e⁻ᶻ) independently to each cell.

Span vectors and label scores. Highlighted cells are kept by the decoder.
Candidate / vectorPERSONORG
Mara Chen(2.4, -1.2) Kept: PERSON91.68% logit 2.4 Selected23.15% logit -1.2
Mara(1.8, -0.8) Overlap rejected85.81% logit 1.8 31.00% logit -0.8
Cedar Labs(-1.4, 2.2) Kept: ORG19.78% logit -1.4 90.02% logit 2.2 Selected

Sigmoid scores are independent, not a softmax distribution. They need not sum to 100%.

2. Keep high scores, reject overlapping spans

This toy decoder sorts eligible candidate-label pairs by score, highest first, and keeps a pair only if its span does not overlap one already kept.

  1. Mara Chen PERSON · 91.68% Keep
  2. Cedar Labs ORG · 90.02% Keep
  3. Mara PERSON · 85.81% Overlap rejected

3. Apply the selected masks

[PERSON] signed for [ORG].

19/ 19Sensitive characters covered

2 spans selected. 19 of 19 sensitive characters covered.

Both full sensitive spans are covered. Increasing the threshold can discard a correct entity.

Authored illustration, not measured model output. The two sensitive spans contain 19 characters, including internal spaces. The vectors and threshold are chosen to make the scoring and selection mechanics inspectable.

Some decisions depend on the candidate's meaning in context as well as its neural score. The evaluated Blackvault Ensemble includes structured candidate review using TypeSafe's Jev. Jev assesses proposed spans and returns constrained type or non-entity decisions; the surrounding system uses these outputs to retain, reject, or relabel candidates and handle uncertain outcomes. Blackvault performs detection without this review stage. Jev supplies contextual review here, while identity linking and replacement are downstream tasks. [15]

This review can be expressed as a decision over a finite action set 𝒜. Let qk(x, a) be the posterior probability of true class k for candidate a in context x, and let C(α, k) assign the cost of an action α under that class. Then the minimum-risk action is [17]:

α⋆(x,a)∈arg min⁡α∈A∑kC(α,k) qk(x,a),qk(x,a)≥0,∑kqk(x,a)=1.\begin{aligned}\alpha^\star(x,a)&\in\underset{\alpha\in\mathcal A}{\operatorname{arg\,min}}\sum_k C(\alpha,k)\,q_k(x,a),\\q_k(x,a)&\geq0,\qquad\sum_k q_k(x,a)=1.\end{aligned}

Retaining an entity, rejecting non-entity text, changing a type, and deferring a decision have different consequences. The action set and cost matrix express those differences. This is a decision-theoretic interpretation of the interface; the evaluated review policy is not claimed to solve this optimization. Substituting estimated probabilities requires calibration in the decision domain [19]. The span scores in Equation 8 establish neither that calibration nor a probability of disclosure.

2.3 Keep one identity consistent across records

The accepted spans identify where a transformation may act. They do not yet establish which mentions must share a replacement. Suppose “Mara Chen,” “M. Chen,” and a related email have been resolved to one identity. Represent the assignments of m = 1Tu accepted mentions to c identity classes by a binary matrix B, with exactly one nonzero entry in each row. Its induced relation H records which mention pairs share an identity:

B∈{0,1}m×c,B1c=1m,H=BBT,Hij=1{mentions i,j share an identity}.\begin{aligned}B&\in\{0,1\}^{m\times c},\qquad B\mathbf1_c=\mathbf1_m,\\H&=BB^{\mathsf T},\qquad H_{ij}=\mathbf1\{\text{mentions }i,j\text{ share an identity}\}.\end{aligned}

This representation enforces transitivity: if two mentions share a class with a third, they share a class with each other. Independent pairwise decisions need not satisfy that property [18]. The task is to infer the right partition and then preserve it during replacement; a coherent but incorrect partition would consistently repeat the original resolution error.

Let R assign the c original identity classes to s fictional identities, where s ≥ c. Giving every class one surrogate and preventing collisions between classes yields an exact preservation condition:

R∈{0,1}c×s,R1s=1c,RT1c≤1s,B~=BR,RRT=Ic,B~B~T=BRRTBT=BBT=H.\begin{aligned}R&\in\{0,1\}^{c\times s},\quad R\mathbf1_s=\mathbf1_c,\quad R^{\mathsf T}\mathbf1_c\leq\mathbf1_s,\\\widetilde B&=BR,\qquad RR^{\mathsf T}=I_c,\\\widetilde B\widetilde B^{\mathsf T}&=BRR^{\mathsf T}B^{\mathsf T}=BB^{\mathsf T}=H.\end{aligned}

The equality follows because the rows of R are distinct one-hot vectors, so RRT = Ic. It states precisely what “the same fictional person across records” means. A surname and an email can render different attributes of that identity while preserving the relation. If two original classes collide, RRT gains off-diagonal entries and unrelated mentions merge. If one original class receives several surrogates, a single assignment matrix R no longer describes the transformation and its links can split.

Nyne's evaluated transformation stage uses deterministic identity resolution, surrogate assignment, and contextual rendering. Equation 14 expresses the invariant that this process seeks to preserve. The pairwise measures in Section 6 define how to count departures from it as merges and splits [14].

2.4 Apply the decisions to an entire corpus

The same construction scales from one sentence to a corpus 𝒳 = (x1, …, xD). Concatenate only its coordinate inventories: let L be the total number of original character positions and N the total number of typed candidates. Each document has its own contextual span and label matrices; stack their scores and selection vectors in the same document order. A sparse incidence matrix A records which characters each candidate covers. Each column lies entirely inside one document's block. Selection then projects from candidate space back into character space:

zX=col⁡d=1Dvec⁡(SdEdT),N=∑d=1Dndkd,A∈{0,1}L×N,Ata=1{character t lies in candidate a},c(u)=1{Au>0}∈{0,1}L,Tu,B,R(X)=(x~1,…,x~D),with assignments B~=BR.\begin{aligned}z_{\mathcal X}&=\operatorname{col}_{d=1}^D\operatorname{vec}(S_dE_d^{\mathsf T}),\quad N=\sum_{d=1}^D n_dk_d,\\A&\in\{0,1\}^{L\times N},\quad A_{ta}=\mathbf1\{\text{character }t\text{ lies in candidate }a\},\\c(u)&=\mathbf1\{Au>0\}\in\{0,1\}^{L},\\\mathcal T_{u,B,R}(\mathcal X)&=(\widetilde x_1,\ldots,\widetilde x_D),\quad\text{with assignments }\widetilde B=BR.\end{aligned}

The indicator is elementwise: a character is covered if at least one selected candidate contains it. This is the coverage vector underlying Equation 4, now expressed directly as a function of the decisions u. Meanwhile, BR assigns every selected mention a surrogate identity. Together they specify where the renderer acts and which fictional identity supplies each replacement.

For document d, sort its md selected spans by original offsets (adj, bdj, ℓdj). Let i(d, j) identify the corresponding row of B, and let ψ render the assigned surrogate in the entity type and surface form of that mention. Writing ⊕ for string concatenation and setting bd0 = 0, the transformed document is:

x~d=⨁j=1md(xd[bd,j−1:adj] ⊕ψ ⁣(ℓdj,(BR)i(d,j),:,xd[adj:bdj]))⊕ xd[bd,md:∣xd∣].\begin{aligned}\widetilde x_d&=\bigoplus_{j=1}^{m_d}\left(x_d[b_{d,j-1}:a_{dj}]\ \oplus\right.\\&\hspace{2.5em}\left.\psi\!\left(\ell_{dj},(BR)_{i(d,j),:},x_d[a_{dj}:b_{dj}]\right)\right)\\&\qquad\oplus\ x_d[b_{d,m_d}:|x_d|].\end{aligned}

Each untouched gap is copied from the original document; each selected span is rendered from its shared identity assignment. Replacement strings may have different lengths because every slice still refers to the source coordinates. When a document has no selected spans, the empty concatenation leaves the whole document intact. Applying this operator to every document produces the transformed corpus without treating repeated mentions as independent replacement problems. Coverage specifies the selected characters; rendering must also change their sensitive values appropriately.

2.5 What success requires across records

Preserving the accepted-mention partition cannot recover an identifier omitted before that partition was built. Return to all annotated mentions of an identity e. Let Ae mean that every such mention is correctly localized and typed, Be that their identity assignments are correct, and Ce that their replacements change the sensitive values coherently. Complete success requires their intersection:

Pr⁡(Ae∩Be∩Ce)=Pr⁡(Ae) Pr⁡(Be∣Ae)⋅Pr⁡(Ce∣Ae∩Be),Pr⁡(incomplete transformation of e)=1−Pr⁡(Ae∩Be∩Ce).\begin{aligned}\Pr(A_e\cap B_e\cap C_e)&=\Pr(A_e)\,\Pr(B_e\mid A_e)\\&\qquad\cdot\Pr(C_e\mid A_e\cap B_e),\\\Pr(\text{incomplete transformation of }e)&=1-\Pr(A_e\cap B_e\cap C_e).\end{aligned}

This factorization retains the dependence between stages. A boundary error changes what the resolver receives; a resolution error changes which surrogate the renderer uses. Multiplying their marginal accuracies would erase those interactions. If an earlier conditioning event has probability zero, complete success is zero. Otherwise the conditional factors describe exactly where the remaining risk enters.

The chain connects representation learning to the output requirements: complete spans must survive selection, accepted mentions must form the right identity classes, and replacement must preserve those classes while changing sensitive values. These identity-level events define the analytical target. The following results report the measured detection and transformation endpoints, with their respective populations and denominators.

The benchmark, the metric,
and the comparison field.

Evaluation begins with a defined population, annotation policy, and endpoint. Selection can overfit a noisy metric, as analyzed by Cawley and Talbot [2] and Varma and Simon [3]. We therefore interpret these measurements at their stated scope, distinguish exact localization from overlap, and examine errors beyond the headline score.

ENSEMBLE

Nyne Blackvault Ensemble

An ensemble evaluated for high-fidelity PII localization on PrivacyBench.

SINGLE MODEL

Nyne Blackvault

A single detector evaluated across PrivacyBench, AI4Privacy, and Nemotron-PII.

3.1 Benchmark populations

We evaluate all five PrivacyBench entity types, then broaden coverage to the 20 native AI4Privacy types and 55 Nemotron-PII types. Each declared cohort retains negative records and every annotated type. The external evaluations use 500 English validation records from AI4Privacy and 500 test records from Nemotron-PII. [8][9]

Table 1. Benchmark populations
BenchmarkPopulationEntity types
PrivacyBench5,198 messages5 entity types
AI4Privacy500 English validation records20 native label types
Nemotron-PII500 test records55 native label types

3.2 Interpreting a reported score

The primary detection results pool counts across records before computing F1. We report 95% bootstrap intervals for Nyne measurements, following the resampling principle of Efron [4]. These quantify sampling uncertainty in the evaluated outputs. Published competitor results are reproduced from their sources; they are not paired reruns on our records. Detailed comparison scope is collected in the source notes.

#1 in exact F1.
The entire field, compared.

Across all nine systems shown, Blackvault Ensemble ranks first in exact-match F1. On 5,198 messages containing 8,805 human-annotated spans, it obtains 94.17 exact F1, a 0.77-point improvement over micro1's published 93.40. Exact precision is 94.44 and recall is 93.90. Blackvault obtains 93.86 exact F1. Blackvault Ensemble's exact-match interval lies above the published reference; Blackvault's higher point estimate has an interval that includes it. These comparisons use reported reference values rather than paired predictions. [7][P1]

FIGURE 02

PrivacyBench detection

+0.77 points

Blackvault Ensemble leads the displayed field and is above micro1’s published exact-match F1.

PrivacyBench exact F1, Nyne and published reference systemsNyne Blackvault Ensemble: 94.17, 95% interval 93.70 to 94.64. Nyne Blackvault: 93.86, 95% interval 93.37 to 94.32. micro1 flow-transform 1.0: 93.40. Tonic Textual: 92.30. Claude Opus 4.8: 89.50. Claude Sonnet 4.6: 88.50. Microsoft Presidio: 86.20. Claude Haiku 4.5: 85.50. GLiNER2 / NeMo Anonymizer: 84.40.02550751001Nyne Blackvault Ensemble94.172Nyne Blackvault93.863micro1 flow-transform 1.093.404Tonic Textual92.305Claude Opus 4.889.506Claude Sonnet 4.688.507Microsoft Presidio86.208Claude Haiku 4.585.509GLiNER2 / NeMo Anonymizer84.40

F1 / score (%) · scale 0 to 100

  1. 1Nyne Blackvault Ensemble94.17

    95% CI 93.70 to 94.64

  2. 2Nyne Blackvault93.86

    95% CI 93.37 to 94.32

  3. 3micro1 flow-transform 1.093.40
  4. 4Tonic Textual92.30
  5. 5Claude Opus 4.889.50
  6. 6Claude Sonnet 4.688.50
  7. 7Microsoft Presidio86.20
  8. 8Claude Haiku 4.585.50
  9. 9GLiNER2 / NeMo Anonymizer84.40

All available field rows · ranked by point estimate · equal displayed scores share a rank

Nyne measurement Published referenceWhiskers: Nyne 95% bootstrap CI
Figure 2. Pooled official Tonic F1 (%). Ranked, zero-based full-field bars show magnitude and uncertainty across the complete reported field. Reference rows reproduce micro1's published table [7]. Whiskers show 95% bootstrap intervals. Comparison scope: P1.

Blackvault Ensemble's overlap F1 is 96.14 [95.79, 96.48], numerically above micro1's 96.00; its interval includes that reference. Blackvault ties micro1 at the displayed overlap F1 of 96.00. The difference between the two readings is scientifically useful: exact matching demands type and boundary fidelity, while overlap asks whether entities were touched. The stronger exact result therefore supports a localization claim. It does not isolate which component caused the gain without a controlled ablation.

Equal-export macro F1 changes Blackvault Ensemble's exact score from 94.17 to 94.10 and Blackvault's from 93.86 to 93.83. The small shift shows that the pooled result is not driven solely by weighting larger exports more heavily. Precision, recall, and both F1 readings are retained in the complete table below.

Table 2. Full precision, recall and F1; aggregation sensitivity
F1 (%) with 95% intervals where available
SystemExact PExact RExact F195% CIOverlap POverlap ROverlap F195% CI
Nyne Blackvault Ensemble94.4493.9094.1793.70 to 94.6496.4295.8796.1495.79 to 96.48
Nyne Blackvault94.6093.1393.8693.37 to 94.3296.7495.2896.0095.64 to 96.34
micro1 flow-transform 1.094.3492.4893.40Not reported96.9595.0796.00Not reported
Tonic Textual94.3090.5092.30Not reported96.5092.6094.50Not reported
Claude Opus 4.893.1086.1089.50Not reported95.4088.2091.70Not reported
Claude Sonnet 4.694.7083.0088.50Not reported97.0085.0090.60Not reported
Microsoft Presidio85.4087.1086.20Not reported87.3089.0088.10Not reported
Claude Haiku 4.595.0077.7085.50Not reported97.6079.9087.80Not reported
GLiNER2 / NeMo Anonymizer82.8086.1084.40Not reported85.2088.6086.90Not reported

Blackvault Ensemble pooled / macro exact F1: 94.17 / 94.10; overlap: 96.14 / 96.08. Blackvault pooled / macro exact F1: 93.86 / 93.83; overlap: 96.00 / 95.96. micro1 does not specify its aggregation; pooled is assumed for the reference comparison.

Scores are rounded directly from the underlying measurements. The final field row preserves the GLiNER2 label used in micro1's detection table.

Leading scores.
Two broader label inventories.

PrivacyBench's five labels cannot by themselves establish broad PII coverage. We retain all 20 AI4Privacy and 55 Nemotron-PII native types on fixed 500-record cohorts, including 1,237 and 4,019 annotated spans respectively. Blackvault is evaluated on both.

Results use strict typed exact-span F1: a detection must recover both boundaries and the correct entity type. The complete field includes micro1, NeMo Anonymizer, NVIDIA's published GLiNER-PII model-card result, and a GLiNER-PII control measured on our cohorts. Published figures retain the metric stated by their source.

FIGURE 03

External benchmark performance

Strict typed exact-span F1 (%) · All available reference systems

AI4Privacy

500 RECORDS
AI4Privacy, strict exact-span F1Blackvault: 92.11, 95% interval 90.39 to 93.71. micro1 flow-transform 1.0: 83.74. NeMo Anonymizer: 73.08. NVIDIA GLiNER-PII (same cohort): 64.36, 95% interval 61.41 to 67.17. NVIDIA gliner-PII (card): 64.00.02550751001BlackvaultNyne measurement92.112micro1 flow-transform 1.0Published reference83.743NeMo AnonymizerPublished reference73.084NVIDIA GLiNER-PII (same cohort)Same-cohort baseline64.365NVIDIA gliner-PII (card)Published reference64.00

F1 / score (%) · scale 0 to 100

  1. 1Blackvault92.11

    Nyne measurement

    95% CI 90.39 to 93.71

  2. 2micro1 flow-transform 1.083.74

    Published reference

  3. 3NeMo Anonymizer73.08

    Published reference

  4. 4NVIDIA GLiNER-PII (same cohort)64.36

    Same-cohort baseline

    95% CI 61.41 to 67.17

  5. 5NVIDIA gliner-PII (card)64.00

    Published reference

All available field rows · ranked by point estimate · equal displayed scores share a rank

Nemotron-PII

500 RECORDS
Nemotron-PII, strict exact-span F1Blackvault: 93.41, 95% interval 92.42 to 94.37. NVIDIA GLiNER-PII (same cohort): 89.66, 95% interval 88.48 to 90.77. micro1 flow-transform 1.0: 88.33. NeMo Anonymizer: 87.84. NVIDIA gliner-PII (card): 87.00.02550751001BlackvaultNyne measurement93.412NVIDIA GLiNER-PII (same cohort)Same-cohort baseline89.663micro1 flow-transform 1.0Published reference88.334NeMo AnonymizerPublished reference87.845NVIDIA gliner-PII (card)Published reference87.00

F1 / score (%) · scale 0 to 100

  1. 1Blackvault93.41

    Nyne measurement

    95% CI 92.42 to 94.37

  2. 2NVIDIA GLiNER-PII (same cohort)89.66

    Same-cohort baseline

    95% CI 88.48 to 90.77

  3. 3micro1 flow-transform 1.088.33

    Published reference

  4. 4NeMo Anonymizer87.84

    Published reference

  5. 5NVIDIA gliner-PII (card)87.00

    Published reference

All available field rows · ranked by point estimate · equal displayed scores share a rank

Nyne measurement Same-cohort baseline Published referenceWhiskers: 95% CI
Figure 3. Nyne and the measured NVIDIA control use the same 500-record cohorts. Published rows are source-reported references; their samples and implementations may differ. Whiskers show 95% intervals where available. Comparison scope.

5.1 Broad PII coverage

Under the strict reading, Blackvault reaches 92.11 [90.39, 93.71] on AI4Privacy and 93.41 [92.42, 94.37] on Nemotron-PII, the highest strict scores in each displayed comparison.

micro1 reports strict-span F1 of 83.74 on AI4Privacy and 88.33 on Nemotron; its NeMo references are 73.08 and 87.84. We preserve that stated metric. The numerical comparison is informative, while the directly measured NVIDIA baseline provides the stronger control on population and scoring. Published rows are not paired reruns. [P2]

Table 3. Full external benchmark results
AI4Privacy · F1 (%)
SystemStrict F195% CI
NVIDIA GLiNER-PII (same cohort)64.3661.41 to 67.17
Nyne Blackvault92.1190.39 to 93.71
micro1 flow-transform 1.083.74Not reported
NeMo Anonymizer73.08Not reported
NVIDIA gliner-PII model card64.00Not reported
Nemotron-PII · F1 (%)
SystemStrict F195% CI
NVIDIA GLiNER-PII (same cohort)89.6688.48 to 90.77
Nyne Blackvault93.4192.42 to 94.37
micro1 flow-transform 1.088.33Not reported
NeMo Anonymizer87.84Not reported
NVIDIA gliner-PII model card87.00Not reported

Rows labelled published reference retain the source's stated strict metric. The measured NVIDIA baseline uses the same cohort and declared scorer as Nyne.

Detection and replacement
are coupled objectives.

Replacing each detected string independently can preserve local readability while breaking corpus-level identity relationships. Coherent transformation must keep a person's related names, handles, and email addresses consistent. We evaluate replacement quality separately from detection to distinguish missed entities from failures to transform a detected entity correctly.

6.1 A decomposable endpoint

Tonic's synthesis scorer selects the same-label prediction with the largest intersection for each gold span. Let G be annotated spans, D those detected, and C those changed and judged coherent. Unchanged replacements receive no credit; missing verdicts remain in the denominator. The three reported quantities are:

RNER=DG,Asynth=CD,Qcombined=CG=RNER Asynth.\begin{aligned}R_{\rm NER}&=\frac D G,\qquad A_{\rm synth}=\frac C D,\\Q_{\rm combined}&=\frac C G=R_{\rm NER}\,A_{\rm synth}.\end{aligned}

This factorization distinguishes a missed detection from an incoherent replacement. Equivalently, the combined error can be partitioned into two disjoint contributions:

1−Qcombined=(1−RNER)⏟missed+RNER(1−Asynth)⏟incoherent.1-Q_{\rm combined}=\underbrace{(1-R_{\rm NER})}_{\text{missed}}+\underbrace{R_{\rm NER}(1-A_{\rm synth})}_{\text{incoherent}}.

For the evaluated Blackvault transformation configuration, recall is 98.07% and conditional accuracy is 99.00%. Using the rounded rates, about 1.93 percentage points of annotated spans are missed and 0.98 points are detected without coherent transformation. The residual is about 2.91%, consistent with the measured 97.09% combined result. This endpoint has no precision term because the generated annotation set is not exhaustive.

6.2 Identity links: two distinct failure modes

Local fluency cannot establish relational consistency. On a fixed set of aligned mentions, let A be the set of unordered mention pairs that share an original identity, and let BT contain pairs assigned the same fictional identity. Pairwise link precision and recall distinguish accidental identity merges from identity splits:

Plink=∣A∩BT∣∣BT∣,Rlink=∣A∩BT∣∣A∣,Nmerge=∣BT∖A∣,Nsplit=∣A∖BT∣.\begin{aligned}P_{\rm link}&=\frac{|A\cap B_T|}{|B_T|},&R_{\rm link}&=\frac{|A\cap B_T|}{|A|},\\N_{\rm merge}&=|B_T\setminus A|,&N_{\rm split}&=|A\setminus B_T|.\end{aligned}

These errors count mention pairs: a merge invents a relationship between different people; a split breaks the connection between two mentions of one person. Pairwise evaluation is established in coreference research [14]. Its pair counts give larger identity clusters more weight; singleton entities and missing mentions need separate coverage accounting. For nonempty pair sets, the equations state a corpus-level consistency objective. The judged transformation scores below measure coherent replacements, not this separate link metric.

FIGURE 04

The highest combined score in the reported field

97.09%

Blackvault + transformation
adaptive after two judged looks
#1 combined score in this comparison

PrivacyBench combined transformation scoreBlackvault + synthesis v1.6: 97.09, 95% interval 96.68 to 97.49. micro1 flow-transform 1.0: 95.46. Tonic Textual + Opus 4.8: 92.00.02550751001Blackvault + synthesis v1.6adaptive after two judged looks97.092micro1 flow-transform 1.0published; development exposure undisclosed95.463Tonic Textual + Opus 4.8published92.00

F1 / score (%) · scale 0 to 100

  1. 1Blackvault + synthesis v1.697.09

    adaptive after two judged looks

    95% CI 96.68 to 97.49

  2. 2micro1 flow-transform 1.095.46

    published; development exposure undisclosed

  3. 3Tonic Textual + Opus 4.892.00

    published

All available field rows · ranked by point estimate · equal displayed scores share a rank

Figure 4. Nyne transformation quality on 8,253 annotated spans, alongside both published comparison systems. Combined quality is the product of detection recall and conditional synthesis accuracy (Equation 18). Detection and transformation are evaluated as distinct endpoints.

6.3 Conditional accuracy and end-to-end success

The evaluated Blackvault transformation configuration records 99.00 [98.76, 99.23] synthesis accuracy and 97.09 [96.68, 97.49] combined (adaptive after two judged looks), alongside micro1's published 98.28 and 95.46. The useful distinction is between conditional replacement quality and total annotated-span success: a coherent replacement cannot recover a span that was never detected.

We use Tonic's synthesis evaluator to assess whether replacements change the source value and remain coherent, with Claude Opus 4.7 providing the stored judgments. Confidence intervals resample document counts while holding those judgments fixed. They describe the reported configuration and do not capture repeated-judge variability. The comparator values are published references, with potentially different evaluation populations and judge configurations.

Table 4. Transformation results
Transformation scores (%)
SystemNER recallSynthesisCombined
Nyne Blackvault + synthesis v1.6adaptive after two judged looks98.0799.0097.09
micro1 flow-transform 1.0published; development exposure undisclosed97.1398.2895.46
Tonic Textual + Claude Opus 4.8published95.0097.0092.00

This task uses 8,253 generated-gold spans, a different denominator from the 8,805 human-annotated detection spans. Synthesis scores do not establish enterprise-wide Transformation Quality Index (TQI), re-identification resistance or guarantees of anonymity.

What the evidence supports.

Three findings emerge. First, exact localization provides the strongest PrivacyBench comparison: Blackvault Ensemble's point estimate and interval exceed the published reference. Second, Blackvault supports broader native-label inventories, with competitive strict scores on AI4Privacy and Nemotron-PII. Third, the transformation evaluation measures detection and coherent replacement together.

Our endpoint is empirical extraction and transformation quality. Differential privacy, for example, constrains output distributions on neighboring datasets, a different mathematical object [5]. The measurements here concern detection, coverage, and coherence; they do not establish such a mechanism-level guarantee.

Source and measurement notes

Three short notes define the scope of the reference comparisons and uncertainty estimates.

P1Published reference comparisons

PrivacyBench field rows and external micro1/NVIDIA model-card references are published values, not paired reruns. Exact samples, aggregation and complete scoring details are not fully specified. Transformation populations and judge configurations may also differ. Nyne intervals describe its measured scores, not paired differences against these references. Both Nyne overlap F1 intervals include micro1's published value.

P2External scoring readings

Strict typed exact-span F1 is the primary measured endpoint. micro1 and NeMo published values retain their source's stated strict-span metric; their exact samples and implementations are not verified here. The Nyne-executed NVIDIA control uses the same 500-record cohorts and native label inventories as Nyne.

P3Conditional uncertainty

Intervals condition on observed outputs and evaluation samples. Blackvault Ensemble resamples documents; transformation resamples documents within source groups; external results resample records. Blackvault's PrivacyBench interval uses 10,000 document resamples across the full evaluated population. Transformation intervals also condition on fixed judge verdicts.

7.1 Author and research responsibility

Michael Fanous is the author and sole researcher for this study, at Nyne Research.

Research report v1.4, September 27, 2026. The tables, metric definitions, and cited sources describe the basis of the reported comparisons.

Methods and prior work.

The references identify the prior work, benchmark definitions, datasets, and published comparisons used in this report.

  1. [1]

    Zaratiana, U., Tomeh, N., Holat, P., and Charnois, T. (2024).

    GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer

    NAACL-HLT, pp. 5364-5376.

  2. [2]

    Cawley, G. C., and Talbot, N. L. C. (2010).

    On Over-fitting in Model Selection and Subsequent Selection Bias in Performance Evaluation

    Journal of Machine Learning Research 11, pp. 2079-2107.

  3. [3]

    Varma, S., and Simon, R. (2006).

    Bias in error estimation when using cross-validation for model selection

    BMC Bioinformatics 7, article 91. doi:10.1186/1471-2105-7-91.

  4. [4]

    Efron, B. (1979).

    Bootstrap Methods: Another Look at the Jackknife

    The Annals of Statistics 7(1), pp. 1-26. doi:10.1214/aos/1176344552.

  5. [5]

    Dwork, C. (2006).

    Differential Privacy

    ICALP II, LNCS 4052, pp. 1-12. doi:10.1007/11787006_1.

  6. [6]

    Tonic AI.

    PrivacyBench: dataset and evaluation implementation

    Detection and synthesis scorers. Replay checked at revisions bc120874 and 7f5f131f.

  7. [7]

    micro1.

    PII transformation for enterprise datasets

    Published reference values, particularly Section 4; accessed September 27, 2026.

  8. [8]

    AI4Privacy.

    Open PII Masking 500k

    Dataset card; English validation cohort in this report.

  9. [9]

    NVIDIA.

    Nemotron-PII

    Dataset card; test cohort in this report.

  10. [10]

    NVIDIA.

    GLiNER-PII model card

    Published strict-F1 reference scores; pinned revision bd23e8ef.

  11. [11]

    Gretel.

    Synthetic PII Finance Multilingual

    Dataset card; source of the observed public examples.

  12. [12]

    Pilán, I., et al. (2022).

    The Text Anonymization Benchmark (TAB): A Dedicated Corpus and Evaluation Framework for Text Anonymization

    Computational Linguistics 48(4), pp. 1053-1101. Entity-level coverage evaluation.

  13. [13]

    Scott, C., and Nowak, R. (2005).

    A Neyman-Pearson Approach to Statistical Learning

    IEEE Transactions on Information Theory 51(11), pp. 3806-3819. Constrained treatment of error types.

  14. [14]

    Moosavi, N. S., and Strube, M. (2016).

    Which Coreference Evaluation Metric Do You Trust? A Proposal for a Link-based Entity Aware Metric

    ACL, pp. 632-642. Analysis of link-based coreference evaluation.

  15. [15]

    TypeSafe AI. (2026).

    Jev and the typed probabilistic decision interface

    Official documentation. Structured decisions and probability outputs; accessed September 27, 2026.

  16. [16]

    Goodfellow, I., Bengio, Y., and Courville, A. (2016).

    Deep Learning: Optimization for Training Deep Models

    MIT Press, Chapter 8. Gradients, curvature, and numerical optimization.

  17. [17]

    Elkan, C. (2001).

    The Foundations of Cost-Sensitive Learning

    IJCAI. Decision costs and optimal classification thresholds.

  18. [18]

    Finkel, J. R., and Manning, C. D. (2008).

    Enforcing Transitivity in Coreference Resolution

    ACL-HLT Short Papers, pp. 45–48. Consistency constraints for identity assignments.

  19. [19]

    Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. (2017).

    On Calibration of Modern Neural Networks

    ICML, Proceedings of Machine Learning Research 70, pp. 1321–1330.

  20. [20]

    Jacot, A., Gabriel, F., and Hongler, C. (2018).

    Neural Tangent Kernel: Convergence and Generalization in Neural Networks

    Advances in Neural Information Processing Systems 31. Parameter-gradient coupling of model outputs.

Research snapshot: September 27, 2026. Scores are percentages; differences are percentage points. Results are rounded directly from the corresponding measurements.

NYNE RESEARCH

Evaluate on the data
your system will encounter.

Discuss an evaluation for your domain, annotation policy, and transformation requirements.

Contact the research team ↗