How Private is Private? A Comparative Study for Face De-Identification

A hierarchical, consistency-based metric HiFD and a demographically balanced benchmark UtilFace for face de-identification.

Hui Wei1,2 · Hao Yu1,2 · Hui Kuurila-Zhang2 · Guoying Zhao1,2

1ELLIS Institute Finland, Finland·2Center for Machine Vision and Signal Analysis (CMVS), University of Oulu, Finland

NeurIPS 2026 · E&D Track

Paper arXiv Code & Toolkit UtilFace Release Leaderboard BibTeX
One face processed by six de-identification methods, with HiFD utility scores at three hierarchy levels and the composite score.
Figure 1. HiFD reveals what the eye cannot: PGD and Adv-Makeup look identical to the original yet score 0.32 and 0.06, Chameleon looks noisy yet keeps the pulse signal, and WeakenDiff passes the macro and micro levels while destroying the physiological one.

Abstract

Face de-identification (FDeID) has emerged as a critical privacy-preserving technology, yet its evaluation remains fundamentally fragmented. Existing protocols rely on inconsistent metrics, heterogeneous datasets, and partial annotation coverage, so methods targeting different utility dimensions are reported on different benchmarks under different metrics, rendering cross-method comparison infeasible.

We revisit FDeID evaluation from both the data and metric perspectives. We introduce UtilFace, a curated, demographically balanced benchmark with high identity diversity, and propose HiFD, a hierarchical face de-identification metric that unifies identity suppression, multi-level utility preservation, and image quality under a single consistency-based paradigm. Using this unified protocol, we conduct a comprehensive comparative study spanning adversarial, GAN-based, and diffusion-based methods, surfacing trade-offs and failure modes that remain invisible under existing protocols.

How private is private? Under HiFD, the best of twelve methods reaches just 0.607 (95% CI 0.593–0.618); nine of twelve score below 0.30 on identity suppression, and only three de-identify more than 85% of faces from a modern recognizer ensemble at a false-accept rate of 10−2.

Leaderboard

Twelve baselines from the paper, ranked under the selected profile. Click any row to reveal the per-axis breakdown (axes defined below in HiFD Metric). The ✓ badge marks entries reproduced by maintainers.

Live data — entries Indexed —
Profile
Paradigm
# Method Paradigm P̄ Q U₁ U₂ U₃ HiFD ✓
Loading… or see the static snapshot if JS is disabled.

Submit to the HiFD leaderboard

Fork the repo, drop one JSON file under docs/page/submissions/, and open a PR. CI validates the schema and harmonic-mean self-consistency; a maintainer reviews and merges. The leaderboard rebuilds automatically.


      

Full instructions: CONTRIBUTING.md

UtilFace Benchmark

UtilFace is built from four large-scale face datasets through a four-stage curation pipeline: identity-aware cleaning, blind face restoration, NIQE-based perceptual filtering, and demographically stratified sampling at the identity level. The result is 99,928 images across 2,069 identities, with near-uniform gender balance (51.5 / 48.5%) and an ethnicity distribution flattened from a 53× to a 2.3× largest-to-smallest ratio. Because HiFD is consistency-based, UtilFace itself requires no attribute annotations.

Stage Operation # Images # Identities
01 · Aggregation Merge 4 source datasets ~69.55M~3.12M
02 · Cleaning & Enhance Identity-aware filtering; GFPGAN / CodeFormer ~20.58M~701,810
03 · Quality filter NIQE < 5.0 ~8.69M ~640,035
04 · Balanced sampling Hierarchical stratified sampling 99,9282,069

HiFD Metric

HiFD organizes facial signals into a three-level utility hierarchy, macro (L₁), micro (L₂), imperceptible (L₃), and integrates them with identity suppression (P̄) and image quality (Q) via a weighted harmonic mean:

$$ \text{HiFD} \;=\; \Bigl( w_P + w_Q + \sum_{\ell=1}^{L} w_\ell \Bigr) \, \bigg/ \, \Bigl( \frac{w_P}{\bar{P} + \epsilon} + \frac{w_Q}{Q + \epsilon} + \sum_{\ell=1}^{L} \frac{w_\ell}{U_\ell + \epsilon} \Bigr) $$

Every component is a consistency score, the similarity between a pretrained estimator's outputs on the original face and its de-identified counterpart, so HiFD requires no per-image attribute annotations. The harmonic mean's vetoing property drives the score toward zero when any single dimension collapses, surfacing imbalanced methods that pass single-axis protocols.

Is HiFD valid?

A consistency-based metric stands or falls on whether estimator consistency tracks true utility. The following results answer it with dedicated experiments:

On the expression-labelled RAF-DB test split de-identified by all twelve methods, the per-image consistency score predicts whether the true label survives with AUC 0.98 over the 28,614 pairs whose original is classified correctly; across methods the consistency sub-score ranks methods as the labels do (Spearman 0.944). On MMPD, ranking by true heart-rate error from the contact sensor agrees with U₃ at Spearman 0.944.

Grouped bars per method: expression consistency next to ground-truth accuracy retained, rPPG consistency next to contact-sensor heart-rate accuracy retained, and privacy score next to de-identification success at a fixed false-accept rate.
Figure 2. HiFD consistency scores (solid) next to ground-truth scores (hatched) for every method, coloured by paradigm. (a) Expression consistency on the de-identified RAF-DB test split vs. the fraction of label accuracy retained: within 0.06 for every method. (b) U₃ on MMPD vs. the fraction of contact-sensor heart-rate accuracy retained: same three best and same three worst methods. (c) P̄ vs. de-identification success at FAR 10−3: only the three methods with P̄ ≥ 0.45 de-identify most faces.

Citation

@inproceedings{wei2026hifd,
  title={How Private is Private? A Comparative Study for Face De-Identification},
  author={Wei, Hui and Yu, Hao and Kuurila-Zhang, Hui and Zhao, Guoying},
  booktitle={Advances in Neural Information Processing Systems},
  year={2026}
}