dorsal/arxiv
View SchemaViTNT-FIQA: Training-Free Face Image Quality Assessment with Vision Transformers
| Authors | Guray Ozgur, Eduarda Caldeira, Tahar Chettaoui, Jan Niklas Kolf, Marco Huber, Naser Damer, Fadi Boutros |
|---|---|
| Categories | |
| ArXiv ID | 2601.05741vv1 |
| URL | https://arxiv.org/abs/2601.05741 |
| License | http://creativecommons.org/licenses/by-nc-sa/4.0/ |
Abstract
Face Image Quality Assessment (FIQA) is essential for reliable face recognition systems. Current approaches primarily exploit only final-layer representations, while training-free methods require multiple forward passes or backpropagation. We propose ViTNT-FIQA, a training-free approach that measures the stability of patch embedding evolution across intermediate Vision Transformer (ViT) blocks. We demonstrate that high-quality face images exhibit stable feature refinement trajectories across blocks, while degraded images show erratic transformations. Our method computes Euclidean distances between L2-normalized patch embeddings from consecutive transformer blocks and aggregates them into image-level quality scores. We empirically validate this correlation on a quality-labeled synthetic dataset with controlled degradation levels. Unlike existing training-free approaches, ViTNT-FIQA requires only a single forward pass without backpropagation or architectural modifications. Through extensive evaluation on eight benchmarks (LFW, AgeDB-30, CFP-FP, CALFW, Adience, CPLFW, XQLFW, IJB-C), we show that ViTNT-FIQA achieves competitive performance with state-of-the-art methods while maintaining computational efficiency and immediate applicability to any pre-trained ViT-based face recognition model.
{
"annotation_id": "72fc21b0-4321-4b60-b876-b4ba0333c8b2",
"date_created": "2026-02-17T05:53:04.827000Z",
"date_modified": "2026-02-17T05:53:04.827000Z",
"file_hash": "89a37212e561346d1fe29c04a9ec2597e84f76574e1c395142ebf651bc54a23c",
"private": false,
"record": {
"abstract": "Face Image Quality Assessment (FIQA) is essential for reliable face recognition systems. Current approaches primarily exploit only final-layer representations, while training-free methods require multiple forward passes or backpropagation. We propose ViTNT-FIQA, a training-free approach that measures the stability of patch embedding evolution across intermediate Vision Transformer (ViT) blocks. We demonstrate that high-quality face images exhibit stable feature refinement trajectories across blocks, while degraded images show erratic transformations. Our method computes Euclidean distances between L2-normalized patch embeddings from consecutive transformer blocks and aggregates them into image-level quality scores. We empirically validate this correlation on a quality-labeled synthetic dataset with controlled degradation levels. Unlike existing training-free approaches, ViTNT-FIQA requires only a single forward pass without backpropagation or architectural modifications. Through extensive evaluation on eight benchmarks (LFW, AgeDB-30, CFP-FP, CALFW, Adience, CPLFW, XQLFW, IJB-C), we show that ViTNT-FIQA achieves competitive performance with state-of-the-art methods while maintaining computational efficiency and immediate applicability to any pre-trained ViT-based face recognition model.",
"arxiv_id": "2601.05741",
"authors": [
"Guray Ozgur",
"Eduarda Caldeira",
"Tahar Chettaoui",
"Jan Niklas Kolf",
"Marco Huber",
"Naser Damer",
"Fadi Boutros"
],
"categories": [
"cs.CV",
"cs.LG"
],
"license": "http://creativecommons.org/licenses/by-nc-sa/4.0/",
"title": "ViTNT-FIQA: Training-Free Face Image Quality Assessment with Vision Transformers",
"url": "https://arxiv.org/abs/2601.05741",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "92e14c5a-4243-454d-9cfd-67f5736cf9f6",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}