dorsal/arxiv
View SchemaVideo Joint-Embedding Predictive Architectures for Facial Expression Recognition
| Authors | Lennart Eing, Cristina Luna-Jiménez, Silvan Mertes, Elisabeth André |
|---|---|
| Categories | |
| ArXiv ID | 2601.09524vv1 |
| URL | https://arxiv.org/abs/2601.09524 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
This paper introduces a novel application of Video Joint-Embedding Predictive Architectures (V-JEPAs) for Facial Expression Recognition (FER). Departing from conventional pre-training methods for video understanding that rely on pixel-level reconstructions, V-JEPAs learn by predicting embeddings of masked regions from the embeddings of unmasked regions. This enables the trained encoder to not capture irrelevant information about a given video like the color of a region of pixels in the background. Using a pre-trained V-JEPA video encoder, we train shallow classifiers using the RAVDESS and CREMA-D datasets, achieving state-of-the-art performance on RAVDESS and outperforming all other vision-based methods on CREMA-D (+1.48 WAR). Furthermore, cross-dataset evaluations reveal strong generalization capabilities, demonstrating the potential of purely embedding-based pre-training approaches to advance FER. We release our code at https://github.com/lennarteingunia/vjepa-for-fer.
{
"annotation_id": "5d0aa6d7-4ecf-4a63-835a-74c928bb4060",
"date_created": "2026-02-17T05:53:19.948000Z",
"date_modified": "2026-02-17T05:53:19.948000Z",
"file_hash": "f6b7a2054fef28f4f5b3ac5051fed503294780d60f00aa46e4b61938899ee09c",
"private": false,
"record": {
"abstract": "This paper introduces a novel application of Video Joint-Embedding Predictive Architectures (V-JEPAs) for Facial Expression Recognition (FER). Departing from conventional pre-training methods for video understanding that rely on pixel-level reconstructions, V-JEPAs learn by predicting embeddings of masked regions from the embeddings of unmasked regions. This enables the trained encoder to not capture irrelevant information about a given video like the color of a region of pixels in the background. Using a pre-trained V-JEPA video encoder, we train shallow classifiers using the RAVDESS and CREMA-D datasets, achieving state-of-the-art performance on RAVDESS and outperforming all other vision-based methods on CREMA-D (+1.48 WAR). Furthermore, cross-dataset evaluations reveal strong generalization capabilities, demonstrating the potential of purely embedding-based pre-training approaches to advance FER. We release our code at https://github.com/lennarteingunia/vjepa-for-fer.",
"arxiv_id": "2601.09524",
"authors": [
"Lennart Eing",
"Cristina Luna-Jim\u00e9nez",
"Silvan Mertes",
"Elisabeth Andr\u00e9"
],
"categories": [
"cs.CV",
"cs.HC"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Video Joint-Embedding Predictive Architectures for Facial Expression Recognition",
"url": "https://arxiv.org/abs/2601.09524",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "e339c35c-3cac-4c05-953c-996719226c2c",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}