dorsal/arxiv
View SchemaRSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn Conversation
| Authors | Peng Chen, Xiaobao Wei, Yi Yang, Naiming Yao, Hui Chen, Feng Tian |
|---|---|
| Categories | |
| ArXiv ID | 2601.10606vv1 |
| URL | https://arxiv.org/abs/2601.10606 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Talking head generation is increasingly important in virtual reality (VR), especially for social scenarios involving multi-turn conversation. Existing approaches face notable limitations: mesh-based 3D methods can model dual-person dialogue but lack realistic textures, while large-model-based 2D methods produce natural appearances but incur prohibitive computational costs. Recently, 3D Gaussian Splatting (3DGS) based methods achieve efficient and realistic rendering but remain speaker-only and ignore social relationships. We introduce RSATalker, the first framework that leverages 3DGS for realistic and socially-aware talking head generation with support for multi-turn conversation. Our method first drives mesh-based 3D facial motion from speech, then binds 3D Gaussians to mesh facets to render high-fidelity 2D avatar videos. To capture interpersonal dynamics, we propose a socially-aware module that encodes social relationships, including blood and non-blood as well as equal and unequal, into high-level embeddings through a learnable query mechanism. We design a three-stage training paradigm and construct the RSATalker dataset with speech-mesh-image triplets annotated with social relationships. Extensive experiments demonstrate that RSATalker achieves state-of-the-art performance in both realism and social awareness. The code and dataset will be released.
{
"annotation_id": "976aefd6-7641-4934-b2b8-ade28029f0cf",
"date_created": "2026-02-17T05:53:24.177000Z",
"date_modified": "2026-02-17T05:53:24.177000Z",
"file_hash": "7125b4a1f8e4bc9e1711616a5bd5f3c7f279cee322c0e2ba89cf3416cb32b96d",
"private": false,
"record": {
"abstract": "Talking head generation is increasingly important in virtual reality (VR), especially for social scenarios involving multi-turn conversation. Existing approaches face notable limitations: mesh-based 3D methods can model dual-person dialogue but lack realistic textures, while large-model-based 2D methods produce natural appearances but incur prohibitive computational costs. Recently, 3D Gaussian Splatting (3DGS) based methods achieve efficient and realistic rendering but remain speaker-only and ignore social relationships. We introduce RSATalker, the first framework that leverages 3DGS for realistic and socially-aware talking head generation with support for multi-turn conversation. Our method first drives mesh-based 3D facial motion from speech, then binds 3D Gaussians to mesh facets to render high-fidelity 2D avatar videos. To capture interpersonal dynamics, we propose a socially-aware module that encodes social relationships, including blood and non-blood as well as equal and unequal, into high-level embeddings through a learnable query mechanism. We design a three-stage training paradigm and construct the RSATalker dataset with speech-mesh-image triplets annotated with social relationships. Extensive experiments demonstrate that RSATalker achieves state-of-the-art performance in both realism and social awareness. The code and dataset will be released.",
"arxiv_id": "2601.10606",
"authors": [
"Peng Chen",
"Xiaobao Wei",
"Yi Yang",
"Naiming Yao",
"Hui Chen",
"Feng Tian"
],
"categories": [
"cs.CV"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn Conversation",
"url": "https://arxiv.org/abs/2601.10606",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "d87d8a0f-02ef-408c-b9fb-00800573e696",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}