dorsal/arxiv
View SchemaEditEmoTalk: Controllable Speech-Driven 3D Facial Animation with Continuous Expression Editing
| Authors | Diqiong Jiang, Kai Zhu, Dan Song, Jian Chang, Chenglizhao Chen, Zhenyu Wu |
|---|---|
| Categories | |
| ArXiv ID | 2601.10000vv1 |
| URL | https://arxiv.org/abs/2601.10000 |
| License | http://creativecommons.org/licenses/by-nc-nd/4.0/ |
Abstract
Speech-driven 3D facial animation aims to generate realistic and expressive facial motions directly from audio. While recent methods achieve high-quality lip synchronization, they often rely on discrete emotion categories, limiting continuous and fine-grained emotional control. We present EditEmoTalk, a controllable speech-driven 3D facial animation framework with continuous emotion editing. The key idea is a boundary-aware semantic embedding that learns the normal directions of inter-emotion decision boundaries, enabling a continuous expression manifold for smooth emotion manipulation. Moreover, we introduce an emotional consistency loss that enforces semantic alignment between the generated motion dynamics and the target emotion embedding through a mapping network, ensuring faithful emotional expression. Extensive experiments demonstrate that EditEmoTalk achieves superior controllability, expressiveness, and generalization while maintaining accurate lip synchronization. Code and pretrained models will be released.
{
"annotation_id": "a4972d7b-fad2-4a3f-bc5d-0997a48aadfb",
"date_created": "2026-02-17T05:53:23.538000Z",
"date_modified": "2026-02-17T05:53:23.538000Z",
"file_hash": "97e22f765af63bffc228946523bfa97fd68b6b2a1bcd022065ade03ae481794b",
"private": false,
"record": {
"abstract": "Speech-driven 3D facial animation aims to generate realistic and expressive facial motions directly from audio. While recent methods achieve high-quality lip synchronization, they often rely on discrete emotion categories, limiting continuous and fine-grained emotional control. We present EditEmoTalk, a controllable speech-driven 3D facial animation framework with continuous emotion editing. The key idea is a boundary-aware semantic embedding that learns the normal directions of inter-emotion decision boundaries, enabling a continuous expression manifold for smooth emotion manipulation. Moreover, we introduce an emotional consistency loss that enforces semantic alignment between the generated motion dynamics and the target emotion embedding through a mapping network, ensuring faithful emotional expression. Extensive experiments demonstrate that EditEmoTalk achieves superior controllability, expressiveness, and generalization while maintaining accurate lip synchronization. Code and pretrained models will be released.",
"arxiv_id": "2601.10000",
"authors": [
"Diqiong Jiang",
"Kai Zhu",
"Dan Song",
"Jian Chang",
"Chenglizhao Chen",
"Zhenyu Wu"
],
"categories": [
"cs.MM",
"cs.CV"
],
"license": "http://creativecommons.org/licenses/by-nc-nd/4.0/",
"title": "EditEmoTalk: Controllable Speech-Driven 3D Facial Animation with Continuous Expression Editing",
"url": "https://arxiv.org/abs/2601.10000",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "51f915ca-f0e1-4bda-b9ca-cdebea6836ad",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}