1CVSSP, University of Surrey, United Kingdom
Realistic, high-fidelity 3D facial animations are crucial for expressive avatar systems in human-computer interaction and accessibility. Existing speech-driven methods optimise in the mesh domain and lack a perceptual signal tied to how visible the speech is, leading to blurry or under-articulated lip movements. We present VisualSpeaker, a novel training framework that introduces a perceptual lip-reading loss derived by passing photorealistic 3D Gaussian Splatting renders of the synthesised avatar through a pre-trained Visual Automatic Speech Recognition (AutoAVSR) model. This bridges the mesh-domain optimisation gap with a direct visual intelligibility signal, without requiring any additional inference-time cost. VisualSpeaker is evaluated on the MEAD dataset and achieves a 56.1% improvement in Lip Vertex Error over prior methods, while enhancing perceptual quality and maintaining mesh-driven animation controllability. The approach also yields more accurate mouthings for sign language avatar systems, where non-manual lip articulations carry grammatical meaning.
Fig. 2 — Visual comparisons for four unseen subjects and sentences from MEAD, highlighting how VisualSpeaker better preserves lip articulation than the baseline. Each subfigure shows three frames, left to right: Pseudo-GT render, VisualSpeaker without the lip-reading loss (ℒread), and VisualSpeaker with full supervision. Alongside these, a zoomed-in crop of the mouth region in the same order, top to bottom, highlights the differences in lip articulation.
| Stage / Method | LVE VOCASET ↓ | LVE MEAD ↓ |
|---|---|---|
| Pretraining | 3.06 | 7.05 |
| VisualSpeaker w/o ℒread | – | 3.85 |
| VisualSpeaker | – | 1.69 |
Lip Vertex Error (LVE, mm — lower is better) across pipeline stages. Removing the perceptual lip-reading loss ℒread raises MEAD LVE from 1.69 to 3.85 mm, the 56.1% gap the method is built to close.
| Stage / Method | PSNR ↑ | SSIM ↑ | LPIPS ↓ |
|---|---|---|---|
| Pseudo-GT Vertices | 20.47 | 0.9126 | 0.1265 |
| Pretraining | 19.48 | 0.9057 | 0.1353 |
| VisualSpeaker w/o ℒread | 19.29 | 0.9077 | 0.1326 |
| VisualSpeaker | 19.32 | 0.9083 | 0.1316 |
PSNR, SSIM and LPIPS for different training stages on the MEAD test set. Pseudo-GT vertices (italic) are an upper bound rather than a competing method. VisualSpeaker improves every perceptual metric over the ablated variant while staying close to the pseudo-GT ceiling.
| Comparison | Realism (%) ↑ | Lip Clarity (%) ↑ |
|---|---|---|
| Ours vs. Baseline | 63.8 ± 8.8 | 66.6 ± 10.5 |
| Ours vs. Pseudo-GT | 34.9 ± 8.0 | 33.6 ± 8.7 |
Percentage of A/B comparisons in which VisualSpeaker was preferred, ± standard deviation. Participants preferred VisualSpeaker over the baseline for both realism and lip clarity; against pseudo-GT vertices it remains behind, as expected of an upper bound.
This work was supported by the SNSF project “SMILE II” (CRSII5 193686), the Innosuisse IICT Flagship (PFFS-21-47), EPSRC grant APP24554 (SignGPT — EP/Z535370/1), and through funding from Google.org via the AI for Global Goals scheme. This work reflects only the authors' views and the funders are not responsible for any use that may be made of the information it contains.
@inproceedings{symeonidisherzig2025visualspeaker,
title = {VisualSpeaker: Visually-Guided 3D Avatar Lip Synthesis},
author = {Symeonidis-Herzig, Alexandre and
Sincan, {\"O}zge Mercano{\u{g}}lu and
Bowden, Richard},
booktitle = {Proceedings of the IEEE/CVF International Conference on
Computer Vision (ICCV) Workshops},
pages = {6725--6734},
year = {2025},
publisher = {IEEE},
doi = {10.1109/ICCVW69036.2025.00694},
}