Abstract Method Qualitative Results Quantitative Results Acknowledgements BibTeX
ICCV Workshop 2025

VisualSpeaker: Visually-Guided 3D Avatar Lip Synthesis

Alexandre Symeonidis-Herzig1,  Özge Mercanoğlu Sincan1,  Richard Bowden1

1CVSSP, University of Surrey, United Kingdom

Paper PDF arXiv
VisualSpeaker qualitative results: ground truth frames, predicted meshes, and photorealistic renders

Fig. 1 — VisualSpeaker qualitative results on the MEAD dataset. Each column corresponds to a different word being spoken. Top: ground truth video frames. Middle: predicted 3D meshes. Bottom: photorealistic 3D Gaussian Splatting renders produced by VisualSpeaker.


Abstract

Realistic, high-fidelity 3D facial animations are crucial for expressive avatar systems in human-computer interaction and accessibility. Existing speech-driven methods optimise in the mesh domain and lack a perceptual signal tied to how visible the speech is, leading to blurry or under-articulated lip movements. We present VisualSpeaker, a novel training framework that introduces a perceptual lip-reading loss derived by passing photorealistic 3D Gaussian Splatting renders of the synthesised avatar through a pre-trained Visual Automatic Speech Recognition (AutoAVSR) model. This bridges the mesh-domain optimisation gap with a direct visual intelligibility signal, without requiring any additional inference-time cost. VisualSpeaker is evaluated on the MEAD dataset and achieves a 56.1% improvement in Lip Vertex Error over prior methods, while enhancing perceptual quality and maintaining mesh-driven animation controllability. The approach also yields more accurate mouthings for sign language avatar systems, where non-manual lip articulations carry grammatical meaning.


Method

Perceptual Lip-Reading Loss via Visual ASR

Visual Intelligibility Signal Rather than optimising only in the mesh domain, VisualSpeaker renders the synthesised avatar with 3D Gaussian Splatting and passes the render through a pre-trained AutoAVSR model to obtain a perceptual lip-reading loss.
3D Gaussian Splatting Rendering Differentiable photorealistic rendering via 3DGS connects the mesh-domain synthesis network to the pixel-domain visual speech recognition model, enabling end-to-end gradient flow from visual intelligibility back to lip shape.
No Inference-Time Overhead The AutoAVSR model is used only during training to compute the perceptual loss. At inference time, VisualSpeaker runs at the same cost as the base synthesis network.

Audio & Visual Feature Extraction

Audio Encoder — wav2vec 2.0 A pre-trained wav2vec 2.0 model encodes raw speech waveforms into rich acoustic representations that capture fine-grained phoneme-level timing.
Visual Encoder — AutoAVSR AutoAVSR extracts visual speech features from video, providing the perceptual supervisory signal that penalises visually unintelligible lip configurations.
Audio–Visual Fusion Audio and visual modalities are fused before the lip synthesis network, enabling the model to leverage complementary cues from both speech acoustics and visual speech.

Lip Synthesis Network & Sign Language Application

Mesh-Driven 3D Facial Deformations The synthesis network predicts vertex-level deformations on a parametric face mesh, preserving full animation controllability and compatibility with downstream avatar pipelines.
Improved Mouthings for Sign Language Avatars More accurate lip articulations directly benefit sign language avatar systems, where mouthings — visible lip movements derived from spoken words — carry grammatical meaning that cannot be inferred from hand shape alone.

Qualitative Results

Qualitative comparison for four unseen MEAD subjects: Pseudo-GT render, VisualSpeaker without the lip-reading loss, and VisualSpeaker with full supervision, each with a zoomed crop of the mouth region

Fig. 2 — Visual comparisons for four unseen subjects and sentences from MEAD, highlighting how VisualSpeaker better preserves lip articulation than the baseline. Each subfigure shows three frames, left to right: Pseudo-GT render, VisualSpeaker without the lip-reading loss (ℒread), and VisualSpeaker with full supervision. Alongside these, a zoomed-in crop of the mouth region in the same order, top to bottom, highlights the differences in lip articulation.


Quantitative Results

56.1%
Reduction in Lip Vertex Error on MEAD — adding the perceptual lip-reading loss cuts LVE from 3.85 mm to 1.69 mm, a 56.1% reduction, driving the model to produce visually intelligible lip shapes while maintaining full mesh-driven animation controllability.

Lip Vertex Error

Stage / Method LVE VOCASET ↓ LVE MEAD ↓
Pretraining3.067.05
VisualSpeaker w/o ℒread–3.85
VisualSpeaker–1.69

Lip Vertex Error (LVE, mm — lower is better) across pipeline stages. Removing the perceptual lip-reading loss ℒread raises MEAD LVE from 1.69 to 3.85 mm, the 56.1% gap the method is built to close.

Visual Quality — MEAD Test Set

Stage / Method PSNR ↑ SSIM ↑ LPIPS ↓
Pseudo-GT Vertices20.470.91260.1265
Pretraining19.480.90570.1353
VisualSpeaker w/o ℒread19.290.90770.1326
VisualSpeaker19.320.90830.1316

PSNR, SSIM and LPIPS for different training stages on the MEAD test set. Pseudo-GT vertices (italic) are an upper bound rather than a competing method. VisualSpeaker improves every perceptual metric over the ablated variant while staying close to the pseudo-GT ceiling.

User Study — Overall Preference

Comparison Realism (%) ↑ Lip Clarity (%) ↑
Ours vs. Baseline63.8 ± 8.866.6 ± 10.5
Ours vs. Pseudo-GT34.9 ± 8.033.6 ± 8.7

Percentage of A/B comparisons in which VisualSpeaker was preferred, ± standard deviation. Participants preferred VisualSpeaker over the baseline for both realism and lip clarity; against pseudo-GT vertices it remains behind, as expected of an upper bound.


Acknowledgements

This work was supported by the SNSF project “SMILE II” (CRSII5 193686), the Innosuisse IICT Flagship (PFFS-21-47), EPSRC grant APP24554 (SignGPT — EP/Z535370/1), and through funding from Google.org via the AI for Global Goals scheme. This work reflects only the authors' views and the funders are not responsible for any use that may be made of the information it contains.


BibTeX

@inproceedings{symeonidisherzig2025visualspeaker,
  title     = {VisualSpeaker: Visually-Guided 3D Avatar Lip Synthesis},
  author    = {Symeonidis-Herzig, Alexandre and
               Sincan, {\"O}zge Mercano{\u{g}}lu and
               Bowden, Richard},
  booktitle = {Proceedings of the IEEE/CVF International Conference on
               Computer Vision (ICCV) Workshops},
  pages     = {6725--6734},
  year      = {2025},
  publisher = {IEEE},
  doi       = {10.1109/ICCVW69036.2025.00694},
}