1CVSSP, University of Surrey, United Kingdom
Sign language production requires more than hand motion generation. Non-manual features, including mouthings, eyebrow raises, gaze, and head movements, are grammatically obligatory and cannot be recovered from manual articulators alone. Existing 3D production systems face two barriers to integrating them: the standard body model provides a facial space too low-dimensional to encode these articulations, and when richer representations are adopted, standard discrete tokenization suffers from codebook collapse, leaving most of the expression space unreachable. We propose SMPL-FX, which couples FLAME's rich expression space with the SMPL-X body, and tokenize the resulting representation with modality-specific Finite Scalar Quantization (FSQ) VAEs for body, hands, and face. M3T is an autoregressive transformer trained on this multi-modal motion vocabulary, with an auxiliary translation objective that encourages semantically grounded embeddings. Across three standard benchmarks (How2Sign, CSL-Daily, Phoenix14T) M3T achieves state-of-the-art sign language production quality, and on NMFs-CSL, where signs are distinguishable only by non-manual features, reaches 58.3% accuracy against 49.0% for the strongest comparable pose baseline.
| Method | How2Sign | CSL-Daily | Phoenix14T | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DTW-PA-JPE ↓ | DTW-JPE ↓ | DTW-PA-JPE ↓ | DTW-JPE ↓ | DTW-PA-JPE ↓ | DTW-JPE ↓ | |||||||
| Body | Hand | Body | Hand | Body | Hand | Body | Hand | Body | Hand | Body | Hand | |
| Prior Work | ||||||||||||
| NAR | 13.94 | 11.80 | – | – | – | – | – | – | – | – | – | – |
| Prog. Trans. | 14.15 | 11.57 | 14.74 | 30.17 | 15.98 | 12.91 | 16.30 | 32.63 | 13.67 | 11.95 | 15.01 | 31.77 |
| Text2Mesh | 13.99 | 13.47 | 15.50 | 32.97 | 13.47 | 12.10 | 13.76 | 30.37 | 13.48 | 12.06 | 14.04 | 31.64 |
| Adv. Train. | 13.78 | 11.17 | – | – | – | – | – | – | – | – | – | – |
| T2S-GPT | 11.48 | 6.39 | 12.65 | 18.44 | 11.94 | 5.93 | 12.32 | 15.43 | 10.38 | 6.47 | 11.65 | 19.09 |
| NSA | 7.83 | 7.33 | – | – | – | – | – | – | – | – | – | – |
| S-MotionGPT | 11.23 | 4.39 | 12.41 | 13.74 | 10.81 | 3.78 | 11.58 | 11.31 | 9.45 | 3.41 | 10.42 | 9.08 |
| SOKE | 6.82 | 2.35 | 7.75 | 10.08 | 6.24 | 1.71 | 7.38 | 9.68 | 4.77 | 1.38 | 6.04 | 7.72 |
| M3T (Ours) | ||||||||||||
| M3T w/o SLT | 4.52 | 2.42 | 5.02 | 8.94 | 4.13 | 1.59 | 4.77 | 9.94 | 3.54 | 1.54 | 4.17 | 9.14 |
| M3T w/ SLT | 4.28 | 2.28 | 4.86 | 9.44 | 3.97 | 1.51 | 4.63 | 9.28 | 3.41 | 1.46 | 4.02 | 8.55 |
Comparison with state-of-the-art SLP (text-to-sign) methods. DTW-PA-JPE measures pose-aligned joint error after rigid normalisation, while DTW-JPE measures raw joint position error; both are reported separately for body and hand joints. Bold marks the best value in each column. M3T reduces body DTW-PA-JPE by 37–29% relative to the strongest prior method (SOKE) across all three benchmarks, with the auxiliary SLT objective improving 9 of 12 metrics over the ablated variant.
| Method | Overall | Confusing | Normal | |||
|---|---|---|---|---|---|---|
| T-1 ↑ | T-5 ↑ | T-1 ↑ | T-5 ↑ | T-1 ↑ | T-5 ↑ | |
| RGB-Based Methods | ||||||
| 3D-R50 | 62.1 | 82.9 | 43.1 | 72.4 | 87.4 | 97.0 |
| DNF | 55.8 | 82.4 | 51.9 | 71.4 | 86.3 | 97.0 |
| I3D | 64.4 | 88.0 | 47.3 | 81.8 | 87.1 | 97.3 |
| TSM | 64.5 | 88.7 | 42.9 | 81.0 | 93.3 | 99.0 |
| SlowFast | 66.3 | 86.6 | 47.0 | 77.4 | 92.0 | 98.9 |
| GLE-Net | 69.0 | 88.1 | 50.6 | 79.6 | 93.6 | 99.3 |
| HMA | 64.7 | 91.0 | 42.3 | 84.8 | 94.6 | 99.3 |
| NLASLR† | 83.1 | 98.3 | – | – | – | – |
| Pose-Based Methods | ||||||
| ST-GCN | 59.9 | 86.8 | 42.2 | 79.4 | 83.4 | 96.7 |
| SignBERT | 67.0 | 95.3 | 46.4 | 92.1 | 94.5 | 99.6 |
| BEST | 68.5 | 94.4 | 49.0 | 90.3 | 94.6 | 99.7 |
| ScalingUp† | 80.2 | 97.5 | 67.2 | 95.7 | 97.5 | 99.8 |
| M3T | 74.1 | 95.1 | 58.3 | 93.0 | 93.5 | 97.6 |
SLR results on NMFs-CSL, separated by RGB and pose modalities. † denotes pre-training on large sign datasets. On the Confusing split — signs distinguishable only by non-manual features — M3T reaches 58.3% Top-1 against 49.0% for BEST, the strongest directly comparable pose baseline, a +9.3 pp gain.
This work was supported by EPSRC grant APP24554 (SignGPT — EP/Z535370/1), EPSRC grant APP78083 (UMCS — UKRI3927), and through funding from Google.org via the AI for Global Goals scheme. The authors acknowledge the use of Isambard-AI National AI Research Resource (AIRR), funded by UK DSIT via UKRI and STFC [ST/AIRR/I-A-I/1023]. This work reflects only the authors' views; the funders are not responsible for any use that may be made of the information it contains.
@article{symeonidisherzig2026m3t,
title = {M3T: Discrete Multi-Modal Motion Tokens for
Sign Language Production},
author = {Symeonidis-Herzig, Alexandre and Low, Jianhe and
Sincan, Ozge Mercanoglu and Bowden, Richard},
journal = {arXiv preprint arXiv:2603.23617},
year = {2026},
eprint = {2603.23617},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
doi = {10.48550/arXiv.2603.23617},
}