1CVSSP, University of Surrey, United Kingdom
Sign language translation (SLT) remains challenging due to its high spatio-temporal complexity, long sequences, and the need to model multiple articulators without relying on gloss annotations. Existing approaches are typically tailored to individual datasets or languages and struggle to scale, while overlooking the relationships between sign languages that could inform more effective cross-lingual transfer.
We present SIGNET, a framework that enables motion-level knowledge transfer for cross-language sign language translation. Our key insight is that, although sign languages differ in grammar and lexicon, pretrained models capture motion-level visual patterns that can be reused across datasets and languages. SIGNET integrates multiple pretrained sign language backbones through an attention-based, hand-prior aggregation mechanism that guides a gated fusion network in dynamically selecting the most relevant experts. Comprehensive experiments on four benchmarks (How2Sign, Phoenix14T, CSL-Daily, and MeineDGS) demonstrate state-of-the-art translation performance, and SIGNET also surpasses prior methods on WLASL for sign language recognition.
Phoenix14T & CSL-Daily — SLT (gloss-free)
| Method | Phoenix14T | CSL-Daily | ||||
|---|---|---|---|---|---|---|
| B-1 ↑ | B-4 ↑ | R ↑ | B-1 ↑ | B-4 ↑ | R ↑ | |
| GFSLT-VLP | 43.71 | 21.44 | 42.49 | 39.37 | 11.00 | 36.44 |
| FLa-LLM | 46.29 | 23.09 | 45.27 | 37.13 | 14.20 | 37.25 |
| Sign2GPT | 49.54 | 22.52 | 48.90 | 41.75 | 15.40 | 42.36 |
| MMSL | 48.92 | 25.73 | 47.97 | 39.55 | 15.75 | 39.91 |
| BeyondGloss | 52.38 | 25.49 | 52.89 | 53.12 | 21.53 | 53.46 |
| PGG-SLT | 54.02 | 27.32 | 52.56 | — | — | — |
| Uni-Sign | — | — | — | 55.89 | 27.42 | 57.95 |
| SIGNET (Ours) | 54.10 | 27.82 | 53.05 | 56.77 | 28.51 | 58.66 |
How2Sign — SLT (gloss-free)
| Method | B-1 ↑ | B-4 ↑ | R ↑ | BLEURT ↑ |
|---|---|---|---|---|
| Uni-Sign | 40.8 | 15.1 | 35.4 | — |
| Geo-Sign | 40.8 | 15.1 | 35.4 | — |
| SSVP-SLT-LSP † | 43.2 | 15.5 | 38.4 | 49.6 |
| SIGNET (Ours) | 41.1 | 15.4 | 36.4 | 49.5 |
† Pre-trained on YT-ASL and How2Sign; requires 64 A100 GPUs for two weeks. SIGNET achieves comparable performance at a fraction of the compute.
MeineDGS — SLT (zero-shot cross-lingual)
| Method | B-1 ↑ | B-4 ↑ | R ↑ | BLEURT ↑ |
|---|---|---|---|---|
| Spotter+GPT | 14.82 | 0.64 | — | 21.62 |
| Spotter+Transformer | 19.50 | 1.68 | — | 19.61 |
| Geo-Sign | 16.65 | 1.55 | — | 29.72 |
| SIGNET (Ours) | 24.20 | 2.75 | 13.6 | 33.58 |
DGS was not seen during pretraining, demonstrating cross-lingual transfer to an unseen sign language.
WLASL — Sign Language Recognition
| Method | WLASL100 | WLASL2000 | ||
|---|---|---|---|---|
| P-I ↑ | P-C ↑ | P-I ↑ | P-C ↑ | |
| NLA-SLR | 91.47 | 92.17 | 61.05 | 58.05 |
| SignRep | — | — | — | 58.89 |
| Uni-Sign | 92.25 | 92.67 | 63.52 | 61.32 |
| Geo-Sign | — | — | 63.64 | 61.80 |
| SIGNET (Ours) | 93.26 | 93.47 | 64.87 | 62.07 |
Per-instance (P-I) and per-class (P-C) Top-1 accuracy on WLASL100 and WLASL2000.
@article{asasi2026signet,
title = {SIGNET: Motion-Level Knowledge Transfer for
Cross-Language Sign Language Translation},
author = {Asasi, Sobhan and Mercanoglu Sincan, Ozge and Bowden, Richard},
journal = {arXiv preprint},
year = {2026},
}