Abstract Method Dimensions Qualitative Results Acknowledgements BibTeX
ECCV 2026

BackTranslation2.0: A Linguistically Motivated Metric to Assess Sign Language Production

Oliver Cory1,  Maksym Ivashechkin1,  Karahan Sahin1,  Oline Ranum1,  Jianhe Low1,
Edward Fish1,  Anton Pelykh1,  Ozge Mercanoglu Sincan1,  Richard Bowden1

1CVSSP, University of Surrey, United Kingdom

Paper arXiv Code (Coming Soon)
BackTranslation2.0 pipeline overview

Fig. 1 — Overview of the BT2 pipeline. Given a source sentence and generated sign output, multi-modal extraction produces a structured sample for segment- and sequence-level analysis. Phase 1 base tools extract lexical, spatial, phonological, motion, and visual evidence stored in a shared memory trace. Phase 2 comparison tools cross-reference this evidence against linguistic expectations to produce grounded judgements. Deterministic aggregation yields per-dimension scores and an overall understandability score, and a final LLM stage writes a grounded natural-language assessment.


Abstract

Sign Languages (SLs) are the primary means of communication for millions of deaf individuals, yet existing evaluation metrics for generated SL remain simplistic and poorly aligned with human judgements. We introduce BackTranslation2.0, a linguistically grounded evaluation metric for text-to-sign translation that moves beyond naïve backtranslation. Our approach adopts an agentic framework in which a deterministic pipeline orchestrates a suite of specialised tools to assess four scoring dimensions — grammatical correctness, phonological accuracy, motion fluency, and generation fidelity — aligned with human rater assessments. Tool outputs are not treated independently: a set of LLM-based cross-referential comparison modules evaluates consistency across tools and checks outputs against linguistic expectations, enabling structured reasoning over grammatical, phonological, and motion-level evidence. Final dimension scores are computed through deterministic weighted formulas over validated tool outputs. To validate BackTranslation2.0, we introduce and evaluate on a British Sign Language (BSL) dataset annotated by native Deaf raters across the same quality dimensions, benchmarking against six baseline metrics. Our method demonstrates strong correlation with human judgements across all dimensions, providing a more comprehensive, interpretable, and linguistically principled evaluation framework for sign language production systems.


Method

BackTranslation2.0 (BT2) is a deterministic agentic framework for linguistically grounded evaluation of text-to-sign translation. Rather than relying on a single end-to-end metric, BT2 runs a fixed two-phase pipeline of specialised tools assessing grammatical correctness, phonological accuracy, motion fluency, and generation fidelity. All intermediate outputs are retained in a shared memory trace, enabling deterministic scoring and an auditable final natural-language assessment.

Phase 1 — Base Evidence Extraction

Pseudo-Gloss LLM proposes BSL pseudo-gloss order from the source sentence.
Sign Spotting Top-K sign-embedding gloss matches per segment.
Directionality SL-GCN indexing/pointing classifier for directional verb agreement.
Topographic Scorer FastDTW wrist-path direction matching for spatial grammar.
Handshape Transformer handshape classifier for left & right hands.
Location Location classifier from SMPL-Xdepth contact-proximity features.
Non-Manuals AU-to-NMF classifier mapping per-frame action units to NMF categories.
Fluency Flow-matching bottleneck fluency score via reference Gaussian in latent space.
Visual Metrics Deterministic smoothness/spectrum score encoding high-frequency energy, jerk, and motion activity.
Hand Fidelity Binary hand-realism logit score (real vs. synthetic classifier over crops).

Phase 2 — Cross-Reference Comparison

Spot-Gloss Comparison Semantic alignment between the pseudo-gloss and sign-spotted candidates using deterministic prefiltering followed by LLM matching on unresolved pairs.
Manual Features Comparison Dictionary-based check of handshape and location against BSL lexical expectations, with optional LLM fallback for low-confidence entries.
Non-Manual Features Comparison Expected NMF markers inferred via LLM or deterministic rules, compared against detected non-manual evidence from Phase 1.
Directionality Comparison Deterministic cross-referencing of SL-GCN indexing predictions and topographic direction scores against grammatical expectations.

Output — Deterministic Scoring & Grounded Assessment

Deterministic Dimension Scores Four dimension scores d = [dgram, dphon, dflu, dfid] are computed via fixed subcomponent weights on a common 0–4 scale. The overall understandability score is their arithmetic mean: u = ¼(dgram + dphon + dflu + dfid). No LLM step modifies these numeric scores.
Final Grounded Assessment A final LLM stage receives the full memory trace together with the fixed deterministic scores, producing an auditable natural-language assessment. This stage is explanation-only and does not recompute or replace any numeric score.

Scoring Dimensions

BT2 evaluates four dimensions aligned with how native Deaf raters assess sign language quality.

Grammatical Correctness Adherence to target sign language syntax, including use of signing space, pronominal indexing, and directional verb agreement. Evaluated via pseudo-gloss order, topographic scoring, and directionality tools.
Phonological Accuracy Correctness of sub-lexical components including handshape, movement, location, orientation, and non-manual features. Assessed via handshape, location, and NMF tools with dictionary cross-referencing.
Motion Fluency Naturalness and temporal smoothness of signing motion — capturing the fluid, continuous articulation expected in native signing. Evaluated via flow-matching bottleneck score and visual motion metrics.
Generation Fidelity Visual quality and anatomical plausibility of the generated output, ensuring the signer representation is realistic and interpretable. Assessed via kinematic visual metrics and the hand-fidelity classifier.

Qualitative Analysis

Qualitative evaluation of BT2 visual-base tools

Fig. 3 — Qualitative evaluation of BT2 visual-base tools. Overlays show predicted class and score, with the top row rewarding good production and the bottom row penalising poor production. Columns (left to right): Handshape, Non-manuals, Hand fidelity, Contact location. Handshape and Non-manuals compare best production against matched corruptions; Hand fidelity contrasts a reference clip with poor-handed synthesis; Contact location shows correct (top) versus missing (bottom) contact via a frontal SMPL-X overlay, depth map, and side sparse view.


Quantitative Results

We evaluate on four complementary benchmark subsets spanning real signing, controlled corruption, synthetic production, and targeted spatial-grammar evaluation. The table below compares BT2 with reference-free and reference-based baselines using direction-normalised Pearson r / Spearman ρ against human judgements — higher values always indicate stronger alignment.

Evaluation Benchmark

Subset Content Purpose
Human SRT 6 signers × 6 sentences Proficiency variation
Known Corruptions 6 sent. × 6 sign. × 3 corrupt. Anomaly sensitivity
Synthetic Productions 6 production systems Production-method evaluation
Directionality / Topography Targeted spatial-grammar set Grammatical placement

Metric – Human Alignment (Table 3)

Method Known Corr. Synthetics (Under.) Synthetics (Qual.) Human SRT
rρ rρ rρ rρ
Reference-based (requires source sequence)
DTW-MPJPE 0.910.50 −0.14−0.03 −0.33−0.37 0.040.00
SiBLEU −0.58−0.50 0.030.31 −0.050.31 0.960.90
SignCLIP p2p
BSL sentence −0.97−1.00 0.850.90 0.820.70 0.690.80
BSL segmentation −0.55−0.50 0.750.70 0.630.60 −0.46−0.40
Multilingual sentence 0.851.00 0.700.70 0.550.60 −0.10−0.40
Multilingual segmentation −0.59−0.50 0.830.70 0.750.60 −0.300.00
SVAE 0.951.00 0.430.26 0.540.31 0.810.80
Reference-free (requires source text only)
BT1
BLEU-4 −0.190.00 −0.38−0.44 −0.15−0.10 0.000.00
BLEURT 0.090.50 −0.350.03 −0.39−0.09 0.150.35
ChrF −0.87−1.00 −0.27−0.09 −0.54−0.37 −0.15−0.35
ROUGE −0.190.00 −0.34−0.44 −0.08−0.10 0.000.00
SignCLIP p2t
BSL sentence 0.250.50 0.030.30 0.170.40 0.560.60
BSL segmentation −0.25−0.50 −0.33−0.30 −0.20−0.10 −0.78−0.70
Multilingual sentence 0.890.50 −0.64−0.60 −0.41−0.30 −0.07−0.10
Multilingual segmentation 0.770.50 −0.37−0.10 −0.50−0.30 −0.11−0.30
BT2 Overall (ours) 1.001.00 0.420.60 0.520.66 −0.43−0.30
BT2 per-component sub-scores
BT2 Phonology 0.280.50 0.480.26 0.35−0.03 0.870.90
BT2 Grammar −1.00−1.00 −0.50−0.14 −0.45−0.26 −0.45−0.60
BT2 Fluency 0.280.50 0.210.37 0.400.54 −0.46−0.40
BT2 Fidelity 0.801.00 0.740.94 0.830.94 −0.81−0.90
Human inter-rater agreement
Gwet's AC2 0.690.69 0.680.69
ICC(2,1) −0.020.51 0.350.07
Within-1 agreement 0.740.79 0.780.75
Quadratic weighted κ −0.030.46 0.350.09
Handshape flag α 0.240.26 0.260.09
Non-manual flag α 0.270.29 0.29−0.03
Word-order flag α 0.300.01 0.01−0.01

BT2 Overall achieves perfect correlation (r = 1.00, ρ = 1.00) on Known Corruptions — matching or exceeding the strongest reference-based alternatives while requiring only source text. BT2 Phonology reaches ρ = 0.90 on Human SRT where proficiency shows through articulatory form; BT2 Fidelity is strongest on Synthetic productions (ρ = 0.94). All values are direction-normalised so higher always indicates better metric–human alignment.

Corruption Sensitivity — Synthetic BSL Corpus (Table 5)

Per-dimension win rates on n = 180 synthetic BSL corpus corruptions (uncorrupted BT2 ≥2.5), conditioned on which classifier fires. Columns measure overall score plus per-dimension scores; rows are corruption type. Bordered cells mark the overall score and the matched target dimension for each corruption.
≥80%   65–79%   55–64%   wrong-way

Corruption n Overall Phon. HS Loc Gram Flu. Fid.
Handshape classifier fires
Handshape138 76% 73% 86% 75% 60% 28% 65%
Location129 75% 76% 82% 74% 61% 52% 59%
Word-order134 73% 73% 82% 73% 83% 42% 66%
Temporal121 79% 71% 79% 75% 60% 55% 66%
Fidelity137 74% 73% 85% 73% 60% 26% 64%
Location classifier fires
Handshape87 69% 68% 87% 98% 57% 22% 62%
Location93 68% 72% 83% 95% 61% 51% 62%
Word-order88 78% 73% 88% 94% 56% 47% 64%
Temporal78 75% 67% 87% 96% 51% 51% 69%
Fidelity88 65% 64% 85% 97% 59% 27% 66%

BT2 correctly localises each corruption to its target dimension across all five corruption types (overall win rate 65–79%). Fluency is deliberately wrong-way for handshape and fidelity corruptions — these alter phonological form without disrupting motion, confirming BT2 does not over-penalise unaffected dimensions.

Sentence-level Directional Predictions (Table 6)

Directional-verb evidence extracted by the directionality comparison tool across clause pairs. Bold words are directional verbs; green pronouns are extracted grammatical-person entities; arrows (1 → 2) denote predicted agreement direction between person loci (1 = 1st, 2 = 2nd, 3 = 3rd person). NDV: no directional verb.

Sentence Predicted Matches GT
I will help you learn some signs but you can ask him 1 → 2, 2 → 3
I will help you learn some signs but I can ask him 1 → 2, 1 → 3
I will help him learn some signs but you can ask him 1 → 3, NDV
She will help you learn some signs but he can ask me NDV, 1 → 3
I will help you learn some signs but you can ask him NDV, NDV

Across all clause pairs, extracted person-locus sequences match the intended directional patterns, illustrating that the directional stack captures clause-conditioned role transfer rather than only surface motion similarity.


Acknowledgements

This work was supported by EPSRC grant APP24554 (SignGPT — EP/Z535370/1), EPSRC grant APP78083 (UMCS — UKRI3927), and through funding from Google.org via the AI for Global Goals scheme. The authors acknowledge the use of the Isambard-AI National AI Research Resource (AIRR), funded by UK DSIT via UKRI and STFC [ST/AIRR/I-A-I/1023]. This work reflects only the authors’ views; the funders are not responsible for any use that may be made of the information it contains.


BibTeX

@article{cory2026backtranslation2,
  title   = {BackTranslation2.0: A Linguistically Motivated Metric
             to Assess Sign Language Production},
  author  = {Cory, Oliver and Ivashechkin, Maksym and Sahin, Karahan and
             Ranum, Oline and Low, Jianhe and Fish, Edward and
             Pelykh, Anton and Mercanoglu Sincan, Ozge and
             Bowden, Richard},
  journal = {arXiv preprint arXiv:2606.28673},
  year    = {2026},
}