SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks
Abstract
Sign language models are typically trained on datasets cap-tured under constrained conditions, with limited viewpoint, background,and signer-identity diversity, leading to poor robustness under real-worlddistribution shifts. We introduce SignNet-1M, a large-scale augmenteddataset spanning ASL, CSL, and German Sign Language (DGS).SignNet-1M synthesizes realistic variations along three axes: (i) novel-view rendering (rotation and zoom) via 3D Gaussian Splatting (3DGS),(ii) scene/ identity editing via diffusion models for background replace-ment and signer substitution while preserving sign motion and linguisticcontent, and (iii) post-rendering augmentations that emulate captureand compression artifacts (e.g., geometric transforms, photometric shifts,mild temporal resampling, and compression) to better match in-the-wildrecordings. Beyond data release, we provide a unified benchmark suiteacross downstream tasks (e.g., translation and recognition) and abla-tions that isolate each augmentation component. Experiments acrossbackbones show that training with SignNet-1M consistently improvesgeneralization under cross-view, cross-background, cross-identity, andpost-rendering shifts, while maintaining strong in-distribution perfor-mance. The dataset, released augmentation components, metadata, andbenchmark resources are available at https://signnet.chatsign.ai/.