Datasets

DGS-Fabeln-1 is a parallel corpus of German text and videos containing German fairy tales interpreted into the German Sign Language (DGS) by a native DGS signer. The corpus contains 573 segments of videos with a total duration of 1 hour and 32 minutes, corresponding with 1428 written sentences. It is the first corpus of semi-naturally expressed DGS that has been filmed from 7 angles, and one of the few sign language (SL) corpora globally which have been filmed from more than 3 angles and where the listener has been simultaneously filmed.

Get it from Zenodo: https://doi.org/10.5281/zenodo.10822096

Publication (LREC 2024): https://aclanthology.org/2024.lrec-main.434

DGS-Fabeln-1-SE if an extension of DGS-Fabeln-1 with Sentiment Estimation. We used fours LLMs plus a majority voting filter to associate a sentiment label (negative, neutral, positive) to each segment of the DGS-Fabeln-1 parallel corpus, and we relied on MediaPipe to extract body/face motion features from videos. All together, this makes DGS-Fabeln-1-SE a unique 517 segments parallel corpus among text, sentiment, motion features, and video. Machine learning experiments achieve a 63.1% accuracy in predicting sentiment from motion features using the explainable XGBoost algorithm; highlighting the equivalent importance of body as well as facial features in estimating emotions in Sign Language.

Get it from Zenodo: https://doi.org/10.5281/zenodo.18879036

Publication (LREC 2026): https://lrec.elra.info/lrec2026-main-748