Paper Reading AI Learner

A Benchmark for Evaluating Machine Translation Metrics on Dialects Without Standard Orthography

2023-11-28 15:12:11
Noëmi Aepli, Chantal Amrhein, Florian Schottmann, Rico Sennrich

Abstract

For sensible progress in natural language processing, it is important that we are aware of the limitations of the evaluation metrics we use. In this work, we evaluate how robust metrics are to non-standardized dialects, i.e. spelling differences in language varieties that do not have a standard orthography. To investigate this, we collect a dataset of human translations and human judgments for automatic machine translations from English to two Swiss German dialects. We further create a challenge set for dialect variation and benchmark existing metrics' performances. Our results show that existing metrics cannot reliably evaluate Swiss German text generation outputs, especially on segment level. We propose initial design adaptations that increase robustness in the face of non-standardized dialects, although there remains much room for further improvement. The dataset, code, and models are available here: this https URL

Abstract (translated)

为了在自然语言处理领域取得合理的进步,了解我们使用的评估指标的局限性非常重要。在这项工作中,我们评估了衡量标准如何适应非标准德语,即没有标准的拼写差异的语言变体。为了研究这个问题,我们收集了来自英语到两种瑞士德语的自动机器翻译的人机翻译数据和人类判断。我们进一步创建了一个方言变化挑战集,对现有指标的性能进行了基准测试。我们的结果表明,现有的指标无法可靠地评估瑞士德语文本生成输出,尤其是在段级别。我们提出了最初的设计适应方案,以增加在非标准德语面对的可靠性,尽管还有很大的改进空间。数据集、代码和模型都可以在这里找到:这个链接

URL

https://arxiv.org/abs/2311.16865

PDF

https://arxiv.org/pdf/2311.16865.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot