Paper Reading AI Learner

Improving BERT for Symbolic Music Understanding Using Token Denoising and Pianoroll Prediction

2025-07-07 08:52:06
Jun-You Wang, Li Su

Abstract

We propose a pre-trained BERT-like model for symbolic music understanding that achieves competitive performance across a wide range of downstream tasks. To achieve this target, we design two novel pre-training objectives, namely token correction and pianoroll prediction. First, we sample a portion of note tokens and corrupt them with a limited amount of noise, and then train the model to denoise the corrupted tokens; second, we also train the model to predict bar-level and local pianoroll-derived representations from the corrupted note tokens. We argue that these objectives guide the model to better learn specific musical knowledge such as pitch intervals. For evaluation, we propose a benchmark that incorporates 12 downstream tasks ranging from chord estimation to symbolic genre classification. Results confirm the effectiveness of the proposed pre-training objectives on downstream tasks.

Abstract (translated)

我们提出了一种类似BERT的预训练模型,用于符号音乐的理解,并在广泛的下游任务中实现了具有竞争力的表现。为了达成这一目标,我们设计了两个新颖的预训练目标,即标记校正和钢琴卷帘预测。首先,我们在音符标记的一部分上添加少量噪声以进行数据扰动,然后训练模型去除这些被污染的标记中的噪声;其次,我们也训练模型从受干扰的音符标记中预测出小节级及局部钢琴卷帘衍生表示。我们认为,这些目标能够引导模型更好地学习特定的音乐知识,例如音高间隔等。 为了评估我们的模型,我们提出了一个包含12个下游任务的新基准测试集,这些任务涵盖了从和弦估计到符号风格分类等多个方面。实验结果证实了所提出的预训练目标在下游任务中的有效性。

URL

https://arxiv.org/abs/2507.04776

PDF

https://arxiv.org/pdf/2507.04776.pdf


Tags
3D Action Action_Localization Action_Recognition Activity Adversarial Agent Attention Autonomous Bert Boundary_Detection Caption Chat Classification CNN Compressive_Sensing Contour Contrastive_Learning Deep_Learning Denoising Detection Dialog Diffusion Drone Dynamic_Memory_Network Edge_Detection Embedding Embodied Emotion Enhancement Face Face_Detection Face_Recognition Facial_Landmark Few-Shot Gait_Recognition GAN Gaze_Estimation Gesture Gradient_Descent Handwriting Human_Parsing Image_Caption Image_Classification Image_Compression Image_Enhancement Image_Generation Image_Matting Image_Retrieval Inference Inpainting Intelligent_Chip Knowledge Knowledge_Graph Language_Model LLM Matching Medical Memory_Networks Multi_Modal Multi_Task NAS NMT Object_Detection Object_Tracking OCR Ontology Optical_Character Optical_Flow Optimization Person_Re-identification Point_Cloud Portrait_Generation Pose Pose_Estimation Prediction QA Quantitative Quantitative_Finance Quantization Re-identification Recognition Recommendation Reconstruction Regularization Reinforcement_Learning Relation Relation_Extraction Represenation Represenation_Learning Restoration Review RNN Robot Salient Scene_Classification Scene_Generation Scene_Parsing Scene_Text Segmentation Self-Supervised Semantic_Instance_Segmentation Semantic_Segmentation Semi_Global Semi_Supervised Sence_graph Sentiment Sentiment_Classification Sketch SLAM Sparse Speech Speech_Recognition Style_Transfer Summarization Super_Resolution Surveillance Survey Text_Classification Text_Generation Time_Series Tracking Transfer_Learning Transformer Unsupervised Video_Caption Video_Classification Video_Indexing Video_Prediction Video_Retrieval Visual_Relation VQA Weakly_Supervised Zero-Shot