- Research Article
- 10.1109/lsp.2026.3668742
CSTNet: A Cross-Scale Transformer Network for Singing Melody Extraction
- Jan 01, 2026
- IEEE Signal Processing Letters
- Jianbin Chen + 3 more +3
Singing melody extraction from polyphonic music is a complex but important task in music information retrieval. Most existing models struggle with adaptive multi-scale feature extraction and long-range dependency modeling, often resulting in common artifacts such as octave errors and pitch fragmentation. To address this, we propose the Cross-Scale Transformer Network (CSTNet), a novel deep learning architecture that dynamically captures and fuses multi-scale temporal and spectral features. The CSTNet introduces two key components: the Multi scale Time-Frequency Aggregation (MF-TFA) module for adaptive local feature representation across resolutions, and Feature Cross Transformer Block (FCTB), which enables explicit cross scale interaction through a global channel-wise cross-attention mechanism. Experimental results show that our proposed method outperforms five compared state-of-the-art methods, achieving overall accuracy scores of 87.3%, 90.5%, and 77.4% on the ADC 2004, MIREX 05, and MEDLEY DB datasets, respectively. Visualization results demonstrate that CSTNet effectively reduces octave errors.
Read more