Overview
We provide audio samples from five randomly selected full-length songs from the test set. For each song, two specific segments are selected according to the predictions of SongSQA: Segment A is predicted as lower quality, and Segment B is predicted as higher quality.
This page combines the audio demo with the supplementary analysis. We first present a detailed case study using the excerpts from Song 1 to illustrate how the framework captures local variations in singing quality. We then summarize the results of a subjective evaluation across all five songs to examine whether the predicted segment-level differences are consistent with human perception.
Case Study
The figure below shows the temporal score curve predicted by SongSQA for a representative song from the test set. Beyond assigning a single overall quality score, the framework reveals clear quality variations across individual segments, reflecting the dynamic nature of vocal performance throughout a complete track.
To further examine whether these local variations are perceptually meaningful, we highlight two representative regions for detailed analysis. These regions correspond directly to the two excerpts of Song 1 provided on this page: Segment A represents a lower predicted quality, and Segment B represents a higher predicted quality.
In the segment receiving a lower score, the vocal delivery appears less stable, particularly in phrases within the lower vocal register, where the descending pitches are not fully supported, resulting in a less stable tonal center. This leads to a weaker sense of intonational precision and phrase support. In addition, the phrasing is comparatively less cohesive, with several note transitions sounding less smoothly connected, which reduces the overall fluency of the performance.
By contrast, the segment receiving a higher score exhibits a more stable pitch contour and a clearer tonal center, especially during sustained notes and phrase transitions. The vocal line appears better supported, with more consistent breath control and a more continuous phrase shape, resulting in a stronger sense of projection and musical coherence.
Taken together, the comparison suggests that the predicted score variations align closely with human perception. As demonstrated in this case study, the framework successfully tracks local variations in vocal delivery that influence human evaluation.
Subjective Evaluation
To further examine whether the segment-level quality differences identified by SongSQA are perceptually meaningful, we conducted a small-scale pairwise subjective evaluation. We randomly selected five full-length songs from the test set. For each song, we selected two audio segments: Segment A, predicted as lower quality, and Segment B, predicted as higher quality. Participants were asked to conduct a blind pairwise evaluation and identify the segment exhibiting better singing performance.
A total of 12 listeners participated in the evaluation, including 8 males and 4 females. To account for varying levels of musical expertise, the participant pool included 8 amateur music enthusiasts and 4 individuals with systematic musical training or professional musical backgrounds.
Across the five songs, 52 out of 60 valid judgments were aligned with the model predictions, corresponding to an agreement rate of 86.7%. Agreement remained consistently high across most song pairs. These results suggest that the segment-level quality differences identified by SongSQA are generally consistent with human perception. Given the modest scale of this study, we regard it as supplementary qualitative evidence supporting the effectiveness of the proposed framework.
| Song ID | Votes for Seg. A | Votes for Seg. B | Agreement |
|---|---|---|---|
| Song 1 | 1 | 11 | 91.7% |
| Song 2 | 0 | 12 | 100.0% |
| Song 3 | 1 | 11 | 91.7% |
| Song 4 | 3 | 9 | 75.0% |
| Song 5 | 3 | 9 | 75.0% |
| Total | 8 | 52 | 86.7% |
Audio Demo
For each song, Segment A corresponds to the lower predicted quality excerpt, and Segment B corresponds to the higher predicted quality excerpt. Reviewers may listen to each pair in any order.
Song 1
Case-study pair corresponding to the highlighted regions in Figure 1.