Search papers, labs, and topics across Lattice.
This paper introduces RPPNet, a two-stage deep learning architecture that generates melodies by leveraging Rhythm-Pitch Primitives (RPPs) with variable structural boundaries, addressing the misalignment between human musical perception and traditional bar-based notation. By deriving RPP groupings from acoustic cues and psychological principles, RPPNet produces melodies that exhibit enhanced long-term structure and musicality compared to existing models. Experimental results demonstrate significant improvements in subjective evaluations, indicating that the psychological representation of musical structure is crucial for effective melody generation.
Melodies generated by RPPNet outperform traditional models by aligning more closely with human perception of musical structure, revealing the importance of psychological factors in music generation.
Existing symbolic music generation models typically use bars as the basic structural unit. However, human perception of musical phrases often does not align with notated bar lines, leading to long-term structural fragmentation. This paper proposes RPPNet-a two-stage deep learning architecture with variable structural boundaries. It first generates variable-length Rhythm-Pitch Primitive (RPP) sequences, where each RPP encodes note count, rhythm, and contour; then decodes the RPP sequences into concrete notes. The grouping of RPPs is automatically derived from acoustic cues, auditory inertia, and similarity perception based on music psychology. Experiments show that melodies generated by RPPNet are superior in both long-term structure and musicality, with significant improvements across all subjective evaluation dimensions. Ablation studies confirm that the performance gain stems from the structural correctness of the psychological representation, rather than from model capacity. This work offers an interdisciplinary perspective for music generation, integrating music theory, computational modeling, and music psychology.