Search papers, labs, and topics across Lattice.
This paper introduces OGR, an end-to-end generative framework for slate recommendation that integrates item-specific semantic and local collaborative information through a novel TUSID construction. By employing list-wise preference planning and a reward-guided conservative policy optimization method, OGR aligns generated slates with user preferences, addressing the limitations of traditional candidate generation and ranking approaches. Experimental results demonstrate significant improvements in recommendation effectiveness, achieving relative NDCG@5 gains of 48.2% and 27.2% on industrial and public datasets, respectively, alongside a notable increase in Effective Views during online testing.
OGR achieves a remarkable 48.2% boost in recommendation effectiveness by seamlessly integrating semantic and collaborative signals into slate generation.
Slate recommendation treats a slate rather than an individual item as the recommendation unit, requiring joint optimization of item interactions and slate utility. Existing approaches typically separate candidate generation from ranking and restrict optimization to retrieved candidates. Generative recommendation with Semantic IDs (SIDs) offers a path to end-to-end recommendation, but existing SID construction often lacks recommendation-aware semantics and effective local collaborative signals, while next-token prediction is misaligned with slate-level objectives. We propose OGR, an end-to-end framework that directly generates ordered slates-"Once Generated, Ranked."OGR first introduces TUSID, which adaptively fuses item-specific semantic and local collaborative information into hierarchical SIDs. It then uses list-wise preference planning and pipelined position-wise SID decoding to model global preferences and inter-item dependencies while generating ordered slates. We further propose SPA, a reward-guided conservative policy optimization method that aligns generated slates with user preferences beyond likelihood imitation. Offline experiments show that OGR outperforms representative baselines, with 48.2% and 27.2% relative NDCG@5 gains on industrial and public datasets, respectively. Online A/B testing on Kuaishou further yields a 1.120% improvement in Effective Views.