Search papers, labs, and topics across Lattice.
This paper introduces MetaStrategy, a novel framework for generative ranking in recommender systems that generates structured, executable ranking strategies rather than direct item sequences. By leveraging a large language model (LLM) to produce a JSON bundle that controls various ranking objectives and constraints, the framework integrates seamlessly with existing predictive models and operational rules. In a real-world deployment on Taobao, MetaStrategy outperformed traditional methods, achieving significant improvements in user engagement metrics such as click page views and transaction amounts during a randomized A/B test.
MetaStrategy achieves a remarkable 27.93% win rate in generative ranking calls while enhancing user engagement metrics significantly, all without increasing response time.
Industrial recommender systems rank heterogeneous content under coupled user, business, commercial, and experience objectives. Existing generative ranking methods typically construct item sequences directly, making them difficult to integrate with mature predictive models, operational rules, and field-level guardrails. We present MetaStrategy, a framework that instead generates a structured, executable ranking strategy. Conditioned on request context, a large language model (LLM) policy emits a typed JSON bundle controlling objective weights, content and category preferences, experience constraints, and position policies. A deterministic validator and compiler instantiate an isolated Generator that competes atomically with incumbents under the list-level Evaluator of the Generator-Evaluator (GE) architecture. We train the policy in a production-path replay environment that re-executes logged requests through the current re-ranking stack without user exposure. The method combines selection, relative-rank, and baseline-lift rewards, a self-competitive curriculum that feeds frequent strategies back as competitors, and Evaluator-routed reward-augmented on-policy distillation that transfers complementary 4B-parameter Teachers into a compact 0.8B-parameter Student. We deploy MetaStrategy in Taobao Homepage Guess You Like through diff-triggered nearline generation; LLM inference remains outside synchronous ranking, with no observable increase in response time (RT). In a seven-day user-randomized online A/B test, MetaStrategy wins 27.93% of treatment-side GE calls and significantly improves click page views (click PV) by 2.11%, item-detail page views (IPV) by 3.12%, and transaction amount by 2.83%.