Search papers, labs, and topics across Lattice.
This paper introduces StrixAE, an intelligent agent designed for audio enhancement that effectively addresses complex distortion couplings in real-world scenarios. By utilizing a multimodal large language model (MLLM) as a controller, StrixAE coordinates various audio enhancement and personalization models, enhancing system robustness and generalization. The agent's two-stage training process, which includes CoT supervised fine-tuning and a novel Audio Perception Reinforcement Learning (APRL), results in state-of-the-art performance across multiple perceptual metrics, outperforming existing solutions.
StrixAE achieves unprecedented audio enhancement performance by integrating a multimodal language model with a tailored reinforcement learning approach, setting a new benchmark in real-world audio restoration.
Audio enhancement in real-world scenarios involves complex distortion couplings and requires personalized enhancement. Existing solutions struggle to address both simultaneously. To improve robustness and enable autonomous operation in such scenarios, we propose StrixAE, an agent based on a multimodal large language model (MLLM). StrixAE leverages the MLLM as a controller to coordinate multiple audio enhancement and personalization models. To further enhance system robustness, reduce artifacts, and improve generalization across diverse real-world scenarios, StrixAE is trained through a two-stage process: first, CoT supervised fine-tuning on AcoustBench to ground basic reasoning and tool invocation; second, Audio Perception Reinforcement Learning (APRL), a reward design specifically tailored for audio restoration pipelines that jointly optimizes format validity, structural coherence, and perceptual quality. Unlike generic RL fine-tuning, APRL introduces structured rewards that enforce executable pipelines and logical section ordering, enabling the agent to produce reliable, interpretable enhancement plans without hallucinated tools. Based on real-world test datasets, our proposed method outperforms most existing open-source and proprietary solutions, achieving state-of-the-art performance across multiple perceptual metrics and demonstrating strong generalization robustness.