Search papers, labs, and topics across Lattice.
This paper introduces the FQTree algorithm for fine-grained quantization-aware training of boosted decision trees (BDTs), addressing the challenges of efficient hardware deployment in latency-critical applications. By employing a novel hardware-oriented leaf-value quantization scheme and integrating it with a compiler-based flow for automatic hardware generation, FQTree significantly reduces hardware costs while maintaining or enhancing model accuracy. Experimental results demonstrate a reduction in LUT usage by 26-57% compared to existing FPGA-based BDT designs, showcasing the effectiveness of the proposed approach.
FQTree slashes hardware costs for boosted decision trees by up to 57% without sacrificing accuracy, revolutionizing their deployment in latency-sensitive applications.
Boosted decision trees (BDTs) are widely used in latency-critical applications, but efficient hardware deployment remains challenging. Existing designs often rely on uniform or manually tuned fixed-point formats, which can introduce unnecessary hardware cost or accuracy loss. This work presents the FQTree algorithm{https://github.com/ecs-bristol/FQTree} for fine-grained quantization-aware training of BDTs, together with the QXGB framework for automatic hardware generation. FQTree introduces a hardware-oriented leaf-value quantization scheme that uses a global quantization step together with a tree-wise shift, enabling compact non-negative integer leaf representations, controlled clipping/pruning, and bias folding to reduce datapath cost. This work further applies this quantization during boosting so that later trees adapt to the errors of the already-quantized ensemble, and then lowers the trained model into low-latency hardware implementations through a compiler-based flow. Results on JSC, MNIST, and NID show that our method reduces LUT usage by 26-57\% compared with the state-of-the-art FPGA-based BDT designs while matching or improving accuracy.