Search papers, labs, and topics across Lattice.
Affiliation:
4
0
5
HBQ achieves state-of-the-art efficiency in LLM inference, delivering up to 2.3x higher area/energy efficiency while maintaining accuracy levels that challenge conventional methods.
Achieving near-baseline accuracy for large-scale models on ReRAM architectures with minimal retraining could revolutionize energy-efficient AI deployment.
Operator-level orchestration can yield up to 3.42x speedup in edge inference by intelligently mapping operators to the best processing units.
MOSAIC uncovers a 46.91% energy savings in heterogeneous NPUs, challenging the dominance of homogeneous designs in AI workloads.