Search papers, labs, and topics across Lattice.
Affiliation:
1
0
2
0
Achieving 1.79x the throughput of unsplit models, this method enables interactive inference of 70B-parameter LLMs on distributed Intel AI PC fleets that individually lack the memory capacity.