Search papers, labs, and topics across Lattice.
66 papers published across 5 labs.
NanoSleep achieves superior sleep stage classification accuracy while being small enough for deployment on wearable devices, striking a crucial balance between performance and efficiency.
Compressing VLMs to a mere 3.7 GB without sacrificing performance could revolutionize mobile AI applications.
QAH enables 4-bit LLMs to outperform their bfloat16 counterparts while reducing memory usage by four times and achieving peak performance seven times faster than traditional methods.
Truncating low-ranked completions can outperform traditional rank-based policies by reducing reliance on brittle top rankings while still achieving strong alignment results.
Achieving a 4% accuracy boost in CNNs without changing datapath width while enhancing energy efficiency by 2.5x could revolutionize industrial visual inspection.
Compressing VLMs to a mere 3.7 GB without sacrificing performance could revolutionize mobile AI applications.
QAH enables 4-bit LLMs to outperform their bfloat16 counterparts while reducing memory usage by four times and achieving peak performance seven times faster than traditional methods.
Truncating low-ranked completions can outperform traditional rank-based policies by reducing reliance on brittle top rankings while still achieving strong alignment results.
Achieving a 4% accuracy boost in CNNs without changing datapath width while enhancing energy efficiency by 2.5x could revolutionize industrial visual inspection.
ML-based data compression can be environmentally sustainable, but only if it surpasses a critical break-even point in carbon savings.
Achieving a 2.3x performance improvement in LLM serving while increasing cache hit rates to over 93% could redefine efficiency benchmarks in large-scale AI deployments.
Fine-tuning with sparse attention can outperform traditional exact attention models while running efficiently on modest hardware.
Compression can lead to a significant loss of critical knowledge while leaving models confidently incorrect, revealing hidden biases that standard metrics fail to detect.
Intermediate context compression can cut GPU energy usage by over 50% while maintaining quality, challenging static compression strategies in edge RAG applications.
Recalibrating the CFG scale in response to CIM noise can restore over 87% of generation quality lost due to nonidealities in Diffusion Transformers.
Achieving real-time EEG auditory attention decoding with a power-efficient ASIC could revolutionize hearing assistance for cochlear implant users in noisy settings.
ReCache achieves a staggering 92.43% reduction in KV-tensor memory while maintaining nearly identical performance in tool-augmented language models.
Daedalus-150M outperforms larger competitors while being optimized for CPU inference, achieving faster decoding and lower memory usage.
Achieving up to 47.26x speedup in long-context LLM serving could redefine efficiency benchmarks in AI inference.
GEAR slashes inference time by up to 2866 times while boosting AUC scores beyond conventional supervised models.
FlashAttention-V achieves up to 42x speedup in transformer inference on CPUs, transforming how we leverage vector architectures for small language models.
INT4 quantization can reduce LLM accuracy by over 12% in high interference scenarios, challenging the assumption that lower precision always maintains performance.
NanoSleep achieves superior sleep stage classification accuracy while being small enough for deployment on wearable devices, striking a crucial balance between performance and efficiency.
WhiteMatter achieves superior performance with fewer resources by allowing each attention layer to dynamically access all previous layer representations, challenging traditional fixed connection patterns.
Compact object detectors can achieve state-of-the-art accuracy without the bulk of larger models, thanks to a novel multi-level knowledge distillation approach.
APEX achieves ANN-equivalent accuracy with 40% energy savings, revolutionizing the efficiency of Spiking Neural Networks in practical applications.
Group-Calibrated On-Policy Distillation boosts long-context reasoning performance by reconciling teacher guidance with verifier feedback, achieving up to a 12-point increase in benchmark scores.
Local AI inference may democratize access, but it also shifts power dynamics, placing control in the hands of hardware vendors and core maintainers.
Achieving 1.79x the throughput of unsplit models, this method enables interactive inference of 70B-parameter LLMs on distributed Intel AI PC fleets that individually lack the memory capacity.
Non-uniform bit allocation can boost recall by up to 18% in low-bit quantization, reshaping how we approach embedding compression.
Selective re-scanning in recurrent networks can drastically reduce memory usage while improving task performance, challenging the conventional wisdom of fixed-size state fidelity.
FESC achieves private long-document inference on a single GPU for sequences up to 2,048 tokens, setting a new standard for efficiency and accuracy in encrypted machine learning.
Achieving over 90% accuracy retention with 50% pruning, DVBP + OB²C outperforms traditional methods by leveraging noise-filtered neuron selection and optimal weight updates.
Achieving near state-of-the-art performance with a 22M encoder, DistillPath-KS16 runs over 25 times faster than its larger counterparts while retaining critical accuracy in pathology tasks.
MoNe slashes compute and memory costs by 80% for long-context inference while enabling Transformers to handle context lengths far beyond their original limits.
Energy-aware knowledge distillation can slash inference energy consumption by up to 90%, challenging the reliability of FLOPs as a metric for sustainability in LLMs.
Achieving up to 11× reduction in storage for dynamic scene streaming without sacrificing quality or speed could revolutionize online video applications.
Achieving a staggering 63.959x external-to-live-staging ratio, BSR revolutionizes how we manage long-context LLM states in memory-constrained environments.
Eliminating semantic redundancy in HGNNs can boost sampling performance by an order of magnitude, transforming mini-batch inference efficiency.
FLEXRec shows that compact LLMs can achieve state-of-the-art recommendation accuracy without the computational burden of larger models.
TileMix achieves a breakthrough in LLM inference by enabling mixed-precision attention that boosts throughput while maintaining long-context quality.
Nexus achieves high-resolution text-to-image generation with a fraction of the computational cost, rivaling leading models in quality.
Achieving a 96% reduction in computation for RAW video restoration with minimal performance loss opens new avenues for efficient video processing in real-time applications.
Pallas cuts service interruption time by up to 89.68 times during LLM inference handovers, transforming mobile AI experiences.
Localized TabICLv2 retains nearly all the accuracy of its predecessor while slashing inference time, making it a game-changer for large-scale tabular data tasks.
Sparsifying activations in collaborative inference may cut costs, but it exposes a hidden privacy risk from the positions of those activations that could enable re-identification.
Index-side pruning can slash retrieval latency by up to 6.6x across diverse engines, while query pruning is largely redundant in modern systems.
SOPD not only outperforms traditional distillation methods but also redefines how we think about trajectory corrections in model training.
Latent-OPD reveals that distilling latent representations at trajectory endpoints can dramatically enhance video reasoning efficiency in LMMs.
A single learned codec can adapt to diverse tasks on-the-fly, achieving near task-specific performance without retraining.
Expanding LLM scheduling from two to multiple priority tiers can yield up to 8.3x faster inference while significantly lowering costs.
Achieving 3.4x throughput over static sharding under skewed loads while ensuring zero task loss during worker failures could revolutionize large-scale data labeling.
Shared multi-agent search in KernelArc outperforms traditional methods, achieving top rankings in GPU kernel optimization tasks while maintaining a fixed candidate budget.
Posits may not be the superior alternative to IEEE floating point for solving chaotic N-body problems, despite their theoretical advantages in precision.
FreeToken transforms personal machines into powerful platforms for running massive AI models, enabling users to deploy frontier-scale intelligence without specialized infrastructure.
Multi-byte prediction accelerates byte-level language model inference without sacrificing performance, achieving a breakthrough in generative task efficiency.
Retained KV cache can undermine rollback consistency in language agents, leading to unexpected model behavior even after logical aborts.
Recovering 87% of accuracy lost to KV eviction could redefine efficiency in reasoning tasks for large language models.
Routing changes in MoE models may not influence behavior as expected, with the routing term contributing less than half of the natural context effect.
Targeted bit-flip attacks can cripple Vision-Language-Action models, reducing their success rates to zero with just a few flips in key layers.
SpecVLA achieves real-time robotic manipulation by enabling long-action-length predictions with timely verification, cutting down latency without sacrificing reliability.
FlashQuant achieves up to 4.18x faster outlier-aware LLM inference by fusing dense and sparse computations, addressing a critical bottleneck in memory efficiency.
Reordering attention and feed-forward operations in LLMs can significantly enhance computational efficiency without sacrificing performance.
EcoVLA boosts energy efficiency for VLA models by up to 236% while ensuring real-time performance, revolutionizing how robotic systems manage inference costs.
KV-Pipe transforms KV reuse into a powerful tool for balancing pipeline workloads, yielding significant efficiency gains in both training and inference.
Tailoring compression strategies to individual client capabilities can cut communication overhead in federated learning by a significant margin.
DeltaLog slashes recurrent-state write traffic by up to 7.83x while boosting decoding speed, revolutionizing how linear attention models manage memory.
Short-context models can achieve superior reasoning performance by leveraging long-context teacher models through innovative token alignment and training strategies.
Fixing deployment bugs in Nanbeige4.2-3B transforms it from a non-functional model to one that can tackle real agentic tasks with a significant performance boost.
Subtracting information from the student rather than adding it to the teacher can yield performance improvements that rival those achieved with privileged data.
A lineage verification method that distinguishes model ancestry with perfect accuracy, even under aggressive checkpoint modifications.