Search papers, labs, and topics across Lattice.
100 papers published across 4 labs.
CoinRAG redefines efficiency in RAG by achieving higher accuracy with lower operational costs through innovative cache reuse strategies.
AutoPrune enables LLMs to autonomously design visual-token pruning strategies, achieving a remarkable 9.9x reduction in FLOPs without sacrificing performance.
WorldTrace redefines memory management in video world models, achieving up to 19.5% better episodic recall without any retraining.
Small LLMs can achieve up to 27.2% accuracy gains by leveraging hierarchical memory from larger teacher agents, reshaping how we think about agent training.
U-OPSD enables LLMs to self-improve without any external supervision, achieving up to 10.7% performance gains on challenging reasoning benchmarks.
CoinRAG redefines efficiency in RAG by achieving higher accuracy with lower operational costs through innovative cache reuse strategies.
AutoPrune enables LLMs to autonomously design visual-token pruning strategies, achieving a remarkable 9.9x reduction in FLOPs without sacrificing performance.
WorldTrace redefines memory management in video world models, achieving up to 19.5% better episodic recall without any retraining.
Small LLMs can achieve up to 27.2% accuracy gains by leveraging hierarchical memory from larger teacher agents, reshaping how we think about agent training.
U-OPSD enables LLMs to self-improve without any external supervision, achieving up to 10.7% performance gains on challenging reasoning benchmarks.
SkillZip achieves a remarkable 3.46x compression ratio while preserving 99.2% of dependencies and 98.7% of verifier reachability, revolutionizing how agent skill libraries can be managed.
BaKron accelerates neural network quantization by harnessing richer curvature information, achieving significant computational efficiency without sacrificing performance.
Adapting LLM inference scheduling to bursty traffic can boost throughput by leveraging real-time request intensity estimation.
Early stopping of accumulation in binary neural networks can slash computation by over 86% with minimal accuracy trade-offs.
Hybridizing autoregressive and speculative decoding, BALANCE boosts task throughput in edge LLM inference while managing latency and memory constraints.
Coordinated parallelism in multi-agent LLM systems can boost accuracy and cut latency, but only if applied judiciously—overdoing it may backfire.
A runtime observability framework reveals how modern memory architectures can be rigorously monitored and quantified, exposing hidden risks in AI model performance.
Flow-Map Distillation achieves superior image restoration by transforming static knowledge transfer into a dynamic flow mapping process, cutting training variance in half.
LiteKD-Net achieves superior image denoising performance on mobile devices while slashing runtime costs, setting a new standard for efficiency in the field.
TensorCast reveals that decoupling tensor management from computation can boost performance by over 90% in multi-turn interactions, challenging the status quo of LLM infrastructure.
SSTQ cuts communication costs in federated learning while ensuring privacy, achieving optimal mean squared error scaling with minimal bit usage.
Achieving 90% sparsity, BnBERT-iPET rivals larger models while drastically reducing computational costs for Bengali NLP tasks.
DIVE achieves an impressive 88.9% reduction in visual tokens without sacrificing performance, redefining efficiency in vision-language models.
RECAP achieves superior network traffic compression by learning rules that outperform expert designs, all while minimizing manual intervention.
Pruning Echo State Networks dynamically can enhance forecasting accuracy while significantly reducing model complexity.
Coordinating cache compression with a process reward can slash token generation by up to 65% without sacrificing accuracy.
Achieving a 3.56× increase in decoding throughput without sacrificing accuracy, BinaryPC revolutionizes efficiency in long-context LLMs.
Output extrapolation in STEP-OPD enables a unified student model to surpass the performance of its specialized teachers across all evaluated tasks.
StaticSegFormer boosts segmentation frame rates by 34% without compromising accuracy, challenging the dominance of dynamic pruning methods.
Retaining seemingly redundant tokens can actually enhance model performance, defying traditional assumptions about token importance in visual processing.
Adversarial attacks can degrade the efficiency of Vision Transformers, but MOAT ensures that performance remains nearly intact, limiting GFLOPs loss to just 3.4%.
A single distilled model can outperform larger heterogeneous teachers by effectively integrating their strengths without interference.
DBLAST significantly enhances the accepted draft length in stochastic decoding scenarios, particularly when the target distribution's entropy is high.
Recoverable eviction can drastically reduce missed attention and improve information retention in long-context decoding, outperforming traditional methods.
EdgeXpert slashes LLM inference latency by over 56% while cutting energy use by nearly 45%, all without sacrificing accuracy.
Relayed key-value caches can boost performance to 100% when private information is needed, but irrelevant relays plummet to just 23-25%.
SpecRoll achieves up to 2.15x faster generation in RL rollouts by cleverly balancing fast and slow adaptation strategies.
Backend choice can distort benchmark scores by nearly 40%, challenging the assumption that model performance is solely a property of the model itself.
Achieving up to 2.72x faster inference times, RAC transforms split LLM deployment by slashing communication bottlenecks without sacrificing performance.
AsymSpec boosts output-token throughput by up to 28 times by cleverly optimizing communication between edge and cloud models.
Deltoris achieves a staggering 34.2× speedup for real-time VLA inference, revolutionizing how embodied AI can operate on edge devices.
Adversarial attacks can exploit input-adaptive optimizations in Vision Transformers, undermining their efficiency without sacrificing accuracy.
SPOT redefines on-policy distillation by ensuring that probing decisions directly enhance downstream reasoning performance, not just teacher alignment.
Hard prompt compressors can leave critical context gaps, leading to a staggering 60% of examples suffering from referential dangling, which severely impacts accuracy in multi-hop question answering.
Bridging disparate model families, Any-OPD achieves a 4.5% increase in performance while reducing model size by 80%.
Reusing KV caches during model swaps can retain up to 98% of prefill accuracy while running 2.7-25x faster than traditional methods.
Prompt design and scoring rules can dramatically alter the perceived reliability of biomedical language models, with calibration errors swinging by over 200%.
CausalOPD reduces the right-label-wrong-reasoning rate from 15.7% to 4.4%, showcasing a breakthrough in accurate causal reasoning for AI models.
Achieving high task performance with ultra-low-rank adapters, SALT recovers accuracy while slashing memory usage by up to 16x.
Accepting selective mismatches in autoregressive decoding can boost throughput by over 15% without any additional training or model adjustments.
Bridging the gap between ANNs and SNNs could revolutionize federated learning on resource-constrained devices, achieving high accuracy without sacrificing efficiency.
SAKI achieves up to 30% better recall than traditional key PCA methods by directly preserving attention scores, not just key variance.
A semantic re-keying strategy in speculative decoding can boost accepted draft lengths by up to 29% while achieving 4.4x faster decoding speeds.
Pruning the vocabulary of multilingual models can lead to a 60% memory savings without sacrificing translation quality, challenging the need for large vocabularies in MNMT.
Pruning 75% of visual tokens without sacrificing performance could redefine efficiency benchmarks for VideoLLMs.
AcceptMoE slashes host-to-device traffic by over 73% while boosting throughput by more than double, all without a major accuracy trade-off.
SparSEEty reveals that even LLMs in secure environments can be vulnerable to token extraction attacks through clever exploitation of activation sparsity.
Reducing visual tokens doesn't always mean faster inference; a pre-vision strategy can significantly cut latency by bypassing preprocessing steps.
KeepAD achieves a 7.9× speedup in zero-shot anomaly detection while preserving crucial defect information, challenging the trade-off between efficiency and accuracy.
ACT achieves a remarkable 6.4x reduction in inference latency while maintaining task success, even on entry-level embedded hardware.
By eliminating the need for iterative optimization, ProtoBlend achieves high-quality video dataset distillation while dramatically reducing computational costs.
SRG transforms dataset distillation by aligning generative samples with the discriminative geometry of self-supervised representations, achieving superior performance across multiple benchmarks.
Achieving superior inference accuracy and resource efficiency in multi-cluster edge intelligence networks hinges on a delicate balance between model pruning and collaborative inference.
CAMTA achieves nearly an order of magnitude improvement in Softmax function approximation while offering unprecedented runtime configurability in hardware.
OmniPack achieves a remarkable 98% performance retention with a staggering 83.3% reduction in computational load, revolutionizing token compression for omni-modal models.
iFAN boosts segmentation accuracy by aligning query competition with mask quality, achieving significant performance gains without extra computational overhead.
SPADE achieves up to 3.40x faster attention processing in video diffusion models while maintaining high-quality outputs, revolutionizing the efficiency of video generation.
SlimVLM achieves unprecedented efficiency in Vision-Language Models by intelligently pruning redundant visual tokens without sacrificing performance.
Targeted optimization of normalization affine parameters can dramatically enhance low-bit quantization performance, breaking the limits of conventional training methods.
ALiBi positional encoding can blind attention heads, drastically impairing token retrieval without impacting standard performance metrics.
vLLM emerges as the leading framework for LLM serving, but developers are missing out on the advantages of multi-framework integration.
PhyAI achieves up to 4.65x speedup in Physical AI tasks by unifying disparate inference processes into a single, efficient runtime.
Relocating the KV cache to processing-near-memory nodes can boost LLM throughput by over 6x while supporting evolving sparse attention methods.
Disaggregating LLM inference stages can boost throughput by up to 75%, reshaping how we design future AI hardware systems.
AdaMX slashes accuracy loss in low-bit LLM inference by 83% while keeping energy costs minimal.
Ternary LLMs can now achieve efficient attention computation without the overhead of high-precision K/V processing, revolutionizing their performance.
Static power consumption can skew efficiency estimates by up to 3.85X, revealing critical oversights in current PIM-GPU design practices for LLM inference.
Adversarial training boosts robustness against input attacks but paradoxically makes models more vulnerable to hardware faults.
Filtering out misleading signals can boost OPD performance by leveraging input-groundedness, leading to more effective model training.
Offline KD can achieve the same training loss as online methods while being 29% faster, revolutionizing how we approach model distillation efficiency.
Gecko achieves private inference in just 0.4-2.2 seconds while maintaining robust security against model extraction attacks.
Hardware-aware training can recover predictive performance in photonic Bayesian neural networks, but only if the required variational family remains representable.
A single shared intervention in the query-key channel can stabilize low-precision transformer training, eliminating the need for individual fault repairs.
Teacher guidance can be strategically enhanced by targeting high-disagreement states, leading to significant performance gains in agentic tasks.
Anomaly detection just got faster—CARE achieves up to 4.8x inference speedup without sacrificing accuracy by intelligently filtering normal data.
xPress boosts acceptance lengths by 30% and decoding throughput by 1.3 times, transforming how block-diffusion drafters handle causal dependencies.
Parent-Conditioned Drafting can boost LLM inference speed by up to 29.5% while increasing the effective acceptance length of generated outputs.
ET-Prune achieves a remarkable balance between efficiency and accuracy, outperforming traditional pruning methods by retaining critical evidence while cutting down on unnecessary tokens.
Bole accelerates hybrid-attention LLMs by up to 4.72 times while slashing memory usage by up to 99 times, transforming the landscape of autoregressive decoding.
The study reveals that maintaining answer accuracy in large models can come at the cost of losing critical reasoning support, highlighting a significant "answer-evidence gap" in KV cache compression.
Retaining nearly all of a model's capability while slashing visual token usage by over 80% reveals a transformative approach to VLM compression.
Token pruning can be done more effectively by directly linking token scores to their utility, achieving high accuracy with significantly reduced computational costs.
SmartGR achieves an 8.6% boost in recommendation performance while slashing inference time by over 2.3 times, tackling unique challenges in generative recommendation systems.
DAPD eliminates privilege illusion in self-distillation, leading to consistent performance improvements across model scales.
Achieving a balance between high detection accuracy and low latency, this IDS framework sets a new standard for cybersecurity in EV charging networks.
Achieving up to 1.8x faster training on consumer GPUs, Meganeura redefines the efficiency of portable AI model deployment across diverse hardware platforms.
TELLER achieves over 80% trace length reduction while maintaining high diagnostic accuracy, revolutionizing root-cause analysis for LLM inference.
Achieving up to 40.3% cost savings in LLM inference through optimal prefix key-value placement could redefine resource allocation strategies in AI deployments.
ISJL strikes an optimal balance between throughput and cost alignment, outperforming traditional batching methods in LLM serving.
Energy consumption in LLM serving can be cut by nearly half without sacrificing performance, thanks to a new framework that intelligently manages GPU frequency scaling.
PrefixShield transforms how multi-tenant LLMs manage shared resources, achieving up to an 84.87% victim cache hit ratio by addressing the admission-responsibility gap.
Achieving 2.19x throughput and 2.03x energy efficiency improvements on FPGA for SSMs by optimizing projection rank could redefine performance benchmarks in latency-sensitive applications.
ARCHead slashes LM-head storage by up to 3.9x while preserving near-optimal performance, revolutionizing how we think about model efficiency.
Triage performance among on-device models is an illusion; even simple agents can achieve perfect scores, undermining the belief that bigger models always outperform smaller ones.
Achieving nearly 13 times faster inference while ensuring privacy could redefine how IoT devices handle sensitive data in real-time applications.