Search papers, labs, and topics across Lattice.
59 papers published across 1 lab.
TP-MPPO achieves up to 87.5% higher goodput for LLM inference in edge networks, revolutionizing how we manage bandwidth and task offloading.
Aggregate latency masks crucial insights into LLM inference efficiency, revealing that backend changes and quantization effects significantly influence performance metrics.
Pruning can significantly undermine sparse autoencoder performance, but activation-aware methods offer a robust alternative that preserves interpretability in LLMs.
Pointwise convolutions, which dominate parameter volume in large-kernel CNNs, can be drastically reduced through a novel group-sharing strategy, enabling efficient deployment on edge devices.
Adaptive self-distillation can boost model performance by over 23 points without extra rollouts, reshaping how we approach teacher-student dynamics in training.
Pruning can significantly undermine sparse autoencoder performance, but activation-aware methods offer a robust alternative that preserves interpretability in LLMs.
Pointwise convolutions, which dominate parameter volume in large-kernel CNNs, can be drastically reduced through a novel group-sharing strategy, enabling efficient deployment on edge devices.
Adaptive self-distillation can boost model performance by over 23 points without extra rollouts, reshaping how we approach teacher-student dynamics in training.
Compressing neural networks just got a theoretical boost with a new framework that integrates weighted graphs and algebraic structures.
Collapse in On-Policy Self-Distillation narrows reasoning paths, revealing critical biases that could undermine model performance.
Low-probability tokens disproportionately influence model updates, and a simple reweighting strategy can significantly enhance performance without sacrificing generalization.
Adjusting adapter rank in QLoRA reveals a critical trade-off between factual acquisition and retention of unrelated capabilities, challenging assumptions about parameter-efficient fine-tuning.
Achieving over 98% accuracy with a compact model, CropCop sets a new standard for plant-health recognition while ensuring data integrity through rigorous auditing.
Foresight pruning can significantly enhance the performance of sparse PINN solvers by focusing on the sensitivity of PDE residuals rather than just output dynamics.
AsymSpec achieves 90% accuracy with 1.7x speedups by leveraging asymmetric context access, redefining efficiency in agentic LLMs.
TP-MPPO achieves up to 87.5% higher goodput for LLM inference in edge networks, revolutionizing how we manage bandwidth and task offloading.
LLMs' personalities are dynamic and layer-dependent, revealing that quantization can significantly disrupt their behavioral consistency.
TOPAS slashes job completion times by up to 49.4% in multi-agent LLM serving by intelligently balancing prefix caching and request scheduling.
Critical visual token selection is driven by a select few attention heads, and leveraging this insight can drastically improve VLM efficiency without sacrificing performance.
APT accelerates high-resolution diffusion models by up to 8.16× while enhancing energy efficiency, revolutionizing the feasibility of real-time generative AI applications.
Achieving 98% computational savings while improving AEC performance reveals the untapped potential of knowledge distillation in audio processing.
Extracting LLM assets from edge AI chips is feasible through laser voltage imaging, exposing critical vulnerabilities in current deployment practices.
Ternary multiplicative adaptation recovers lost performance in quantized models while maintaining extreme efficiency, outperforming traditional low-bit methods.
Achieving over five times improvement in Latency-Error-Energy metrics could redefine efficiency standards for edge-deployable virtual sensing systems.
ResiSpec redefines speculative decoding efficiency, achieving nearly double the speed of current methods without sacrificing output quality.
Quantization can enhance Bangla language model deployment, but the architecture and quantization method can dramatically influence performance, with some models losing over half their accuracy.
AgentSpec slashes response times for LLM agents by addressing high rejection rates and optimizing token budgets, outperforming existing methods.
MaST achieves nearly double the tracking speed of its predecessor while setting new benchmarks for accuracy in lightweight object tracking.
Tailored noise in multi-teacher distillation can yield student models that rival those trained on real data, even with just 1K images.
Achieving a staggering 54.67× speedup in text-to-video-audio generation without sacrificing quality could revolutionize real-time multimedia applications.
Correction capability varies dramatically across parameter groups, with normalization-affine parameters offering the most efficient path to improved quantization robustness.
Achieving up to 2.35× faster inference with only 19–28% of the original KV cache, VisCache redefines efficiency in Vision Large Language Models.
Retaining the largest attention weights in KV-cache eviction is nearly optimal, challenging the perceived complexity of selection strategies.
Targeted mixed-precision strategies can cut CFD simulation time and energy costs by over a third without sacrificing accuracy.
Depth pruning doesn't have to mean sacrificing accuracy—SHIFT-LLM recovers lost performance with minimal overhead.
FLINT transforms LLM inference by integrating high-bandwidth flash, overcoming memory constraints that limit model deployment and performance.
Aggregate latency masks crucial insights into LLM inference efficiency, revealing that backend changes and quantization effects significantly influence performance metrics.
Jointly applying sparsity, quantization, and low-rank approximations can yield up to 5.66% better accuracy than the best existing methods for LLMs.
Degrading LLM service to save costs can actually backfire, leading to increased churn and inflated demand during peak times.
ProxyFormer achieves a staggering 0.7 million token context length while retaining up to 95% retrieval accuracy, revolutionizing the scalability of generative models.
Achieving up to 3.68x faster inference and 5.12x fewer FLOPs, ChebBooster redefines efficiency in Diffusion Transformers without the need for additional training.
Sigmoid attention, while less effective in dense modeling, dramatically improves KV-cache eviction performance, challenging conventional wisdom about attention mechanisms.
RoI slashes the parameter overhead of semi-structured sparsity by up to 8.75 times, paving the way for more efficient large language model deployment.
Closing nearly 90% of performance gaps in low-bit quantized models with a compact repair codec could redefine efficiency in language model deployment.
Amortizing distillation across model size and variant can produce a continuum of optimized LLMs with a single distillation run, revolutionizing model deployment efficiency.
Pruning data with MCL not only boosts efficiency but also reveals hidden semantic structures that traditional methods overlook.
E2S-Pruner retains up to 98% of performance while drastically improving inference speed, redefining efficiency in vision-language models.
SA-RSQ achieves a remarkable balance between compact storage and high-quality representation, leading to significant performance boosts in real-world recommender systems.
SelFusion enables diffusion language models to surpass traditional autoregressive models in generation quality by leveraging a novel self-distillation approach.
WnW reduces GPU memory usage to 20% of audio tokens without sacrificing accuracy, challenging the limitations of existing KV cache methods in long-form speech processing.
Dynamic visual evidence retrieval can accelerate multimodal decoding by over 2x while enhancing draft acceptance rates.
Pruning models can lead to smaller, more accurate prediction sets without sacrificing reliability, as shown by CPP's impressive performance on large-label tasks.
ROBBIN achieves nearly 90% attack success while preserving over 83% accuracy, making it a game-changer for reliable backdoor attacks across varying DRAM devices.
Pruning visual tokens based on head alignment can retain nearly all performance while drastically reducing computational costs.
Intermediate knowledge distillation can dramatically improve performance in data-scarce environments, turning the conventional wisdom on its head.
Robots can now learn and adapt in real-time, overcoming inference delays that previously stymied reinforcement learning effectiveness.
Achieving a semantic demand lower bound of 43.59375 GiB reveals that traditional memory limits can be surpassed without sacrificing execution accuracy in AI inference.
Achieving a 2.00x reduction in weight bandwidth while decoding at 5.94 tokens/s could redefine efficiency benchmarks for autoregressive models on CPU architectures.
Reclaiming memory from idle KV cache during decode phases yields negligible latency improvements, challenging assumptions about prefill chunk sizes in LLM serving.
Predictive delta prefetching can achieve up to 12x faster LLM inference on edge devices, even when models exceed memory limits.
NOVA's innovative architecture delivers 4.5x higher throughput and 69.8% lower latency for hybrid LLMs, pushing the boundaries of memory processing efficiency.
Targeted approximation in floating point multipliers can yield up to 92% hardware footprint savings without sacrificing CNN accuracy.
Compressing VLMs to a mere 3.7 GB without sacrificing performance could revolutionize mobile AI applications.
QAH enables 4-bit LLMs to outperform their bfloat16 counterparts while reducing memory usage by four times and achieving peak performance seven times faster than traditional methods.