Search papers, labs, and topics across Lattice.
90 papers published across 6 labs.
Short-context models can achieve superior reasoning performance by leveraging long-context teacher models through innovative token alignment and training strategies.
Fixing deployment bugs in Nanbeige4.2-3B transforms it from a non-functional model to one that can tackle real agentic tasks with a significant performance boost.
Subtracting information from the student rather than adding it to the teacher can yield performance improvements that rival those achieved with privileged data.
A lineage verification method that distinguishes model ancestry with perfect accuracy, even under aggressive checkpoint modifications.
RMM reveals that optimizing attention-side computations can lead to substantial runtime gains in Transformer models without sacrificing accuracy.
Short-context models can achieve superior reasoning performance by leveraging long-context teacher models through innovative token alignment and training strategies.
Fixing deployment bugs in Nanbeige4.2-3B transforms it from a non-functional model to one that can tackle real agentic tasks with a significant performance boost.
Subtracting information from the student rather than adding it to the teacher can yield performance improvements that rival those achieved with privileged data.
A lineage verification method that distinguishes model ancestry with perfect accuracy, even under aggressive checkpoint modifications.
RMM reveals that optimizing attention-side computations can lead to substantial runtime gains in Transformer models without sacrificing accuracy.
Certified caching can boost edge image classification speed by 1.65x without compromising reliability.
HBF can significantly boost LLM serving efficiency by enabling more expert replicas and reducing loading times, all while preserving critical execution paths.
DARTree achieves a staggering 9.73× speedup in autoregressive decoding while accepting nearly 99% more tokens per round than existing methods.
Despite promising theoretical foundations, RoPE-aligned rotations fail to improve quantization accuracy, highlighting a critical misalignment in current methods.
GCache achieves a remarkable 2.17x speedup in video diffusion while enhancing visual quality, challenging the effectiveness of traditional caching heuristics.
Online inference in QTD can now be performed efficiently without the need to store entire trajectories, revolutionizing memory management in distributional reinforcement learning.
Achieving nearly 2,000-fold context compression, BAPS enables pretrained tabular models to handle million-row datasets without retraining.
SNIPER achieves a remarkable CRAFT score of 0.98, ensuring near-perfect adherence to compression budgets while maintaining model performance.
INT4 quantization can distort the architecture landscape, but a zero-shot FP32 surrogate offers a surprising advantage in maintaining Pareto optimality.
vToken slashes KV memory retention by over 70%, dramatically boosting throughput and concurrency in large language model serving.
Coverage-driven token pruning can significantly enhance the efficiency of 3D VLMs without sacrificing reasoning capabilities.
SPADE slashes cloud model calls by 76% while preserving accuracy, revolutionizing the deployment of large language models in edge environments.
Expert-aligned drafting outperforms contrastive-aware methods, leading to up to 12x faster proposal paths in decoding.
Misalignment in speculative decoding can lead to a staggering 21-frame error in ASR, but innovative tracking methods can significantly boost efficiency and accuracy.
Aligning teacher supervision with causal constraints leads to state-of-the-art performance in autoregressive video generation.
Activation calibration can make or break predictive performance in low-bit quantization for financial forecasting, especially when moving from 8-bit to 4-bit models.
SoftWater slashes quantization error by up to 8.3x while maintaining near-lossless performance in language models, revolutionizing softmax layer efficiency.
TREX enables compact models to match or exceed the performance of large foundation models while dramatically speeding up inference times.
OPD may improve sampling efficiency, but it risks making previously solvable problems unsolvable, challenging the notion of true capability expansion in LLMs.
Consolidator transforms how memory is retained and accessed, boosting recall of updated information by over 42 percentage points without sacrificing short-term performance.
BoltNet achieves state-of-the-art plant species identification accuracy with a fraction of the model size, making it ideal for on-device deployment.
ASD reveals that leveraging a teacher's optimization trajectory can significantly enhance student model performance by actively suppressing shortcut features.
XYZFlow achieves up to 8.5X speed improvements in generative modeling without compromising image quality, redefining the efficiency landscape in high-fidelity image generation.
Proactively committing mid-entropy pivot positions can accelerate dLLM decoding by up to 18 times while improving accuracy.
Language-Conditional Dequantization recovers up to 83% of the perplexity gap for non-Latin languages, challenging the notion that quantization is uniformly detrimental across languages.
Test-time harnesses can nearly double the performance of weaker models, transforming how we think about capability transfer in AI.
A single-layer speech enhancement model outperforms naive architectures and achieves competitive quality with a significant speedup through progressive knowledge distillation.
Compressing a reliable large model via quantization yields Small Language Models that are not only more trustworthy but also more adaptable than those trained from scratch.
Achieving nearly 5x model compression with minimal quality loss, HAMP-LIC sets a new standard for efficient learned image compression across heterogeneous hardware.
FQTree slashes hardware costs for boosted decision trees by up to 57% without sacrificing accuracy, revolutionizing their deployment in latency-sensitive applications.
User-assisted collaborative inference can slash dedicated resource usage while boosting performance as demand scales.
INT8 support on NVIDIA's Blackwell Ultra GPU is effectively non-existent despite being listed in specifications, revealing a critical gap between hardware promises and practical usability.
APEX achieves over 99% overlap accuracy in expert prefetching, slashing per-token latency by up to 26% while enhancing energy efficiency for edge MoE inference.
Achieving near-zero overhead in real-time detection pipelines could revolutionize edge-deployed vision systems by enabling efficient concurrent model execution without sacrificing accuracy.
Lonic achieves up to 66.28x energy efficiency improvements over leading GPUs, revolutionizing the training landscape for spiking neural networks.
Replacing SSDs with High-Bandwidth Flash in LLM serving can paradoxically slow down performance by over 5 times due to mismatched workload characteristics.
Achieving a staggering reduction in hardware footprint while maintaining near-perfect accuracy, Uni-SFU redefines the efficiency of activation function implementations in neural networks.
SIEVE retains 97.5% of performance while reducing visual token usage to just 11.1%, revolutionizing efficiency in vision-language models.
Pruning 50% of channels in RGB-infrared object detectors can actually boost performance by 0.6% mAP, challenging conventional wisdom about redundancy.
iBKD outperforms traditional distillation methods by preserving spatial grid structures, enabling Vision Transformers to excel even with limited training data.
Legal definitions of inference in EU digital law diverge significantly, revealing a gap that could expose AI systems to unanticipated regulatory scrutiny.
Gated VLA-Cache recovers lost accuracy in real-time control while slashing compute costs by leveraging model uncertainty.
MISA-T boosts rollout throughput by over 53% while preserving workload integrity, revolutionizing how RL pipelines manage heterogeneous demands.
Prioritizing KV cache entries by importance can achieve over 93.7% accuracy during edge LLM handovers, significantly optimizing bandwidth usage.
Energy and latency can diverge by 3x under high computational demand, necessitating platform-specific models for accurate CNN inference cost predictions.
The optimal vocabulary size for LLMs can shift dramatically based on serving conditions, with potential divergences from training norms by up to 16x.
ReRound outperforms standard quantization techniques by resolving midpoint ambiguity, achieving superior accuracy in low-bit weight quantization for small LLMs.
A single-vector visual document retrieval system that is 15.6 times smaller and an order of magnitude faster than existing multi-vector models while retaining high accuracy.
Sorting prompts by their reliability can dramatically enhance the effectiveness of on-policy distillation, leading to superior performance in complex tasks.
Compressing tabular models by 85% without sacrificing performance could revolutionize how we deploy foundation models in resource-constrained environments.
Eliminating the lower bound on distillation loss could redefine the effectiveness of quantized models in label-scarce environments.
A novel I/O-aware reformulation of wavelet convolution slashes memory usage and accelerates training speed, making it a game-changer for deep learning efficiency.
Adversarial tenants can reconstruct private prompts with 100% success using timing attacks, but KVGov effectively neutralizes this threat while preserving cache efficiency.
Hand-written PTX kernels can outperform WMMA by up to 98.7x for INT4 operations, revealing critical insights into when low-level optimizations are worth the added complexity.
Achieving up to 99% of the theoretical maximum speed-up in looped language models could revolutionize inference efficiency in AI applications.
WDL-OPD boosts MATH500 accuracy from 0.630 to 0.685, showcasing a powerful new approach to stabilizing on-policy distillation.
LITEWAY slashes model size and energy consumption for wearable human activity recognition without sacrificing performance, achieving up to 9.52x size reduction.
Teacher-student mismatch can lead to flawed outputs, but TIDE's innovative correction method boosts reasoning accuracy by over 200% in challenging scenarios.
Stacking language models into a single nested architecture can cut training costs by 36% while maintaining competitive performance.
DUET achieves a groundbreaking balance in video generation, delivering high-quality visuals while preserving twice the diversity of traditional methods.
ICBQ not only cuts perplexity in quantized models but also salvages performance where traditional methods falter, redefining efficiency in model compression.
A staggering 63.2% of compressed runs show low or partial coverage, revealing critical gaps in KV-cache compression effectiveness.
Visual blind spots can be transformed into powerful self-supervision signals, leading to substantial performance gains in multimodal language models.
REST achieves few-step image generation that rivals traditional 40-step methods while slashing training costs by over 75%.
UnionSparse achieves up to 3.46x faster low-bit sparse LLM inference on edge GPUs by optimizing metadata handling, challenging the status quo in model efficiency.
Achieving SLOs in edge-cloud caching can be done with significantly lower capacity and cost through a novel hybrid segmented policy that adapts to workload patterns.
Achieving nearly 5x speedup on RISC-V processors, RVANNS revolutionizes approximate nearest neighbor search by optimizing both vector representation and graph traversal.
Achieving over 20% performance improvement in data-free knowledge distillation, UniDFKD eliminates reliance on architecture-specific priors, paving the way for more robust model training across diverse architectures.
A-PACK reveals that deferring audio pruning can lead to a 78% reduction in prefill costs while boosting performance in omni-modal LLMs.
Tying scales in PTQTP leads to a uniform nine-level quantizer that matches official serving performance while reducing file size and improving decoding speed.
Fine-tuning LLMs with a carbon-aware objective can yield task accuracy improvements while minimizing carbon emissions, but the effectiveness varies by task structure.
Speculative decoding can be up to 8.49 times faster than autoregressive methods with LibraSpec's novel marginal-gain-driven optimization.
LLMVisor achieves up to 4.4x improvement in latency attribution accuracy for multi-tenant LLMs, revealing hidden inefficiencies in GPU resource usage.
Injecting world-awareness into speculative decoding leads to a 1.5x speedup in task success rates while reducing near-contact failures by nearly 19%.
Reducing dispatch count is the key to unlocking efficient LLM inference in WebGPU, not improving kernel quality.
Calibration errors in LLM compression can be mitigated, leading to significant accuracy gains without the need for retraining.
FlashBoot slashes weight loading times from 20.1 seconds to just 0.4 seconds, revolutionizing how large models are deployed at scale.
DistillCache retains over 94% accuracy on long-context tasks while slashing memory usage by 75%, outperforming traditional methods and setting a new standard for efficient LLM inference.
BAMU achieves a remarkable MOS improvement in speech quality by dynamically allocating quantization resources based on frame complexity, outperforming traditional fixed-depth codecs.
Achieving over 90% performance retention with a staggering 20x KV cache compression could redefine efficiency in long-context audio inference.
OasisKV achieves up to 2.1x throughput gains in LLM inference while using significantly less memory, challenging the limits of current HBM constraints.
CoinRAG redefines efficiency in RAG by achieving higher accuracy with lower operational costs through innovative cache reuse strategies.
AutoPrune enables LLMs to autonomously design visual-token pruning strategies, achieving a remarkable 9.9x reduction in FLOPs without sacrificing performance.
WorldTrace redefines memory management in video world models, achieving up to 19.5% better episodic recall without any retraining.
Small LLMs can achieve up to 27.2% accuracy gains by leveraging hierarchical memory from larger teacher agents, reshaping how we think about agent training.