Search papers, labs, and topics across Lattice.
96 papers published across 4 labs.
On-policy distillation is massively data-overfed: a single training prompt recovers most full-dataset performance gains, while just 16 prompts saturate 98.9% of reachable state space to match full-data distillation.
Foundation models are practically useless for lossless time-series compression, yet shifting to error-bounded lossy regimes lets them outperform classical predictors across the board by collapsing in-band residual costs to zero.
Prematurely abandoned in modern scaling recipes, layer dropout can actually slash LLM pre-training compute by 25% while natively unlocking 1.5x faster inference through zero-shot layer skipping.
Models can autonomously bootstrap their own dense token-level supervision simply by extrapolating the trajectory of their own RL updates away from a trailing checkpoint.
Long-horizon reasoning traces repeatedly backtrack to early planning steps via predictable query clusters, revealing that a handful of representative "beacon" vectors can guide 5.8× KV cache compression with near-zero quality loss.
Foundation models are practically useless for lossless time-series compression, yet shifting to error-bounded lossy regimes lets them outperform classical predictors across the board by collapsing in-band residual costs to zero.
Prematurely abandoned in modern scaling recipes, layer dropout can actually slash LLM pre-training compute by 25% while natively unlocking 1.5x faster inference through zero-shot layer skipping.
Models can autonomously bootstrap their own dense token-level supervision simply by extrapolating the trajectory of their own RL updates away from a trailing checkpoint.
Long-horizon reasoning traces repeatedly backtrack to early planning steps via predictable query clusters, revealing that a handful of representative "beacon" vectors can guide 5.8× KV cache compression with near-zero quality loss.
Video diffusion models fail at few-step camera control not because of capacity limits, but because discretization errors bend the sampling trajectory—a failure mode resolved here to unlock 25× faster novel-view synthesis.
A two-stage OPD-then-RL approach outperforms traditional methods by leveraging the strengths of both on-policy distillation and reinforcement learning without the interference seen in joint optimization.
Free pause tokens boost language model performance without increasing context length or latency, achieving significant gains with minimal training overhead.
Query-independent eviction signals can maintain 100% retrieval accuracy in KV caches while compressing data without any measurable cost.
SBW achieves over 6000x faster watermarking than previous methods while maintaining robust detection guarantees and minimal overhead.
The intricate relationship between prime-power spectra and critical-line geometry reveals new insights into integer quantization and its implications for number theory.
Compressed models may achieve high accuracy but suffer a dramatic drop in adaptability during test-time, revealing a critical trade-off in model deployment.
ALRA achieves a notable accuracy improvement over existing distillation methods by intelligently combining student and teacher token selections, revealing the importance of local relational alignment in model training.
Instruction duplication boosts diagnostic accuracy in language models by nearly 3 percentage points, significantly reducing failure rates without the need for retraining.
Classical data types can be quantised in a way that preserves their structural properties, unlocking new insights for quantum programming.
Iapetus achieves a remarkable 91.6% task completion rate while slashing mean latency and battery draw by over 50% in satellite-based ViT inference.
On-policy distillation is massively data-overfed: a single training prompt recovers most full-dataset performance gains, while just 16 prompts saturate 98.9% of reachable state space to match full-data distillation.
Achieving over 2x throughput improvements in attention mechanisms could redefine efficiency benchmarks for large-scale model training.
Task-adaptive pruning can cut model size by nearly 25% while boosting accuracy in histopathology tasks, challenging the notion that bigger always means better.
By reallocating KV cache resources based on attention head specialization, SGD-KV cuts memory usage by 75% while handling contexts up to 1M tokens with state-of-the-art accuracy.
Achieving 2.8x compression in on-device speech tokenizers without sacrificing accuracy could revolutionize mobile AI applications.
PACodec slashes bitrate by 30% while preserving decoding quality, thanks to its innovative parallel quantization approach.
DSAQuant reveals that aligning quantization training with the stage-wise nature of video diffusion can drastically enhance visual fidelity in text-to-video generation.
Halving frame resolution costs virtually nothing in long-video MLLMs, but reinvesting those saved visual tokens to double temporal frame counts yields an immediate 2–3 point accuracy gain—all while a 30-year-old sparse approximation algorithm matches state-of-the-art custom selectors.
Recurrent linear attention does not compound quantization drift over long contexts—delta-rule updates and non-linear gates actively overwrite and compress noise, allowing an entire 27B hybrid model to run in NVFP4 W4A4 with near-zero quality loss up to 64K tokens.
Sub-threshold recurrent updates silently freeze low-precision hidden states—causing up to a 300x error spike—unless historical quantization residuals are actively fed forward across time.
Complex scoring algorithms for KV cache compression are largely unnecessary: protecting the initial prompt and dropping reasoning tokens completely at random matches state-of-the-art accuracy with up to 43% higher serving throughput.
Auxiliary draft models are no longer necessary for speculative decoding: distilling lightweight diffusion heads directly into standard LLMs yields lossless 3× generation speedups even at peak batch sizes.
Recurring frontier API calls can be replaced by one-minute compile-time distillation, turning plain-text prompts into local, versionable neural functions that achieve 83.6% accuracy on benchmarks where instantaneous compilers fail entirely.
Language models can now self-select relevant context, slashing attention costs by over 50% while maintaining performance.
Switching to E5M3 block scaling not only simplifies FP4 pretraining but also yields significantly improved training and validation performance.
Achieving 2.27x faster decoding for 2-bit LLM weights could redefine efficiency benchmarks in low-bit quantization methods.
Value projection layers are the Achilles' heel of billion-parameter models, revealing surprising sensitivity that could reshape compression strategies.
Minimizing distortion, not maximizing codebook utilization, is the key to achieving superior reconstruction fidelity in visual tokenization.
XMerge outperforms existing depth-compression techniques, achieving top rankings in task performance while enabling aggressive layer removal without sacrificing quality.
AVIS recovers nearly 70% of accuracy loss from quantization while ensuring real-time performance on lunar rovers, even under radiation constraints.
Trajectory geometry can be harnessed to dramatically enhance diffusion model sampling efficiency without the need for retraining, yielding significant improvements in output quality.
Model size is a poor predictor of answer quality in retrieval-augmented systems, challenging long-held assumptions in AI deployment.
SCX Router achieves a 1.5% performance boost over the best fixed model by intelligently selecting the most suitable LLM for each task in real-time.
Training on the entire deployed matrix can unlock up to 81.4% of a model's capacity that was previously unreachable, leading to unprecedented performance gains in low-rank distillation.
GaLe achieves exact-inference performance with a staggering 65% speedup and 90% RAM reduction, revolutionizing how we deploy models on embedded devices.
Teacher confidence can be the key to unlocking reliable predictions in autoregressive vision language models, leading to substantial performance gains across multiple benchmarks.
FastMTP speculative decoding doubles the decoding speed on low-budget GPUs while maintaining high accuracy in document parsing.
CoverPruner reveals that optimizing for representational coverage can significantly enhance the accuracy of pruned vision-language models, especially when faced with aggressive compression.
AceSpec transforms edge-cloud LLM inference by turning catastrophic network stalls into efficient local memory lookups, achieving a remarkable 3.52× speedup even in low-bandwidth environments.
RACE-AIMC cuts energy use by 69% while maintaining accuracy comparable to digital systems by intelligently selecting the best AIMC accelerator and certifying its reliability.
As conversation lengths increase, LLM systems can achieve a hit ratio that converges to a predictable limit, revolutionizing memory management strategies.
Achieving up to 65.5% parameter reduction and nearly 2x inference speedup for 3D point cloud models without needing source code is a game-changer for edge deployment.
Pruning LLMs can amplify biases, but Debias-SparseGPT shows that you can reduce these biases while maintaining performance, even under aggressive sparsity constraints.
MLLMs do not need full-depth forward passes on every incoming frame: indexing video streams solely within early transformer layers slashes prefill latency by 52x without degrading downstream comprehension.
Because mode-seeking reverse KL aggressively amplifies incorrect teacher signals, gating dense distillation on verifier-scored teacher probes systematically outperforms uniform distillation while reclaiming massive amounts of idle teacher compute.
Repurposing speculative decoding backbones for output-length prediction slashes short-request latency by nearly 35%.
Achieving significant network compression without sacrificing performance, LRNBA allows for the creation of deeper and wider models under the same parameter budget.
Global allocation of precision budget in LLMs can improve accuracy by 21-52 points compared to local layer-specific repairs, defying conventional wisdom about quantization damage.
FPGAs could revolutionize Transformer model deployment by offering superior energy efficiency and latency compared to traditional hardware.
High-norm outlier tokens, often preserved by existing methods, are actually redundant and can be pruned without sacrificing performance—SinkPruner proves it with an 89% token reduction while retaining over 96% of model effectiveness.
Hidden preferences can be stealthily transferred during model distillation, but targeted regularization can significantly curb this effect without sacrificing performance.
SFAD achieves a remarkable 2.48x speedup in inference while enhancing contextual faithfulness, addressing one of the most pressing challenges in large language models.
Retaining 79.3% of full performance with just 32 visual tokens reveals that spatial coverage trumps traditional importance metrics in multimodal models.
Structured pruning is the first stage where driving capabilities degrade, challenging the effectiveness of standard compression techniques in automated driving.
Achieving over 24 times better energy efficiency in LLM inference could redefine the scalability of AI applications.
RaDiCal reveals that relying solely on attention saliency can misguide token pruning, leading to a 39-45% reduction in FLOPs while enhancing ranking performance across key datasets.
Achieving a 98% reduction in AI traffic without sacrificing performance challenges conventional wisdom about the necessity of precise data representation in VLMs.
CRISP recovers up to 28% improvement on retrieval tasks while achieving a remarkable 5.30x speedup in attention computation for long-context LLMs.
Switch Distillation not only enhances reasoning performance by up to 71% but also preserves factual recall, challenging the conventional trade-off in knowledge distillation.
Cognitive load can increase energy costs per token by a significant margin, revealing the hidden costs of prompt design in mobile LLM applications.
Tailoring iteration counts for nonlinearity approximations can cut inference latency in encrypted language models by over 40%.
Photonic interconnects can deliver up to 5.8x latency improvements for large-scale AI inference, revolutionizing how we handle complex models.
mzCache slashes Time-to-First-Token by up to 5.5× in multitasking environments, revolutionizing on-device LLM responsiveness.
FORGE recovers 93% of accuracy gains from advanced adaptation methods while running efficiently on integer-only microcontrollers, revolutionizing deployment strategies for edge AI.
Achieving nearly 4x energy savings and over 4x latency improvements for LLMs on edge devices without sacrificing performance opens new avenues for deploying advanced AI in real-world applications.
MERGED achieves a 13.79% boost in PR-AUC while being 6x cheaper than larger models, revolutionizing entity resolution in dynamic business environments.
Language models process context more effectively through direct continuous embeddings than human-readable text, beating uncompressed context accuracy at 7.7× compression while slashing inference latency by up to 9×.
Structural and magnitude pruning can retain significant functional redundancy, revealing that model compression strategies may overlook critical parameter directions.
Cross-model KV sharing can boost accuracy and cut prefill costs dramatically, challenging the notion that KV states are strictly model-local.
DynaNDE accelerates MoE inference by leveraging dynamic scheduling and expert reuse, resulting in up to 2.6x throughput improvements.
Precomputed memory in language models can degrade significantly unless rebuilt frequently and updated with specifically phrased corrections.
Speculative decoding can be significantly accelerated by training models to anticipate verification outcomes, leading to faster and more efficient inference.
OPD's effectiveness hinges less on teacher supervision than previously thought, with a new method achieving a staggering 263% relative gain without any teacher input.
Efficient evaluation methods can drastically alter the conclusions drawn about model behavior, revealing hidden vulnerabilities in AI benchmarking.
Fine-tuning low-bit models can be both efficient and deployment-faithful, thanks to a new optimization technique that leverages code surrogate gradients.
Achieving 4x smaller compression budgets without sacrificing performance, TopoCompress redefines long-context handling in language models.
Q-Strata outperforms existing quantization methods by directly optimizing inter-block dependencies, leading to superior model performance with lower perplexity.
Sparse, quantized linear-attention models can outperform dense counterparts on neuromorphic hardware, achieving up to 37× higher throughput and 16× lower power consumption.
Acceptance histograms reveal hidden efficiency gains in speculative decoding, leading to a novel post-training method that boosts token generation significantly.
SkillZip Pro slashes skill bundle size by 38% without sacrificing performance, revolutionizing how self-evolving agents manage their execution paths.
NN-PPI enables small language models to rival the accuracy of larger models, making automated fact-checking both scalable and affordable.
Over 90% of faithfulness changes are negative under INT4 quantization, revealing a hidden risk in compressed retrieval-augmented generation systems.
Centering visual token features can enhance diversity detection but may paradoxically reduce the effectiveness of pruning—unless you leverage both centered and raw geometries together.
CHIPSMORE accelerates LLM inference by achieving 2.38x higher throughput and 27x better energy efficiency compared to Nvidia H100, without duplicating weights for concurrent requests.
RSLM achieves up to 4x compression in ANN systems without sacrificing recall, revolutionizing the efficiency of high-dimensional search tasks.
LaMoC achieves a surprising 2.5% reduction in perplexity while enhancing task accuracy, redefining the standards for modular compression in LLMs.
HBQ achieves state-of-the-art efficiency in LLM inference, delivering up to 2.3x higher area/energy efficiency while maintaining accuracy levels that challenge conventional methods.
DRLM achieves up to 51% faster inference and 67% reduced queuing delays in edge environments, all while maintaining accuracy.
CASTER achieves superior performance in frozen model adaptation while requiring 18x less state, challenging the conventional reliance on gradient-based methods.
JITterFlip reveals that targeting the JIT serving control plane can amplify attack effects by over 7 million times, exposing a major vulnerability in LLM inference systems.
IDA-OPD not only boosts diversity in distilled models but does so while cutting computational costs, challenging the notion that more data always leads to better performance.