Search papers, labs, and topics across Lattice.
98 papers published across 4 labs.
OracleZoom is presented, an on-policy distillation-inspired, reference-constrained framework that trains on its trajectory while carrying the last ground-truth evidence beyond the supervision boundary, and achieves the state-of-the-art SR quality across zooming scales.
This work shows that specialist optimization implicitly selects from this latent trajectory space, and establishes a new view of specialist training: when gold reasoning is absent, tuning choices directly control the latent supervision passed to downstream models.
This study introduces StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework.
Native unified modelling is position as a promising path towards systems that perceive, reason and create within a fully end-to-end framework through SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture.
A branch-and-bound algorithm tailored to the resulting robust sparse portfolio problems is developed, together with a new pruning rule that can discard exponentially many candidate portfolios in a single step.
This work shows that specialist optimization implicitly selects from this latent trajectory space, and establishes a new view of specialist training: when gold reasoning is absent, tuning choices directly control the latent supervision passed to downstream models.
This study introduces StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types within a unified framework.
Native unified modelling is position as a promising path towards systems that perceive, reason and create within a fully end-to-end framework through SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture.
A branch-and-bound algorithm tailored to the resulting robust sparse portfolio problems is developed, together with a new pruning rule that can discard exponentially many candidate portfolios in a single step.
It is found that cache performance depends on transfer granularity, intermediate memory use, and when transfers enter the request schedule, not only on device bandwidth, not only on device bandwidth.
Comparing full-precision and quantized forward passes, and identifying two mechanisms that characterize pretrained quantization robustness, these mechanisms explain why quantization error that passes through numerous layers can still produce only small output changes, and are verified across models and quantization settings.
This work revisits RAG compression from a data-mining perspective by aggregating historical query--document--model interactions into reusable evidence views and proposes Reusable Evidence View Aggregation (REVA), a framework that mines the target generator's historical attention traces into a document-keyed, budget-agnostic score store.
A phase-decoupled, model-calibrated controller that hypothesizes that the optimal power setting is a property of the deployed (model, quantization, engine, hardware) combination rather than of the GPU class, that each lane warrants its own profile, and that converting SLO headroom into energy safely requires latency-gated calibration under a runtime SLO guard rather than a fixed recipe.
Extending LILA to dynamically allocate sparsity budgets via KS-scores yields state-of-the-art generative preservation at moderate compression, while uncovering fundamental single-layer architectural bottlenecks at higher compression regimes.
Evaluation on a representative multimodal benchmark demonstrates that EMMI can reduce the communication payload by 32x while maintaining comparable downstream accuracy, resulting in up to a 3.4x reduction in estimated end-to-end inference latency under bandwidth-constrained edge conditions.
This paper proposes OmniKVQuant, a training-free framework that enables 2-bit KV caches while substantially preserving performance across seven audio-visual benchmarks and provides a fused Triton decode kernel that unpacks the 2-bit cache during attention, so no dense FP16 cache is ever built.
It is shown that modality-unified prototypical contrast facilitates better modality invariance by jointly and simultaneously optimizing similarity relation within and across-modality.
Experiments on 4DMatch and 4DLoMatch show that both variants produce more accurate correspondences than the compared methods and improve downstream registration, with larger gains in low-overlap cases.
Negative Self-Distillation (NSD) is introduced, a new framework that optimizes LLMs by diverging from flawed reasoning rather than imitating privileged solutions, and consistently outperforms OPSD and other label-free, self-bootstrapping reinforcement learning (RL) baselines.
Experiments show that ZipCodec substantially outperforms existing streaming codecs at comparable bitrates in both reconstruction and downstream tasks, while operating at a significantly lower frame rate.
A four-coefficient parameterization lambda_t = sigma(a * h_t + b * u(x) + c + d * gap_t) in which direction-aligned proxies of EOPD and ToDi appear as one-dimensional (1D) restrictions, and which adds multi-channel composition and an explicit bias as further degrees of freedom.
HerALD (High-fidelity Exemplar Retrieval with Adaptive Landmark Distillation with Adaptive Landmark Distillation), a gradient-free graph condensation framework that adapts the node scoring and feature selection in the condensation pipeline to the graph's measured heterophily, is proposed.
The proposed greedy algorithm with alternating updates guarantees the four-peak distribution required for stable 2-bit clustering of each factor and admits closed-form initialization of cluster centers, removing the multi-restart k-means bottleneck of prior work.
FlexComp is proposed, a method-agnostic framework that decouples the ratio from both training and deployment and matches separately trained fixed-ratio specialists with minimal degradation across ICAE, 500xCompressor, and SAC on MRQA.
SpecGuard is introduced, an inference-time backdoor detector that repurposes speculative decoding at zero added model-computation cost and doubles as a free, always-on signal for detecting backdoored LLM behavior.
Experimental results demonstrate that the KuaiRP models not only match the current state-of-the-art proprietary models in role-playing fidelity within the authors' target domains, but also successfully recover general agent capabilities, maintaining extremely low deployment costs.
Uncertainty DMD is proposed, a simple uncertainty-injection framework that restores stochasticity at two key stages of AR generation: a timestep perturbation for the first chunk to increase first-chunk diversity, and a stochastic cache-writing mechanism for later chunks to preserve uncertainty in autoregressive conditioning.
PATTON, a PIM runtime that integrates production LLM serving engines with commodity PIM, achieves an average 1.95x speedup and 4.83x higher energy efficiency over evaluated baselines, requires no PIM processing-unit modifications, and maintains a KV cache hit rate comparable to the native GPU KV cache in vLLM.
X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning, is introduced.
Qualitative generalization across trajectory, scene, object, embodiment, and viewpoint shifts highlights Pelican-Sim's potential as a general-purpose world model simulator.
This survey covers inference-efficiency mechanisms for visual and audiovisual VideoLLMs that report concrete reductions in parameter count, FLOPs per input, latency, memory, or visual and audio token count, and organize methods by the pipeline stage at which they act.
Progressive semantic refinement no longer requires linear codebook memory growth: reusing a compact set of shared knowledge bases across residual stages achieves multi-depth transmission with a strictly bounded storage footprint.
Sub-7ms edge inference on an embedded FPGA is sufficient to diagnose 21 distinct electrical fault modes in high-frequency avionics grids with under 180,000 parameters.
Most LLM unlearning dissolves the moment a model is compressed for production, but isolating updates to high-significance layers prevents forgotten data from resurfacing under 4-bit quantization.
Winning lottery tickets at up to 95% sparsity can be harvested essentially for free simply by piggybacking iterative magnitude pruning onto the natural retraining loop of active learning.
Supervised compression can shrink 768-dimensional language embeddings down to just 5 dimensions without accuracy loss, allowing 5-qubit variational circuits to match full-scale 384-dimensional classical baselines.
Multi-source Bayesian network fusion usually destroys inference tractability, but imposing hard treewidth constraints via evolutionary search preserves consensus graph structure without sacrificing computational feasibility.
Generative crystal structure prediction no longer requires thousands of sequential sampling steps: learning average probability flows yields state-of-the-art structural and space-group fidelity in just 1 to 5 evaluations, generating 10,000 candidates in under a minute.
Prefill latency can be slashed by 80% without sacrificing long-context retrieval accuracy by pairing concatenation-aware fine-tuning with selective KV recomputation.
Softmax doesn't require explicit exponentials: directly mapping shifted scores to block-scaled E2M1 codes slashes vector-stage latency by 40% while exceeding standard MXFP4 accuracy on frontier LLMs and VLMs.
When valid solutions form disconnected manifolds, standard conditional means fail 100% of the time; deterministic equilibrium refinement recovers near-perfect validity without the trajectory jitter of stochastic diffusion.
Shrinking a 102M-parameter nnU-Net by 81× costs barely 2.5% in segmentation Dice while actually boosting lesion-level detection F1 by over 5 points.
Benchmark-average rankings mask massive sample-level disagreement among vision pruning techniques: dynamically routing inputs across existing pruning methods yields a 26.9% relative accuracy jump over any single fixed strategy.
Eliminating the language model output projection matrix in favor of geodesic decoding on a Poincaré ball cuts perplexity by over 50% compared to tied Transformers and SSMs at sub-million parameter scale.
Gating multi-teacher distillation on candidate correctness induces catastrophic label collapse and zero minority-class recall while failing to outperform simple hard filtering or improve factual grounding.
Unrepaired KV caches can perform worse than no cache at all, showing that zero-overhead position shifts break down the moment a query demands cross-attention between multiple sources.
Standard Gaussian teacher assumptions hide severe optimization traps: hidden node alignment alone dictates whether gradient descent fails into boundary versus interior local minima.
Treating distillation targets as dynamic on-policy decisions rather than static teacher outputs prevents catastrophic error propagation when adapting compact vision-language models to out-of-distribution multimodal data.
Stopping reasoning models early via intermediate answer consensus routinely truncates self-correction, because repeated agreement on a partial trajectory reflects prompt persistence rather than completed reasoning.
Event-driven neuromorphic execution slashes language model decode energy to 0.044 Joules per token by directly converting activation sparsity into skipped memory traffic rather than just idle compute cycles.
Hyper-parallel decoding on a distilled 4B LLM matches frontier model information extraction accuracy while operating at an 8% compute footprint.
LLMs can parse complex clinical and security logs using just 1% to 2% of the original context length without sacrificing predictive performance or explanation faithfulness.
Directly conditioning video diffusion on audio wastes massive capacity on static background and identity pixels—routing control transitively through a causal motion latent distilled under a single frozen video teacher achieves real-time streaming at 15.4 FPS with zero fidelity loss.
A training-free framework that combines a query-independent layout-derived interaction prior with instruction relevance and feature novelty to rank visual evidence according to both potential future utility and diversity, allowing retained visual evidence to shrink monotonically across budgets while remaining reusable throughout the trajectory.
Standard Gumbel-Sigmoid pruning fails in 3DGS because it forces binary decisions before importance ranks can stabilize—swapping it for a simple linear activation cuts primitive count by up to 3.6x while actually increasing rendering quality.
Interactive communication yields zero minimax advantage in 1-bit distributed mean estimation: purely non-adaptive queries match the adaptive rate for any heavy-tailed moment condition $k > 1$.
The foundational $+1$ SNR Pythagorean identity of rotation-based quantization holds strictly per-realization, but counterintuitively breaks down under rotation-averaging at finite dimensions before recovering asymptotically as coordinates Gaussianize.
On-policy LLM distillation does not actually need precise advantage magnitudes: retaining merely the directional sign of token advantages matches standard distillation performance while Total Variation smoothing eliminates late-stage training instability.
Scalar rewards in on-policy distillation tell tokens to gain or lose probability without specifying where that mass should actually go—framing distillation as explicit pairwise probability transport eliminates this background leakage and beats reverse-KL across reasoning tasks.
SequenceO1 achieves efficient ultra-long sequence modeling by compressing user histories while maintaining high performance, revolutionizing recommendation systems at scale.
Shared KV caches inherently leak private context when gated by semantic relevance alone, but binding biometric probes directly to cache blocks caps unauthorized memory access below 2% without sacrificing sub-30-token prefill speeds.
Dynamic layer-skipping fails on complex reasoning tasks because standard routers suffer from depth amnesia, but tracking execution history across layers lets Llama 3.1-8B bypass over a quarter of its parameters with zero performance loss.
LLM embedding models waste over 40% of their inference compute processing prefix tokens that become completely redundant at deeper layers, enabling aggressive, training-free token pruning with virtually zero retrieval degradation.
Heuristic KV cache pruning routinely fails under heavy compression because it ignores softmax curvature, but treating attention as a nonlinear Gaussian channel reveals exactly which tokens preserve information capacity.
ReMoMask-2, a structure-aware RAG framework coupling Hierarchical Bidirectional Momentum contrastive learning to align global and part-level features with text; Semantic Spatial-Temporal Attention (SSTA) for topology-aware fusion; and Topology Structured Masking (TSM) to force robust part-level grounding via adaptive masking are introduced.
Pretrained 3D Gaussian scenes can be compressed 100-fold post-hoc without any per-scene optimization, beating existing simplification baselines by 1.3 dB PSNR at 12x the speed.
Neural image compression can finally rival classic formats in raw throughput, achieving 2000 FPS decoding and practical 20 FPS encoding while matching JPEG's rate-distortion performance.
This work presents TontaubeV1, a model that preserves natural prosody while enabling streaming from a single consumer GPU, and is designed primarily for English and German, with additional multilingual support.
Spatio-temporal masking during student self-rollouts breaks the reverse-KL mode-collapse trap in autoregressive video diffusion, eliminating visual over-smoothing without relying on real video data.
Throwing away learned attention in favor of pure 3D spatial coverage allows VLMs to retain over 93% of their multi-view 3D reasoning performance while slashing visual token budgets by ~92%.
Speculative decoding drafters no longer need to be trained from scratch per model: target-agnostic pretraining on pruned small LMs produces a single, reusable backbone that outperforms bespoke drafters by up to 22.7% across completely different target architectures.
Contrary to widespread assumptions, RoPE is not responsible for attention sinks and outlier activations; the true culprits are causal-mask self-concentration and value-non-mixing at the sequence start.
Selective homomorphic encryption yields 10x inference speedups without model retraining, but early global mixing in modern vision backbones completely obliterates the savings.
FP8 Ozaki II emulates FP64 matrix multiplication by tensor-core products over a CRT residue system; converting the operands into residue planes costs integer-pipe and memory resources before tensor instructions issue.
A controlled measurement study of self-hosted LLM inference across edge and near-edge deployment nodes: an NVIDIA Jetson AGX Orin and a near-edge server with CPU-only and GPU-enabled inference modes, highlighting that compute-side inference metrics alone can lead to suboptimal placement for latency-sensitive interactive web services.
FlexSpIM, a digital CIM architecture supporting arbitrary operand resolution and shape within a unified storage for weights and neuron states, is introduced, enabling a layer-level hybrid weight- and output-stationary dataflow, maximizing operand reuse and reducing costly on- and off-chip data movement during SNN execution.
PENDA (processing element via norm-of-difference architecture), which leverages the law of cosines to recast multiplications as squared-difference operations to preserve exactness while optimizing hardware.
HDA-MoE is presented, a framework that optimizes MoE execution on NMP architectures through hybrid parallel deployment and runtime scheduling and integrates an offline hybrid parallel mapping algorithm with an online dynamic and adaptive scheduling mechanism to reduce communication overhead while improving computation utilization.
DiffLUT-Net is presented, an FPGA-native network connected by six-input LUTs that are trained from scratch, demonstrating the effectiveness of jointly learning LUT functions and sparse connectivity for compact FPGA-native inference.
ActionSplice is introduced, an inference framework that formulates this problem as Counterfactual State Transport (CST), a lightweight corrector that transports the interrupted backbone-native representation toward the matched state induced by the revised action at the same solver step.
This work introduces Grouped Value Attention (GVA), which stores grouped values and reconstructs content keys with a learned linear map at inference, which can be absorbed into the query, eliminating the need to materialize content keys in the intended decode path.
Chain-of-thought context can be halved without sacrificing accuracy by using the geometric trajectory of hidden states to selectively compress exploratory reasoning steps into continuous latent tokens.
Jointly pruning and quantizing weights under a single Bayesian objective breaks the traditional trade-offs of sequential compression pipelines, squeezing modern LLMs like Llama 3.2 and Qwen 2.5 further without catastrophic accuracy degradation.
Co-training speculative draft models directly inside 122B, 256K-token distributed RL runs removes the massive rollout bottleneck without causing pipeline stalls or context-parallel memory blowups.
Fixed-size chunked prefill forces an unnecessary compromise between decode latency and launch overhead: dynamically sizing chunks to fit active decode deadlines boosts serving goodput by up to 3.3× under tight latency SLOs.
Escalating to a larger LLM is counterproductive when it corrupts correct answers, meaning optimal cascading requires routing on net rescue-versus-harm rather than model uncertainty.
OracleZoom is presented, an on-policy distillation-inspired, reference-constrained framework that trains on its trajectory while carrying the last ground-truth evidence beyond the supervision boundary, and achieves the state-of-the-art SR quality across zooming scales.
Foundation models are practically useless for lossless time-series compression, yet shifting to error-bounded lossy regimes lets them outperform classical predictors across the board by collapsing in-band residual costs to zero.