Search papers, labs, and topics across Lattice.
We track OpenAI, DeepMind, Anthropic, and 17 other labs daily - with AI-powered summaries, trend charts, and a weekly digest.
We read everything so you don't have to. One email, zero noise.
We read everything so you don't have to. One email, zero noise.
Scaling generalized zero-shot recognition to massive vocabularies does not require re-architecting backbones: treating seen-versus-unseen gating as a downstream Monte Carlo calibration problem boosts unseen accuracy by over 20% in just 15 dimensions.
Benchmark-average rankings mask massive sample-level disagreement among vision pruning techniques: dynamically routing inputs across existing pruning methods yields a 26.9% relative accuracy jump over any single fixed strategy.
Agent memory breaks when models are forced to simultaneously interpret, update, and rewrite facts—decoupling maintenance into explicit pairwise relation classification and role-based fusion yields up to a 29.8 percentage point accuracy leap.
Industry-standard GPU memory sanitizers harbor major blind spots across core memory spaces, finally quantifiable through 149 targeted CUDA stress tests.
Interactive communication yields zero minimax advantage in 1-bit distributed mean estimation: purely non-adaptive queries match the adaptive rate for any heavy-tailed moment condition $k > 1$.
Deterministic weather models can match dedicated probabilistic ensembles within 0.04–0.13 CRPSS at a 10-day lead time simply by perturbing raw network weights at inference without any retraining.
Agent memory can be compressed by 50% with virtually zero performance degradation (retaining up to 99.7% accuracy) and a 2x retrieval speedup by structuring historical context into event-centric maximum spanning trees.
We read everything so you don't have to. One email, zero noise.
While long-tail knowledge is notoriously hard for LLMs to retain, structurally popular facts suffer the worst collateral damage during knowledge updates and act as super-spreaders of downstream hallucinations.
Neural image compression can finally rival classic formats in raw throughput, achieving 2000 FPS decoding and practical 20 FPS encoding while matching JPEG's rate-distortion performance.
Standard motion benchmarks fail inside the cabin because drivers mostly sit still—until right before a maneuver, when arm movement spikes 3.4x.
Video-level labels alone can reliably drive clinical video segmentation: coarse 3D-CAM cues paired with MedSAM2 achieve 94.5% recall and double small-polyp localization without a single manual frame annotation.
Turning published scientific figures back into executable Python remains a brittle frontier for top VLMs, with even frontier models routinely failing on multi-component layouts, precise axis semantics, and domain typography.
Upper-bounding the number of distribution shifts in sequential data is provably impossible without parametric assumptions, but conformal inference can uniquely deliver valid, non-trivial lower bounds.
Safety monitors miss critical risks like sandbagging and data leaks not because LLMs lack capability, but because hyperproperty detection fundamentally demands an executed second trace and an explicit comparative procedure.
Moving beyond passive prompt-and-generate workflows to genuine human-AI collaboration requires solving open technical challenges in mutual agency, real-time latent steering, and dynamic evaluation.
Multimodal agents do not need external verifier modules to handle noisy retrieval—targeted RL rewards can teach a 7B model to internally audit search results and approach proprietary-tier multi-hop reasoning on just 5,000 training samples.
MultiHuSE is introduced, a multi-modal dataset comprising 2,407 high-definition videos of 50 demographically diverse actors performing 1,463 text samples across four psychological humour styles (affiliative, aggressive, self-enhancing, and self-deprecating), as well as neutral content that provides empirical support for psychological theories linking humour and emotion.
Looped flows are proposed, an approach that allows solving harder problems by spending more computation through a finer temporal grid and enables multiple valid predictions from different initial noise samples.
Standard differential privacy reductions drastically underestimate how robust ensemble methods actually are, missing an exact connection between algorithmic stability and the spectral norm of the ensembling covariance operator.
Extreme label skew degrades federated multimodal performance nearly three times more than a sevenfold increase in participating clients, exposing data heterogeneity rather than network scale as the primary failure mode for decentralized clinical AI.
Entangling $t$ quantum copies during measurement improves state tomography efficiency by at most a factor of $\sqrt{t}$, showing that reaching optimal collective sample rates strictly requires entangling $\Omega(r^2)$ samples regardless of classical adaptivity.
Sub-7ms edge inference on an embedded FPGA is sufficient to diagnose 21 distinct electrical fault modes in high-frequency avionics grids with under 180,000 parameters.
Single-model 3D Gaussian Splatting can match the calibrated uncertainty bounds of a 10-model ensemble at 280 FPS simply by factorizing spatial rendering geometry from view-dependent difficulty under a conformal framework.
We read everything so you don't have to. One email, zero noise.
Autonomous vehicles can maintain millimeter-wave beam alignment even when critical sensors drop out by leveraging generative cross-modal reconstruction across camera, LiDAR, and radar feeds.
Even "leak-free" protein interaction benchmarks are riddled with non-biological shortcuts that models readily exploit—and standard negative-sampling heuristics often make the bias worse.
LLM research agents consistently lose to a naive moving-average baseline at forecasting scientific trends—not due to poor reasoning, but because asking models to predict the future systematically biases their search policies away from the most recent literature.
Single-device benchmarks mask a critical failure mode: current GUI agents break down rapidly as soon as tasks require maintaining state and transferring intermediate results across operating systems.
Treating distillation targets as dynamic on-policy decisions rather than static teacher outputs prevents catastrophic error propagation when adapting compact vision-language models to out-of-distribution multimodal data.
Fluent surface text masks deep grammatical fragility: across 600K test cases, leading LLMs consistently break down when forced to execute fine-grained Arabic morphosyntactic control, particularly under cliticization and rare inflections.
We read everything so you don't have to. One email, zero noise.
Directly conditioning video diffusion on audio wastes massive capacity on static background and identity pixels—routing control transitively through a causal motion latent distilled under a single frozen video teacher achieves real-time streaming at 15.4 FPS with zero fidelity loss.
High-resolution, multi-view consistent 3D texturing no longer requires costly per-scene optimization or custom fine-tuning: frozen 2D diffusion models can directly synthesize production-ready texture atlases with baked shadows at an 80% speedup.
Standard pixel augmentations frequently corrupt delicate vision-language alignment, but injecting diffusion-style isotropic noise directly into embedding spaces breaks through the longstanding performance ceiling of stacked CutMix, Mixup, and RandAug recipes.
Diffusion-generated thermal faces can effectively break the multi-modal data bottleneck in biometrics, outperforming models trained on scarce real pairs without requiring costly image translation at inference time.
Pruning GUI agent screenshots typically cripples long-horizon execution because discarded visual context cannot be recovered from the KV-cache, but framing token retention as a nested, coverage-aware admission problem allows massive visual compression without sacrificing future task utility.
Standard Gumbel-Sigmoid pruning fails in 3DGS because it forces binary decisions before importance ranks can stabilize—swapping it for a simple linear activation cuts primitive count by up to 3.6x while actually increasing rendering quality.
We read everything so you don't have to. One email, zero noise.
We read everything so you don't have to. One email, zero noise.