Search papers, labs, and topics across Lattice.
24 papers published across 2 labs.
The results suggest that fine-tuning is not just about how much a model changes, but how that change is spent, and that changing the accessible directions can qualitatively alter the outcome of fine-tuning.
The spectral residual is characterized as a useful but domain-sensitive inductive bias for structural connectivity screening, andarse scaling extends to 20,000 nodes and separates one-time spectral setup from amortized screening cost.
This work compares country-continent questions with noun, adjective, and code answers while keeping several fitted measurements distinct across Qwen, Llama, and Gemma to separate early readability, natural strength, causal steering, and later content dependence.
Whether medical research is keeping pace with the systems it evaluates is asked, and Rigour and currency are in tension, and that tension reflects model selection rather than research timelines.
This work introduces AgentActionBench, a process-oriented benchmark for evaluating agent-based experiment reproduction across ML and AI4Science domains, and uses an MCP-based Action Recorder to capture agents' behaviour throughout the reproduction process and evaluates the resulting traces with paper-specific rubrics.
The results suggest that fine-tuning is not just about how much a model changes, but how that change is spent, and that changing the accessible directions can qualitatively alter the outcome of fine-tuning.
The spectral residual is characterized as a useful but domain-sensitive inductive bias for structural connectivity screening, andarse scaling extends to 20,000 nodes and separates one-time spectral setup from amortized screening cost.
This work compares country-continent questions with noun, adjective, and code answers while keeping several fitted measurements distinct across Qwen, Llama, and Gemma to separate early readability, natural strength, causal steering, and later content dependence.
Whether medical research is keeping pace with the systems it evaluates is asked, and Rigour and currency are in tension, and that tension reflects model selection rather than research timelines.
This work introduces AgentActionBench, a process-oriented benchmark for evaluating agent-based experiment reproduction across ML and AI4Science domains, and uses an MCP-based Action Recorder to capture agents' behaviour throughout the reproduction process and evaluates the resulting traces with paper-specific rubrics.
This work derives the first model of WhatsApp Web's implementation of the Signal protocol and the most detailed model to date of Signal's original protocol, and reveals previously undocumented differences between the original libsignal library and WhatsApp's fork.
It is argued that achieving reproducibility in TEEs requires a holistic development approach that extends beyond individual developers and calls for stronger commitments - rather than treating TEEs as a"security badge".
This work systematically reviews FSL approaches for NIDS published from 2022 to 2026 with PRISMA 2020-like reporting to search ACM Digital Library, IEEE Xplore, and Scopus, and compares reported performance.
A descriptor of GEMM reduction order is introduced, and the first black-box reconstruction of a closed-source library's arithmetic for bit-level correctness is performed, including the first black-box reconstruction of a closed-source library's arithmetic for bit-level correctness.
An information-weighted cross-entropy loss that rescales token-level contributions using TF-IDF statistics, emphasizing semantically informative tokens while down-weighting ubiquitous ones is presented, offering a lightweight and principled way to mitigate memorization without disrupting standard training dynamics.
Two consistent dissociations between representation-level alignment and behavioral expression are reported, plus a common failure under position perturbations, characterize representation-behavior dissociation in a high-signal setting rather than establishing universality across models or persona pairs.
It is argued that agents lower the cost of maintaining tests, commit histories, repository structure, instructions, and decision records while making their benefits immediate while making their benefits immediate.
Multilingual speech foundations leave massive performance on the table: specialized monolingual Whisper models beat Whisper-large-v3 across 77 of 102 languages while slashing character error rates by nearly 3x.
This work studies the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions.
Natural-language instructions can now handle both zero-shot speech synthesis and surgical acoustic editing within a single unified model, operating at 4-step distilled inference speeds without classifier-free guidance.
Domain-specialized RL experts can be unified without negative transfer by distilling their feedback directly onto student-generated trajectories, resolving the multi-task optimization bottleneck that limits single video foundation models.
A controlled measurement study of self-hosted LLM inference across edge and near-edge deployment nodes: an NVIDIA Jetson AGX Orin and a near-edge server with CPU-only and GPU-enabled inference modes, highlighting that compute-side inference metrics alone can lead to suboptimal placement for latency-sensitive interactive web services.
Video generative priors only translate into robust physical control when paired with explicit world-to-action information routing and synchronized joint denoising rather than standard monolithic fine-tuning.
DataFlex-RL, an evaluation platform for comparing choices under a common GRPO recipe, is introduced, finding that changing the data policy measurably changes the training process but does not produce a reproducible improvement over uniform training.