Search papers, labs, and topics across Lattice.
Open-weight model releases, reproducibility, model licensing, and community-driven AI development.
#24 of 24
2
The spectral residual is characterized as a useful but domain-sensitive inductive bias for structural connectivity screening, andarse scaling extends to 20,000 nodes and separates one-time spectral setup from amortized screening cost.
This work compares country-continent questions with noun, adjective, and code answers while keeping several fitted measurements distinct across Qwen, Llama, and Gemma to separate early readability, natural strength, causal steering, and later content dependence.
Whether medical research is keeping pace with the systems it evaluates is asked, and Rigour and currency are in tension, and that tension reflects model selection rather than research timelines.
This work introduces AgentActionBench, a process-oriented benchmark for evaluating agent-based experiment reproduction across ML and AI4Science domains, and uses an MCP-based Action Recorder to capture agents' behaviour throughout the reproduction process and evaluates the resulting traces with paper-specific rubrics.
This work derives the first model of WhatsApp Web's implementation of the Signal protocol and the most detailed model to date of Signal's original protocol, and reveals previously undocumented differences between the original libsignal library and WhatsApp's fork.
It is argued that achieving reproducibility in TEEs requires a holistic development approach that extends beyond individual developers and calls for stronger commitments - rather than treating TEEs as a"security badge".
This work systematically reviews FSL approaches for NIDS published from 2022 to 2026 with PRISMA 2020-like reporting to search ACM Digital Library, IEEE Xplore, and Scopus, and compares reported performance.
A descriptor of GEMM reduction order is introduced, and the first black-box reconstruction of a closed-source library's arithmetic for bit-level correctness is performed, including the first black-box reconstruction of a closed-source library's arithmetic for bit-level correctness.
An information-weighted cross-entropy loss that rescales token-level contributions using TF-IDF statistics, emphasizing semantically informative tokens while down-weighting ubiquitous ones is presented, offering a lightweight and principled way to mitigate memorization without disrupting standard training dynamics.
Two consistent dissociations between representation-level alignment and behavioral expression are reported, plus a common failure under position perturbations, characterize representation-behavior dissociation in a high-signal setting rather than establishing universality across models or persona pairs.
It is argued that agents lower the cost of maintaining tests, commit histories, repository structure, instructions, and decision records while making their benefits immediate while making their benefits immediate.
Multilingual speech foundations leave massive performance on the table: specialized monolingual Whisper models beat Whisper-large-v3 across 77 of 102 languages while slashing character error rates by nearly 3x.
This work studies the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions.
Natural-language instructions can now handle both zero-shot speech synthesis and surgical acoustic editing within a single unified model, operating at 4-step distilled inference speeds without classifier-free guidance.
Domain-specialized RL experts can be unified without negative transfer by distilling their feedback directly onto student-generated trajectories, resolving the multi-task optimization bottleneck that limits single video foundation models.
A controlled measurement study of self-hosted LLM inference across edge and near-edge deployment nodes: an NVIDIA Jetson AGX Orin and a near-edge server with CPU-only and GPU-enabled inference modes, highlighting that compute-side inference metrics alone can lead to suboptimal placement for latency-sensitive interactive web services.
Video generative priors only translate into robust physical control when paired with explicit world-to-action information routing and synchronized joint denoising rather than standard monolithic fine-tuning.
Seizure detectors may perform well on clean data but can falter dramatically under real-world conditions, revealing a critical gap in pre-deployment assessments.
The first open Armenian LLM, arm-gemma-e4b, not only outperforms all predecessors but also highlights the critical balance between fluency and knowledge retention in low-resource language models.
WebXR could be the key to a more accessible and sustainable Metaverse, challenging the dominance of commercial game engines.
High-fidelity image synthesis does not require paired text from day one: pre-training visual priors on uncaptioned images before multimodal alignment beats conventional joint training pipelines to establish a new open-source DiT benchmark.
Governance metrics in agricultural open-source software reveal surprising independence from actual security risks, challenging assumptions about the sector's vulnerability.
Uncovering recurring implementation patterns in LLM codebases could revolutionize how developers approach building and optimizing AI applications.
Boundary-mutation testing uncovers critical vulnerabilities in secret scanners, revealing that some rules can fail completely under realistic context variations.
MutMem V2 achieves portable integrity and reproducibility in persistent agent memory, setting a new standard for cryptographic authorization in AI systems.
Weak governance in standards-setting organizations is pushing geospatial coordination towards proprietary platforms, threatening the very essence of open standards as public goods.
Traditional knowledge graph matchers fall short, with LLMs outperforming them by a significant margin in threat report evaluations.
A leading-order effective field links empirical dynamics to neural computation, revealing how human-AI interactions exhibit reproducible and interpretable patterns.
Disagreement among evaluators can be systematically bounded, revealing critical insights into the reliability of natural language task assessments.
Hand-written behavioral clock gating fails at gate level, while tool-inserted ICG cells achieve robust power savings across all simulation corners.
Achieving competitive performance in full-key side-channel attacks on uncropped datasets with a simple transformer model could revolutionize the field by making advanced techniques more accessible.
Feedback effectiveness in LLM code repair is not universally applicable across programming languages, challenging previous assumptions about its reliability.
Experimental records in autonomous driving can now be seamlessly integrated with real-time operational conditions, enhancing interpretability and reuse across teams.
Achieving 95.7% precision in parameter extraction, this open-weights model challenges the reliance on commercial systems for reproducible agentic workflows in materials science.
Generating synthetic populations with full reproducibility from public data sources could revolutionize urban studies and demographic modeling.
Only 8.7% of reproducibility-relevant mutations in ML repositories are detected by current validation workflows, revealing a critical oversight in safeguarding research integrity.
Automatic reproducibility for 32% of Maven Central packages reveals critical flaws in manual curation efforts and sets a new standard for software reliability.
Trust and traceability in AI-enabled scientific discovery are now recognized as critical pillars for future scientific computing ecosystems.
Partitioning algorithms with similar entanglement costs can incur drastically different execution penalties, revealing hidden trade-offs that could reshape DQC compiler design.
Achieving near state-of-the-art performance for under $7K opens the door for cost-effective language model training accessible to the broader research community.
SAMpLE transforms the integration of machine learning in virtual prototyping, enabling seamless evaluation of diverse models without cumbersome re-implementation.
Claim-locked reporting boosts the accuracy of LLM-generated statistical reports by over 37%, ensuring that evidence integrity is maintained throughout the writing process.
VietAIDetector achieves superior detection of AI-generated Vietnamese text without requiring any domain-specific training data, setting a new standard for language-specific AI content verification.
Only 6.5% of neuro-symbolic AI studies can be reproduced from their published artifacts, exposing a severe reproducibility crisis in the field.
The open-source satellite software ecosystem is not only growing in popularity but is also marked by a surprising diversity of goals and programming languages that could redefine development practices in the field.