Search papers, labs, and topics across Lattice.
22 papers published across 2 labs.
Cypress-based end-to-end testing can achieve high reliability and maintainability, even as web applications evolve.
ARG scaffolding boosts GPT-5's success rate from 9.4% to 49.0% in real-world ML tasks, while Modular setups suffer from high specification gaming.
Axon achieves up to 107% speedup on JAX, revolutionizing how LLMs can be efficiently deployed across different frameworks without sacrificing optimization.
Replication studies in usable security and privacy often modify original research significantly, raising questions about their validity and the standards of the field.
LLMs can autonomously generate prompts that rival expert-written ones, but still fall short in accurately interpreting scientific context and discovering literature.
Cypress-based end-to-end testing can achieve high reliability and maintainability, even as web applications evolve.
ARG scaffolding boosts GPT-5's success rate from 9.4% to 49.0% in real-world ML tasks, while Modular setups suffer from high specification gaming.
Axon achieves up to 107% speedup on JAX, revolutionizing how LLMs can be efficiently deployed across different frameworks without sacrificing optimization.
Replication studies in usable security and privacy often modify original research significantly, raising questions about their validity and the standards of the field.
LLMs can autonomously generate prompts that rival expert-written ones, but still fall short in accurately interpreting scientific context and discovering literature.
Fine-tuned small language models can outperform massive counterparts by leveraging high-quality filtered data, challenging the notion that bigger is always better in machine translation.
Local AI inference may democratize access, but it also shifts power dynamics, placing control in the hands of hardware vendors and core maintainers.
Local LLM-based SSH honeypots can achieve superior shell emulation accuracy with the right prompting and fine-tuning strategies, but their effects can conflict in unexpected ways.
Flama unifies API development and LLM services in a single, type-driven framework that streamlines production workflows and enhances developer efficiency.
Achieving an 80% reduction in delay estimation error while being over 300 times smaller and twice as fast than existing methods could revolutionize pre-route analysis in IC design.
ArguLens achieves 82.6% accuracy in essay classification while providing interpretable feedback, challenging the status quo of opaque scoring systems.
An adaptive ensemble-size rule can cut memory usage by 37 MB and fit time by 0.4 seconds while maintaining classification accuracy in time series analysis.
Human expertise remains essential in the agentic coding process, even as productivity in simulation library development accelerates significantly.
Verifiability of software artifacts in decentralized ecosystems is severely limited by metadata gaps, revealing a critical need for systemic improvements to enhance trust in distributed builds.
Only 4 out of 35 configurations achieved meaningful performance in extracting structured information from high-risk documents, underscoring the challenges of using open-source models in critical applications.
TabPFN-Rel not only tops the leaderboard on RelArena-$\alpha$ but also challenges the notion that specialized architectures are always superior to flattened relational databases.
The best commercially-licensed retrieval system still trails behind its free counterpart, revealing a hidden "commercial tax" that could impact enterprise decisions.
Over 80% of research software projects have conflicting metadata across their self-descriptions, risking fragmented credit and provenance in the scientific community.
Trustworthiness in open-source dependencies can now be quantified with a single score that reveals critical insights into software supply chain security.
Extreme fidelity loss reveals critical vulnerabilities in long-horizon tasks that standard accuracy metrics overlook.
UI-Mate-27B not only sets a new open-weight benchmark in GUI automation but also doubles the success rate of long-horizon tasks with just one demonstration.
A lineage verification method that distinguishes model ancestry with perfect accuracy, even under aggressive checkpoint modifications.