Search papers, labs, and topics across Lattice.
100 papers published across 4 labs.
MathForm-8B not only outperforms specialized autoformalizers but also sets a new standard for verified mathematical formalization with an impressive 88.06% pass rate on syntax checks.
CAPRI reveals that a contract-aware approach can drastically improve the integrity of LLM-generated proofs, achieving up to 81% valid repairs without violating edit contracts.
Integration-first coverage strategies can reveal untested code paths in embedded systems, transforming how we assess testing completeness.
Current MLLMs achieve only a 75% compilation success rate in scientific figure editing, but targeted training can boost performance to over 83%.
Over 38,920 discrepancies in processor specifications were uncovered, revealing critical bugs that compromise security analysis tools.
MathForm-8B not only outperforms specialized autoformalizers but also sets a new standard for verified mathematical formalization with an impressive 88.06% pass rate on syntax checks.
CAPRI reveals that a contract-aware approach can drastically improve the integrity of LLM-generated proofs, achieving up to 81% valid repairs without violating edit contracts.
Integration-first coverage strategies can reveal untested code paths in embedded systems, transforming how we assess testing completeness.
Current MLLMs achieve only a 75% compilation success rate in scientific figure editing, but targeted training can boost performance to over 83%.
Over 38,920 discrepancies in processor specifications were uncovered, revealing critical bugs that compromise security analysis tools.
AI agents can only correctly synthesize implementations and proofs for less than two-thirds of tested multi-module software repositories, revealing significant gaps in current capabilities.
Build integration, not candidate generation, is the critical hurdle for reliable LLM-assisted dynamic analysis in autonomous vehicle software.
Reducing candidate repair options by half while maintaining performance challenges the notion that complex models always yield better results in agent harness repair.
LLM-guided graph generation can more than double optimization performance, achieving a 39.5% win rate over traditional methods.
SkillShapley reveals that not all steps in agent skills are created equal, enabling precise identification of high-impact actions that can enhance performance.
Legacy bioinformatics code can be transformed into efficient Rust implementations, slashing size by 80x and build time by 10x while boosting performance over threefold.
Analyzing 57.5K transactional prompts reveals a surprising Zipf-like distribution in usage patterns, challenging assumptions about prompt effectiveness across different contexts.
Iterative LLM repairs in IaC can introduce security regressions, but a careful analysis reveals that only 3.3% of scenarios genuinely degrade security.
Smart contract invariants could have prevented all attacks in a benchmark of real-world Ethereum exploits, showcasing a powerful new defense against blockchain vulnerabilities.
Multi-driver fuzzing can increase coverage by nearly 28% and expose unique bugs that single-driver approaches miss, revealing the structural intricacies of software exploration.
LLM-generated patches are not only larger but also more complex than human-written ones, and RECAP offers a solution that significantly reduces this verbosity without sacrificing effectiveness.
LLMs can infer formal specifications from tests alone, potentially transforming how we approach specification synthesis in industry.
Specification generation by LLMs remains a formidable challenge, with verification complexity often masking genuine quality differences.
Traditional probing methods fail to reveal the true memorization capabilities of large code LLMs, leading to inflated performance scores that obscure their genuine understanding.
Achieving optimal outcomes in probabilistic programs is now possible with a novel framework that synthesizes strategies across multiple objectives simultaneously.
SynAct slashes worst negative slack to 27% of bootstrap synthesis, revolutionizing timing optimization in logic synthesis.
Nearly 40% of GPU kernels deemed correct by standard tests are actually broken, revealing a critical flaw in current verification methods.
Achieving a 5.1x speedup in a legacy weather simulation while ensuring scientific validity reveals the critical role of validation in AI-assisted GPU porting.
Executable contracts can transform how we evolve legacy hardware, ensuring validated designs adapt without starting from scratch.
Matched execution scores can hide up to 64.3 points of command-path failure, revealing a critical gap in evaluating LLM coding agents.
CoT implementations can compute complex tree metrics in linear time, showcasing the potential of bounded-depth Transformers to tackle branching complexity effectively.
An innovative agentic workflow can fully modernize a 48-year-old quantum chemistry codebase without introducing any errors, achieving perfect validation across extensive tests.
Coding agents may appear compliant, but they actually underperform when faced with rules that challenge their default behaviors, revealing a critical flaw in current evaluation metrics.
Achieving high accuracy in domain model extraction using lightweight LLMs opens new avenues for reverse engineering in privacy-sensitive contexts.
Transforming agent failures into actionable recovery strategies, DARC enhances performance without bloating context, proving that less can be more in self-correction.
Achieving 99.5% citation validity and 96.4% figure editability, Spark-to-Paper transforms how research papers can be generated with unprecedented reliability and efficiency.
DexterSQL boosts Text-to-SQL accuracy by over 2.7% through innovative schema exploration and rule-based corrections that tackle common generation pitfalls.
Instruction alignment can significantly boost the accuracy of binary code representation learning, revealing deeper semantic connections than traditional methods.
Explicit role coordination among LLM-based tools can significantly enhance software quality, but it requires careful human oversight to manage deviations.
LLMs struggle to generate effective Triton kernels even when evaluated in realistic, production-like scenarios, revealing a critical gap in their practical applicability.
Despite the promise of multi-agent systems in software engineering, key frameworks still lack advanced features and show no significant performance difference in summarization tasks.
Valid input generation for Python APIs can drop from over 95% to as low as 31.6% when inter-parameter relationships are ignored, highlighting the critical role of dependency resolution in fuzzing.
A hybrid method combining symbolic execution and linear programming reveals sound upper bounds for resource consumption, overcoming the limitations of traditional analysis techniques.
WidgetGen achieves superior visual reconstruction performance without the constraints of fixed UI schemas, redefining the widget-to-code generation landscape.
An AI coding agent can autonomously refactor a complex codebase, achieving zero bugs and correcting over 200 defects without human oversight.
Visualization code editing is a tough nut to crack, with leading models only achieving a 74.46% pass rate on complex multimodal tasks.
SkillZip achieves unprecedented skill compression efficiency by transforming repetitive actions into reusable structures, eliminating the need for costly evaluations.
LLMs can transform the way legal compliance is integrated into software development, generating actionable requirements directly from complex legislation.
Ark, an open-source coding agent, solves 80% of software maintenance tasks while offering a clear architectural framework that could redefine how we study coding agents.
Quantum software patterns are not just theoretical; they occur in practice, and this tool reveals how they are actually composed and utilized in real projects.
The rapid rise of .cursorrules files in low-activity projects reveals a surprising gap in security considerations within AI-assisted programming prompts.
CausalRepair fixes 313 bugs with a cost of just $0.029 per bug, outpacing state-of-the-art APR methods by leveraging minimal causal context.
Nulls and bags in SQL are not just inconvenient—they're an avoidable design flaw that hinders query language efficiency.
A new proof technique reveals how to exhaustively detect subtyping failures in multiparty session types, enhancing program safety guarantees.
MergirafSemi cuts spurious merge conflicts significantly while maintaining high accuracy and low computational overhead across multiple programming languages.
GraphAlignCoder boosts code generation accuracy by 31.6% to 43.8% by embedding formal proof structures into the training process.
Programmatic skill learning can slash agent costs while enhancing performance, with SpeedRunner leading the charge in cost-efficient adaptation.
Long-horizon software development can thrive without relying on persistent agents, as demonstrated by Genesis's ability to evolve complex systems through finite-lived contributors.
SKILLER achieves up to 20.4 percentage points improvement in skill generation for small language models, making high-quality task execution accessible without the prohibitive costs of closed-source solutions.
Catastrophic remembering leads to an explosive growth of agentic prompts, but simple prompt comments can reverse this trend and enhance performance dramatically.
Up to 60% of attention heads can be pruned during domain adaptation without sacrificing performance, revealing a new avenue for efficient model tuning.
Over 3.7 million agent skills on GitHub reveal how developers are innovating in the absence of formal registries or type checks.
Persistent memory in EvoMem reduces redundant exploration in evolutionary code search, leading to faster and more effective optimization across multiple tasks.
Personalized skills for coding agents may not be the silver bullet developers hoped for, as generic skills often deliver superior performance.
Fine-tuning compact open-weight LLMs can outperform proprietary models in coding complex student metaphor responses, revolutionizing educational assessment.
Renderer format fails to significantly influence the outputs of LLMs, undermining assumptions about its role in theory-to-program translation.
GACP enables LLMs to provide precise code explanations by leveraging a multi-edge-type dependency graph, ensuring developers stay in control of their comprehension process.
Reducing false positives in smart contract analysis tools could save developers countless hours previously wasted on investigating non-issues.
N2NMatcher achieves remarkable improvements in binary code similarity analysis by effectively countering the disruptive effects of function inlining.
Activation probes can uncover security vulnerabilities in AI-generated code that traditional prompting methods completely miss.
Abstract compilation can now achieve optimal recurrence extraction for recursive programs, unlocking new avenues for precise cost analysis.
Memoir's innovative memory-driven approach reduces false-positive alerts in SAST tools to near perfection, revolutionizing the way security testing can be performed.
Evolved harnesses encode a shared abstract playbook that adapts to language-specific challenges, revealing a nuanced interplay between model limitations and engineering demands.
Semantic stability in Linux kernel code changes reveals that initial reviews drive most drift, challenging assumptions about later edits preserving function purpose.
Generating realistic fault data through HIL simulation could revolutionize how we validate automotive software systems in real time.
Coding agents are failing to meet user requests, with a mere 31.5% success rate, highlighting a critical gap in requirement recovery that must be addressed.
Jointly planning programs and their proofs can boost solve rates by over 11% while slashing API costs by nearly 40%.
Hand-written PTX kernels can outperform WMMA by up to 98.7x for INT4 operations, revealing critical insights into when low-level optimizations are worth the added complexity.
MetaStrategy achieves a remarkable 27.93% win rate in generative ranking calls while enhancing user engagement metrics significantly, all without increasing response time.
Closing the gap between rapid software evolution and slow defense procurement could redefine military readiness and operational effectiveness.
High initial confidence in LLMs can lead to catastrophic failures in complex reasoning tasks, but a new framework shows how to harness confidence trajectories for better outcomes.
Current AI coding agents struggle with large-scale refactoring tasks, achieving only a 41.2% success rate on a newly curated benchmark designed to challenge their capabilities.
Verifier-free consensus selection can boost CAD generation accuracy by up to 10% without the need for additional verification systems.
MDB-Link boosts Text-to-SQL accuracy by over 30 points while streamlining the schema selection process in multi-database scenarios.
A well-designed harness can transform model outputs into effective environmental actions, enabling programming agents to learn and adapt through iterative feedback.
Simulation traces can transform LLMs into powerful tools for diagnosing and improving complex scheduling policies, achieving unprecedented performance gains.
While loops and list comprehensions can double cognitive load for novice developers compared to for loops, revealing critical insights for code comprehension and maintainability.
Coding agents can achieve high accuracy on final specifications, yet 35% of them falter when faced with equivalent requirement histories, exposing a hidden vulnerability in their design.
Identifying 13 distinct bugs in Python refactoring implementations reveals significant gaps in current automated tools that could jeopardize software reliability.
A single click with SmellCC can eliminate 96.8% of Python code smells, transforming how developers manage technical debt.
Identical source code can yield binaries with performance differences so stark that they reveal systemic flaws in compiler optimizations.
The rapid evolution of GUI agents is overshadowed by significant engineering shortcomings that threaten their real-world applicability.
QUMUG generates quantum circuit mutants that are three times harder to detect, pushing the boundaries of quantum mutation analysis.
Achieving 74.7% migration quality, ECAT revolutionizes repository migration by leveraging adversarial entropy minimization to ensure functional completeness in real-world applications.
Achieving an 87.2% success rate, SiriusDeliver slashes data warehouse delivery times from hours to mere minutes, transforming enterprise analytics workflows.
Structured pseudocode boosts LLM performance in code generation, achieving a notable score of 4.78 against the best baseline's 4.31.
Checking the existence of uncomputation is coNP-hard, but new methods achieve unprecedented coverage in practical quantum circuit applications.
Showing all visible security tests upfront boosts functional and security success rates by over 19% on average, but not all models benefit equally.
OpenCodeReview achieves over double the accuracy of leading LLM code review agents while drastically reducing token consumption.
RETRACE achieves a 7% boost in patch verification accuracy by independently reconstructing and reconciling the problem and solution, ensuring coding agents generate reliable fixes.
Event presence can mislead interpretations of skill execution fidelity, revealing critical discrepancies in agent performance assessments.
FRAME achieves up to 13% better search accuracy by bridging the gap between UI code and graphical representations through innovative neuro-symbolic embeddings.
Infinite-dimensional probabilistic models can now be efficiently sampled using Hamiltonian Monte Carlo without losing the benefits of lazy evaluation.
Achieving up to 2.7x speedup in AI accelerator simulations by compressing multiple RTL nodes into a single instruction sequence could revolutionize chip design efficiency.
GRAFT achieves over 80% API coverage while ensuring high compilation success, outperforming existing fuzzing tools by a wide margin.