Search papers, labs, and topics across Lattice.
100 papers published across 5 labs.
Proactively synthesizing project-specific issues can drastically enhance LLM agents' ability to resolve software problems effectively.
Continuous monitoring of developer efficiency is now possible even in non-employment contexts, revealing how barriers evolve over time.
PhysCaP enables robots to actively infer hidden physical properties, achieving superior manipulation efficiency without additional sensors.
New recursive measures for operations in VDM dramatically enhance proof obligation generation, ensuring greater model consistency.
Technical debt in software development is escalating, and without a shift to integrated testing and specification, we risk compounding vulnerabilities in AI systems.
PhysCaP enables robots to actively infer hidden physical properties, achieving superior manipulation efficiency without additional sensors.
New recursive measures for operations in VDM dramatically enhance proof obligation generation, ensuring greater model consistency.
Technical debt in software development is escalating, and without a shift to integrated testing and specification, we risk compounding vulnerabilities in AI systems.
Quality-Diversity search can yield unexpected winners from diverse ancestry, challenging the dominance of sequential champion approaches in program evolution.
Achieving 100% requirement coverage and Gherkin generation accuracy, this testing pipeline could redefine how automotive software is validated across decentralized systems.
Relying on a single oracle for feedback can inflate perceived gains in LLM test generation by nearly 15 percentage points, masking the true effectiveness of evolution strategies.
Execution traces can transform VDM-SL specifications into analyzable state-based models, enhancing both validation and debugging capabilities.
ChatGPT effortlessly solved all tested Qiskit homework assignments, revealing critical vulnerabilities in quantum software education assessments.
MidTool reveals that dedicated mid-training can significantly enhance LLMs' ability to utilize tools effectively, outperforming traditional post-training approaches.
Agents rely more on internal notes than traditional documentation, raising questions about the relevance of current documentation practices in automated coding environments.
Achieving a Micro-F1 score of 0.8085, this framework effectively balances adaptation and retention in LLMs for smart contract vulnerability detection.
Fine-tuning on synthetic data allows for a dramatic boost in code retrieval performance, achieving a balanced macro nDCG@10 of 0.5992 in a challenging domain.
UniLang enables pretrained LLMs to seamlessly integrate machine-native symbols, outperforming traditional models in diverse structured prediction tasks.
BreakGuard reveals that LLM-generated tests can detect over 30% of breaking changes in client applications, offering a cost-effective solution for maintaining software reliability.
Verifying compiler optimizations can transform the reliability of code generation, ensuring both performance and predictable compile times.
Hippogriff reveals a novel way to maintain general recursion in dependent type systems without sacrificing typechecking efficiency.
PRAXIS reveals that systematically extracting and surfacing tacit knowledge can dramatically enhance LLM performance in domain-specific code generation, outperforming traditional methods.
Achieving up to 414x speedup for complex image processing tasks, ParaWeb transforms how developers can harness parallelism in web applications.
Repo0 redefines code generation by enabling agents to construct entire software architectures from scratch, achieving unprecedented functionality and reliability.
Even the best coding agents fail to repair scientific software effectively, with pass rates below 50%, revealing critical gaps in their capabilities.
A novel deductive verification framework for weighted programming reveals how to express and automate verification conditions for complex quantitative models.
Achieving a 32% reduction in energy consumption while cutting waiting times by 30% could redefine operational efficiency in data centers powered by LLMs.
AlphaClifford consistently outperforms state-of-the-art synthesis heuristics by reducing gate counts while using a less expressive gate set, revolutionizing Clifford circuit optimization.
A contract-aware proof-repair tool can restore verification integrity in complex theorem proving, but challenges in operational end-to-end verification remain.
RoomWright transforms indoor scene synthesis by prioritizing functional usage, enabling interactive environments that are ready for real-world robotic applications.
Exploitation, not exploration, is the critical bottleneck in test-time scaling for language models, with selection processes yielding near-random results despite rich candidate pools.
Reward function design just got a major upgrade—MLREF achieves 25.2% better performance by reusing and evolving reward components across iterations.
AI-assisted mentorship can transform K-12 engineering education by empowering students to lead real-world projects while benefiting from undergraduate expertise.
Many SAST tools sacrifice detection capabilities for performance, but CAUSEC reveals that the underlying assumptions may not hold true, challenging the status quo in security analysis.
OdinEval reveals that even in niche programming languages, LLMs can achieve impressive repair accuracy, with top models scoring over 66% in resolving defects.
LLM-generated tests are less effective on poorly maintained code, revealing a surprising link between code health and token efficiency.
Mobile application repair performance varies dramatically across LLM agents, with success rates ranging from 22% to 90% depending on the evaluator used.
Users can now program robots with confidence, thanks to a system that combines LLMs and visual feedback to ensure intent alignment and reusability.
NeuroAssertion doubles the number of assertions and mutation coverage, transforming hardware verification by ensuring critical design behaviors are not overlooked.
LLMs may not be the silver bullet for automated program repair, as their integration can lead to worse performance than traditional methods.
Runtime for complex project-scheduling simulations can be slashed from over 1,200 seconds to under 200 seconds using agentic AI optimizations, saving researchers significant computational resources.
FACET achieves unprecedented task synthesis quality by preserving source intent and ensuring executable state consistency, leading to more reliable terminal agents.
Proactively synthesizing project-specific issues can drastically enhance LLM agents' ability to resolve software problems effectively.
SemaPLC achieves a remarkable 52.2% dynamic behavior score, setting a new standard for verifying the operational integrity of PLC code generation.
Acceptance in Code World Models certifies sample consistency but often overlooks critical events, leading to substantial planning failures.
Evolving procedural content generators with reusable programming primitives boosts fitness scores across multiple classic games, revealing a powerful synergy between abstraction and evolutionary search.
Tying every view of a digital artifact to its version can boost knowledge-work agent performance by over 12 points on critical tasks.
LLMs can significantly outperform traditional symbolic methods in PDDL model repair, but their reliability still falters in complex domains.
Human expertise remains essential in the agentic coding process, even as productivity in simulation library development accelerates significantly.
Achieving an 86% reduction in nonconformance bugs by standardizing tests across a dozen programming languages reveals the power of a unified testing approach.
REChart slashes reasoning token usage by 79% while achieving state-of-the-art chart-editing performance, tackling the "overthinking" problem in large reasoning models.
Bridging semantic logic with geometric constraints, aDSL enables agents to create complex 3D structures more reliably than traditional LLM approaches.
Skill Optimizers trained through execution feedback can outperform traditional models by over 9 points, revealing a critical gap in agent learning methodologies.
TraceSQL achieves superior SQL verification accuracy while offering unprecedented transparency into the decision-making process of the model.
Even state-of-the-art MLLMs fail to accurately reconstruct academic documents, revealing a critical gap in machine understanding of scientific knowledge.
AdaRare achieves higher median edge coverage than traditional AFL++ by intelligently coordinating multiple control mechanisms, revealing critical vulnerabilities in firmware that were previously undetected.
The integration of data engineering and software engineering practices could fundamentally redefine how we approach the software lifecycle in AI systems.
Execution validation of LLM-inferred dependencies boosts REST API testing success rates to 88.1%, revealing critical failures that traditional methods miss.
A unified message model can revolutionize how we automate and integrate complex embedded systems by providing a clear, formal basis for serial communication.
SNIPTEST reveals that targeted fuzzing of code slices can significantly enhance the validation of static analysis warnings, achieving over 54% success in identifying true vulnerabilities.
Energy-aware knowledge distillation can slash inference energy consumption by up to 90%, challenging the reliability of FLOPs as a metric for sustainability in LLMs.
COMMITGUARD narrows down hundreds of potential bugs to just a handful of actionable reports, validating over 70% of them as real issues.
An LLM coding agent achieved a 100% success rate in robot manipulation with 46% fewer steps than traditional methods, revolutionizing how we approach task learning without human demonstrations.
Achieving up to 44.9x speedup in concolic execution for WebAssembly could revolutionize how we analyze and test complex software systems.
LEGO-RL boosts coding agent performance by up to 8.4% while ensuring robust training signals and execution reliability.
ICD-Deepresearch outperforms traditional methods, achieving a 51% usefulness rating from physicians, highlighting its potential to enhance clinical decision-making.
Viewing technical debt as a strategic asset rather than a liability could revolutionize how startups approach software experimentation and risk management.
Prior audit-repair episodes can significantly lower false alarm rates in language model verifiers, challenging assumptions about accumulated-message effects.
LACE transforms AutoML by enabling the evolution of fully executable Python pipelines, allowing practitioners to directly edit and reuse generated models.
Coordination among AI agents can be optimized by leveraging shared files, reducing communication overhead by 42% in message-heavy tasks, but simply adding a coordinator offers no advantage. WHY_IT MATTERS: These insights could transform how we design multi-agent systems, emphasizing the importance of task structure and communication strategies in enhancing collaborative efficiency.
SPC achieves a staggering 97.4% accuracy in text-to-SQL tasks, leaving traditional DDL-to-SQL methods in the dust.
State transformers can simplify the mechanization of distributed programming, ensuring deadlock freedom while abstracting away local complexities.
Visual observations can be transformed into editable city layouts with a surprising level of detail and accuracy using a multimodal large language model.
Executable Code Knowledge enables coding agents to carry validation evidence directly within code, achieving perfect precision and recall in patch tasks where traditional methods fail.
SOPD not only outperforms traditional distillation methods but also redefines how we think about trajectory corrections in model training.
Credentialing systems risk becoming obsolete as generative AI alters the landscape of skill assessment, with old competition medals losing 82% of their predictive power.
Proprietary LLMs can match human educators in delivering relevant answers from curated lecture videos, potentially transforming how programming education leverages AI.
LLMs may inherently encode vulnerability signals in their activations, enabling lightweight, model-native vulnerability detection that rivals traditional methods.
Over a third of Java libraries harbor hidden dependencies that can introduce breaking changes and security vulnerabilities, yet remain unnoticed by developers.
Choosing a vibe coding tool isn't just about productivity; it involves navigating complex trade-offs in code quality that can impact long-term maintainability.
Trustworthiness in open-source dependencies can now be quantified with a single score that reveals critical insights into software supply chain security.
P2's approach to vulnerability remediation achieved up to a 69% reduction in findings, challenging the assumption that better code generation models always yield the best security outcomes.
Spec-Driven Test Generation boosts bug detection rates by nearly 10% by making LLMs reason about code contracts before generating tests.
Continuous monitoring of developer efficiency is now possible even in non-employment contexts, revealing how barriers evolve over time.
TDD-Agent transforms how LLMs generate code by making tests integral to the development process, leading to higher correctness and more effective tests.
Missing critical facts during coding tasks leads to systematic failures across multiple models, highlighting the importance of coherence in repository-scale coding agents.
As AI automates coding, the real challenge shifts to ensuring that human specifications are accurate and verifiable, revealing a critical paradox in software development.
Generative AI systems can outperform average students in programming assessments, but they still falter on advanced OOP concepts and graphics tasks.
Automated techniques can uncover 166 hidden faults in real-world REST APIs, revealing critical vulnerabilities in HTTP semantics adherence.
Minor control failures in AI-driven building operations can lead to irreversible energy waste and occupant discomfort, necessitating a new software engineering framework.
Newer LLMs may ace build tasks, but they struggle with delivering correct app behaviors, especially when specifications are involved.
Slice transforms continuous probabilistic programs into discrete forms, enabling exact inference where previous methods fail.
Closing the loop on robot manipulation failures, VLCP rewrites control code in real-time, achieving a tenfold increase in success rates over traditional methods.
GoalEvolve achieves a 30.67% improvement in post-route TNS while simultaneously reducing leakage and dynamic power, showcasing a transformative approach to physical design algorithm evolution.
Achieving up to 19x performance improvements on resource-constrained workloads by harnessing the power of GPUs could redefine how we approach software efficiency.
Energy optimization in circuit synthesis can be dramatically improved with Renesis, achieving up to 91% of the default energy across various benchmarks.
Transforming operational telemetry into actionable repair context, ORCA outperforms traditional methods in automated program repair for microservices.
Achieving over 81% accuracy in actionable software requirements extraction, LadderTeam automates a traditionally burdensome process, revolutionizing user feedback collection.
Kozuchi Agent achieves a remarkable 74.8% success rate in software repair tasks, setting a new standard for open-weight agents in the field.
Despite high pull request acceptance rates, a staggering 93% of valuable commits in software forks remain unmerged, highlighting a critical gap in collaborative development.
SMTpip resolves Python dependency conflicts 6.9 times faster than pip by leveraging SMT-based inference for interpreter-aware environment construction.
LLMs struggle with procedural database programming, revealing critical gaps in their capabilities that traditional benchmarks overlook.
Layered memory-safety defenses in C can incur hidden performance penalties and compatibility issues that native languages like Rust and Go avoid entirely.
Normalizing vague software requirements can dramatically enhance traceability, but too much normalization risks undermining high-quality specifications.
AutoSQL outperforms traditional methods by over 20% in SQL template extraction, revealing hidden efficiencies in ORM codebases.