Search papers, labs, and topics across Lattice.
AI-driven code generation, program synthesis, automated debugging, and software engineering with LLMs.
#11 of 24
4
CalibForge reveals that adversarial calibration can transform terminal task generation, leading to up to 30.04 percentage points improvement in model performance.
GSE achieves up to 180% improvement in recall for coding agents, revolutionizing how skills are evolved and reused in automated programming tasks.
Programmatic tool calling outperforms traditional JSON tool calling in 11 out of 14 language models, showcasing a significant leap in efficiency and robustness.
Python remains the overwhelming choice for code generation in LLMs, but many selections are based on convenience rather than project needs, revealing critical flaws in model reasoning.
Cheap models can recover early evolutionary progress, enabling a shift in budget allocation that dramatically enhances LLM-driven algorithm discovery.
Game-hopping proofs can now be mechanized in Lean with unprecedented clarity and integration, enabling more robust cryptographic security verification.
RustGo prunes 78.49% of irrelevant paths, accelerating bug discovery in unsafe Rust code while uncovering 13 previously unknown vulnerabilities.
AssertMate outperforms existing LLM-based assertion generation methods, achieving higher accuracy and bug detection through a novel multi-agent approach.
Quantum programs can be optimized for real-world noise, revealing that classical probabilistic branching is essential for achieving optimal performance.
Optimizing developer assignments based on expertise can boost task completion efficiency by over 70% in open-source projects.
Software engineers are increasingly dependent on LLMs, risking overreliance that could undermine traditional practices like peer consultation and documentation.
Causal memory can boost execution accuracy in Text-to-SQL tasks, but its effectiveness is context-dependent and not universally superior to other retrieval methods.
CodeGrep slashes token usage and rounds by over 15% while maintaining high resolve rates, revolutionizing how LLM coding agents handle file retrieval.
SkillZip achieves a remarkable 3.46x compression ratio while preserving 99.2% of dependencies and 98.7% verifier reachability, revolutionizing how we manage agent skill libraries.
Autogrammar can automatically learn context-free grammars that boost language model performance, achieving near-perfect precision and tripling execution speed on DSL tasks.
ARIA can achieve a staggering 94.5% success rate in implanting covert backdoors in customized LLMs while ensuring high task performance.
AgentExecutor outperforms existing methods by achieving up to 94% code coverage while slashing execution time by over 80%.
Iterative self-repair methods may hinder fault detection, but DCAware's dual-context approach reveals faults more effectively without the computational burden.
Fine-tuning with planning-aware trajectories can significantly enhance agent performance across diverse scaffolds, overcoming the limitations imposed by conventional training environments.
WasmMend achieves a remarkable 70% fix rate for discrepancies between WebAssembly and native binaries, showcasing the power of divergence-guided reasoning in automated repair.
Small LLMs can achieve the same fault diagnosis accuracy as larger models, challenging the assumption that bigger is always better in automotive software validation.
JDomInO keeps Java code and domain models in sync, preventing the costly drift that undermines effective Domain-Driven Design.
RepoOMP achieves up to 8.96x speedup in parallelizing hotspots while reducing agent-side token costs by nearly 68%.
ShimGen not only matches but surpasses manually-designed protocols in consistency, revealing critical performance gains in heterogeneous memory systems.
Emerging AI tools are reshaping software engineering education, revealing significant curricular trends that could redefine how future developers are trained.
SuperScout matches the best-performing model's solve rate while slashing costs to one-fifth, revolutionizing how we approach coding agent routing.
Language models are far more capable in scientific coding than previously thought, with corrected evaluations revealing accuracy improvements of up to 92%.
Most coding agents fail to proactively fix bugs without issue reports, revealing a critical gap in their capabilities.
Candidate ordering, not just model confidence, is the key to effective code deletion in AI systems, leading to a significant boost in verified coverage.
Execution consistency can cut misleading feedback in code generation by over 87%, transforming how LLMs self-correct.
FineMote reduces perception-to-decision latency in robotic systems by statically orchestrating control firmware for tree-structured device models, achieving better timing behavior with minimal overhead.
COMPAS boosts code generation performance by 15% while slashing costs by over 86%, revealing the critical interplay between task difficulty and optimization choices.
Formalizing the soundness of symbolic execution tools reveals that path-merging can significantly enhance software reliability without sacrificing correctness.
Training-free anomaly detection and repair can drastically enhance the validity of CAD models without the computational burden of retraining large generative models.
LLMs struggle to match expert-level performance in GPU communication tasks, with top models achieving only 30.7% success in generating efficient code.
Static discovery of cryptographic assets can reveal hidden vulnerabilities and streamline post-quantum migration in software systems.
Reasoning Core achieves unprecedented performance in completion-supervised reasoning tasks, outpacing existing procedural datasets and revealing critical design insights for effective training.
LLMs can revolutionize hardware security by autonomously identifying vulnerabilities in Verilog designs before they become embedded in silicon.
Strings alone can dramatically enhance secret detection, achieving over 80% semantic retention with only a third of the context.
Coding agents are alarmingly susceptible to malicious skill files, with exploitation rates exceeding 95% in some cases.
Replacing just 12% of traditional training data with OctoLong's curated code contexts leads to substantial improvements in long-range retrieval and state tracking for language models.
RepairFormer repairs structured inputs with an 88% success rate while preserving 94% of the original content, revolutionizing how we handle corrupted data files.
Efficient resource estimation for fault-tolerant quantum programs can lead to substantial savings and improved algorithm performance without sacrificing programmability.
AdaptAgent achieves superior code adaptation by leveraging a multi-agent approach that mirrors real developer practices, outperforming traditional methods in semantic correctness.
Blind and low-vision developers face critical accessibility barriers in AI tools, with three key issues dominating the landscape.
Automatic tracking of quantum compiler provenance could revolutionize how researchers evaluate and optimize quantum transpilation processes.
Achieving up to 3.14x speedup in pointer analysis by leveraging compiler optimizations could revolutionize static analysis performance benchmarks.
LLMs exhibit a troubling tendency to prioritize code generation over genuine architectural understanding, revealing a critical gap in their evaluation metrics.
Compiling recursive logic programs into quantum annealers could revolutionize how we solve complex computational problems by ensuring optimal solutions are reached efficiently.
Agent Plans in open-source repositories reveal critical insights into how AI coding tools can be effectively guided through structured task-oriented artifacts.