Search papers, labs, and topics across Lattice.
100 papers published across 5 labs.
MergeSE can recover cross-domain performance from specialized models in under five seconds, revolutionizing how software engineers handle model merging without retraining.
Merging models can yield a cross-domain clone detector that generalizes better to unseen AI-generated duplicates while maintaining high performance.
Hybrid nested optimization outperforms traditional methods by effectively decoupling structural design from parameter tuning in LLM-driven evolutionary algorithms.
A self-developing agent that autonomously improves its own coding capabilities while navigating operational safety challenges has set new performance benchmarks.
Self-improving coding agents can achieve unprecedented efficiency and generalizability by learning from multiple task trajectories simultaneously.
Hybrid nested optimization outperforms traditional methods by effectively decoupling structural design from parameter tuning in LLM-driven evolutionary algorithms.
A self-developing agent that autonomously improves its own coding capabilities while navigating operational safety challenges has set new performance benchmarks.
Self-improving coding agents can achieve unprecedented efficiency and generalizability by learning from multiple task trajectories simultaneously.
Complex search behaviors typically requiring explicit controllers can emerge naturally from an agent's reasoning process, leading to substantial performance gains in optimization tasks.
SkillZip achieves a remarkable 3.46x compression ratio while preserving 99.2% of dependencies and 98.7% of verifier reachability, revolutionizing how agent skill libraries can be managed.
CalibForge reveals that adversarial calibration can dramatically enhance the effectiveness of training data for terminal agents, leading to unprecedented performance improvements on standard benchmarks.
GSE achieves up to 180% improvement in recall for coding agents, revolutionizing how skills are evolved and reused in automated programming tasks.
Programmatic tool calling outperforms traditional JSON tool calling in 11 out of 14 language models, showcasing a significant leap in efficiency and robustness.
Python remains the overwhelming choice for code generation in LLMs, but many selections are based on convenience rather than project needs, revealing critical flaws in model reasoning.
Cheap models can recover early evolutionary progress, enabling a shift in budget allocation that dramatically enhances LLM-driven algorithm discovery.
Game-hopping proofs can now be mechanized in Lean with unprecedented clarity and integration, enabling more robust cryptographic security verification.
RustGo prunes 78.49% of irrelevant paths, accelerating bug discovery in unsafe Rust code while uncovering 13 previously unknown vulnerabilities.
AssertMate outperforms existing LLM-based assertion generation methods, achieving higher accuracy and bug detection through a novel multi-agent approach.
Quantum programs can be optimized for real-world noise, revealing that classical probabilistic branching is essential for achieving optimal performance.
Optimizing developer assignments based on expertise can boost task completion efficiency by over 70% in open-source projects.
Software engineers are increasingly dependent on LLMs, risking overreliance that could undermine traditional practices like peer consultation and documentation.
Causal memory can boost execution accuracy in Text-to-SQL tasks, but its effectiveness is context-dependent and not universally superior to other retrieval methods.
CodeGrep slashes token usage and rounds by over 15% while maintaining high resolve rates, revolutionizing how LLM coding agents handle file retrieval.
Autogrammar can automatically learn context-free grammars that boost language model performance, achieving near-perfect precision and tripling execution speed on DSL tasks.
ARIA can achieve a staggering 94.5% success rate in implanting covert backdoors in customized LLMs while ensuring high task performance.
AgentExecutor outperforms existing methods by achieving up to 94% code coverage while slashing execution time by over 80%.
Iterative self-repair methods may hinder fault detection, but DCAware's dual-context approach reveals faults more effectively without the computational burden.
WasmMend achieves a remarkable 70% fix rate for discrepancies between WebAssembly and native binaries, showcasing the power of divergence-guided reasoning in automated repair.
Small LLMs can achieve the same fault diagnosis accuracy as larger models, challenging the assumption that bigger is always better in automotive software validation.
JDomInO keeps Java code and domain models in sync, preventing the costly drift that undermines effective Domain-Driven Design.
RepoOMP achieves up to 8.96x speedup in parallelizing hotspots while reducing agent-side token costs by nearly 68%.
ShimGen not only matches but surpasses manually-designed protocols in consistency, revealing critical performance gains in heterogeneous memory systems.
Emerging AI tools are reshaping software engineering education, revealing significant curricular trends that could redefine how future developers are trained.
AI-generated C++ code incurs 5-8% more compute costs and increased review efforts due to its distinct quality profile, but targeted feedback can substantially improve its performance.
Fine-tuning on planning-aware trajectories can enhance model performance across diverse CLI environments, mitigating the pitfalls of scaffold-specific training.
SuperScout matches the best-performing model's solve rate while slashing costs to one-fifth, revolutionizing how we approach coding agent routing.
Language models are far more capable in scientific coding than previously thought, with corrected evaluations revealing accuracy improvements of up to 92%.
Most coding agents fail to proactively fix bugs without issue reports, revealing a critical gap in their capabilities.
Candidate ordering, not just model confidence, is the key to effective code deletion in AI systems, leading to a significant boost in verified coverage.
Execution consistency can cut misleading feedback in code generation by over 87%, transforming how LLMs self-correct.
FineMote reduces perception-to-decision latency in robotic systems by statically orchestrating control firmware for tree-structured device models, achieving better timing behavior with minimal overhead.
COMPAS boosts code generation performance by 15% while slashing costs by over 86%, revealing the critical interplay between task difficulty and optimization choices.
Formalizing the soundness of symbolic execution tools reveals that path-merging can significantly enhance software reliability without sacrificing correctness.
Training-free anomaly detection and repair can drastically enhance the validity of CAD models without the computational burden of retraining large generative models.
LLMs struggle to match expert-level performance in GPU communication tasks, with top models achieving only 30.7% success in generating efficient code.
Static discovery of cryptographic assets can reveal hidden vulnerabilities and streamline post-quantum migration in software systems.
Reasoning Core achieves unprecedented performance in completion-supervised reasoning tasks, outpacing existing procedural datasets and revealing critical design insights for effective training.
LLMs can revolutionize hardware security by autonomously identifying vulnerabilities in Verilog designs before they become embedded in silicon.
Strings alone can dramatically enhance secret detection, achieving over 80% semantic retention with only a third of the context.
Coding agents are alarmingly susceptible to malicious skill files, with exploitation rates exceeding 95% in some cases.
Replacing just 12% of traditional training data with OctoLong's curated code contexts leads to substantial improvements in long-range retrieval and state tracking for language models.
RepairFormer repairs structured inputs with an 88% success rate while preserving 94% of the original content, revolutionizing how we handle corrupted data files.
Efficient resource estimation for fault-tolerant quantum programs can lead to substantial savings and improved algorithm performance without sacrificing programmability.
AdaptAgent achieves superior code adaptation by leveraging a multi-agent approach that mirrors real developer practices, outperforming traditional methods in semantic correctness.
Blind and low-vision developers face critical accessibility barriers in AI tools, with three key issues dominating the landscape.
Automatic tracking of quantum compiler provenance could revolutionize how researchers evaluate and optimize quantum transpilation processes.
Achieving up to 3.14x speedup in pointer analysis by leveraging compiler optimizations could revolutionize static analysis performance benchmarks.
LLMs exhibit a troubling tendency to prioritize code generation over genuine architectural understanding, revealing a critical gap in their evaluation metrics.
Compiling recursive logic programs into quantum annealers could revolutionize how we solve complex computational problems by ensuring optimal solutions are reached efficiently.
Agent Plans in open-source repositories reveal critical insights into how AI coding tools can be effectively guided through structured task-oriented artifacts.
Counterexample-based fault localization techniques can dramatically enhance debugging in verification-aware languages, outperforming traditional state-based methods.
Mosaic's innovative modular reasoning approach enables substantial performance gains in bit-precise program verification, outperforming existing solvers.
The firmware update intended to fix safety issues also introduced critical security flaws, revealing the complexities of patching legacy systems.
Recursive task synthesis not only slashes generation costs to $0.05 per task but also produces increasingly complex challenges that boost model performance by up to 10 points on key benchmarks.
Equivalent prompts can yield identical code outputs, ensuring semantic robustness in AI coding workflows despite minor changes in specifications.
AutoSND uncovers more effective and interpretable network dismantling heuristics by transforming execution evidence into actionable structural policies.
LLMs are transforming PDE workflows, but their effectiveness is hampered by data scarcity and the challenge of applying simulations to real-world scenarios.
Achieving up to 2,500x speed improvements, string2string Studio revolutionizes string-to-string analysis by making complex algorithms accessible and interactive in the browser.
Mapping SQL errors to conceptual gaps reveals that students' misunderstandings are often more profound than mere syntax issues, offering a new lens for educational feedback.
CLEAR achieves a staggering 130.7% improvement in vulnerability detection by harnessing causal knowledge graphs to navigate complex dependencies in code.
ReBug achieves nearly 50% success in reproducing web GUI bugs from natural-language reports, transforming how developers address software maintenance challenges.
Regrading model outputs can shift correctness labels by 9% and double the performance spread, revealing critical flaws in conventional evaluation methods for LLM code generation.
Stale comments in the Linux kernel mislead maintainers, but ReCite offers a robust solution that identifies and repairs these inconsistencies with an impressive 89% utility in suggestions.
Persistent state and localized search in TraceCAD can dramatically enhance CAD generation quality and reliability, revealing a critical dependency in agentic design.
Current coding agents falter in preserving content integrity while reconstructing UI regions, revealing critical gaps in their iterative coding capabilities.
Early failure prediction can save up to 20.4% of execution tokens while improving resolution rates in software engineering tasks, transforming how agents handle long trajectories.
NotDec outperforms Ghidra by achieving a 100% recompilation success rate and recovering 85.33% of struct member accesses, setting a new standard for WebAssembly decompilation.
Coding agents boost task completion rates by over 30%, but at the cost of diminishing public knowledge resources for future developers.
LiveEvalBench reveals that evaluating web generation requires a dynamic, collaborative approach that reflects the interactive nature of frontend development.
MergeSE can recover cross-domain performance from specialized models in under five seconds, revolutionizing how software engineers handle model merging without retraining.
Artifact-anchored memory not only preserves verification claims but also ensures they remain valid even as upstream content changes, outperforming traditional methods in critical coding tasks.
Static analysis tools miss 93.6% of reuse opportunities due to hardware incompatibilities, but a new RAG pipeline accurately identifies reusable functions with 97.5% validation accuracy.
Debugging Emfrp applications just got easier—now you can seamlessly navigate between high-level abstractions and low-level C/C++ code.
EffiHolmes achieves a remarkable 15% improvement in function-level accuracy for time inefficiency localization, setting a new standard in the field.
CURATE transforms workflow management by enabling seamless composition, reuse, and deployment of code, bridging the gap between generation and execution.
POVGEN generates Proofs-of-Vulnerability for 78.98% of cases, revealing flaws in existing patches and uncovering new vulnerabilities, all while minimizing costs.
Merging models can yield a cross-domain clone detector that generalizes better to unseen AI-generated duplicates while maintaining high performance.
Task-aware evaluation reveals that LLMs prioritize Correctness and Abstraction over Conciseness and Fluency, challenging traditional human-centric summary metrics.
Novice developers using AgentForge significantly improved their software engineering skills and critical collaboration with AI, despite facing varying interaction challenges.
High-level quantum programming can significantly reduce cognitive load and errors by treating computations as structured compositions of quantum registers.
Random exploration outperforms LLMs in TUI testing, revealing that model choice may not be as critical as previously thought.
MLLMs exhibit a staggering 80.22% bias towards incorrect outputs when faced with repeated UI patterns, revealing a critical flaw in their code generation capabilities.
Achieving a 9.92 percentage point boost in functional correctness for code generation without changing the backbone model architecture is a game-changer for LLM applications.
HyperFL achieves up to 16.7% better fault localization accuracy by adapting to the unique characteristics of diverse issue reports in real-time.
AI adoption can either reduce or exacerbate community smells in software teams, depending on whether the work is specialized or coordinated.
Smaller models can achieve up to 95.7% structural success in enterprise workflow generation, challenging the notion that only larger models are viable for production use.
PSC achieves up to 700x faster grammar-constrained decoding, making LLMs as efficient as unconstrained decoding for structured output generation.
LLMs can achieve a staggering 94.8% correctness in generating compiler optimizations that traditional methods miss, revealing a new frontier in program analysis.
Models can identify abstract structures in isolation but fail dramatically when asked to translate that understanding across different contexts, revealing a critical gap in AI's reasoning abilities.
Self-evolving coding agents can revolutionize software development by learning from past interactions, but they also face significant challenges in reliability and safety.
A hands-on framework that transforms AI education in power systems, making complex concepts accessible to newcomers and practitioners alike.
Translating LTL to LTLf+ unlocks efficient finite automata techniques for a wide array of AI applications without sacrificing complexity.
Harness-R1 achieves a remarkable 9.3 percentage point increase in task success rates by learning to edit agent harnesses in response to failure trajectories.
Shifting from solution-centric to information-centric decision-making, Iris achieves unprecedented performance in autonomous ML engineering tasks.
A novel digital twin approach reveals vulnerabilities in AArch64 machine code without needing source code, achieving comprehensive detection with clear explanations.