Search papers, labs, and topics across Lattice.
AI-driven code generation, program synthesis, automated debugging, and software engineering with LLMs.
#14 of 24
2
Achieving spatiotemporal composability allows components to be dynamically composed without side effects, revolutionizing how software systems manage dependencies and interactions.
A coding agent can maintain persistent world states and generate coherent visualizations, revolutionizing how we model complex environments.
An autonomous AI agent achieved 99.5% of the optimal solution for cell-edge power control while slashing inference costs by 600x, revolutionizing the role of researchers in algorithm design.
Narcissus achieves a 40% success rate on challenging ARC tasks by intelligently leveraging context-aware program proposals, far surpassing traditional methods that rely on static guidance.
KOPE's ability to retain and leverage past optimization experiences leads to a 1.54x speedup in kernel optimization tasks compared to existing methods.
Self-evolving coding agents can unwittingly amplify malicious skills, with self-poisoning rates soaring to 86.7% under tailored conditions.
Existing debugging methods for multi-agent systems fail to reliably reproduce and repair failures, but a new symptom-driven approach shows a dramatic improvement in effectiveness.
A self-evolving RCA harness outperforms traditional methods by reusing general agent capabilities, achieving a remarkable 59.0% accuracy in diagnosis.
Game development offers a powerful alternative to traditional reward signals, enabling RL systems to leverage executable environments for superior feedback and data generation.
Multi-agent collaboration in code generation can boost functional and security correctness by over 19 percentage points, reshaping how we approach secure coding.
RotDroid uncovers 94 previously unknown GUI rotation bugs in Android apps, highlighting a critical gap in current testing methodologies.
A lightweight defense can significantly reduce backdoor threats in RTL code generation without the heavy cost of full model retraining.
Adversarial attacks can mislead code search tools by altering identifiers, reducing retrieval accuracy by up to 77% without changing code functionality.
SkillShield slashes malware-generation severity from 3.37 to 0.58, showcasing a new frontier in prompt-space security for coding agents.
99% of Dockerfiles in a major enterprise are misconfigured, yet 83% of functional clusters contain high-quality reference implementations that could drastically improve security.
EAVA not only predicts software vulnerabilities but also provides crucial evidence, making automated assessments more reliable and actionable for security experts.
Achieving over 80% coverage in SQL testing, DBcover transforms how we approach reliability in RDBMSs by intelligently leveraging contextual reasoning.
LLMs show a stark performance drop in unit test generation when evaluated in realistic repository contexts, revealing a gap that could hinder practical deployment.
Keystroke-level logs reveal critical early indicators of student struggle, outperforming traditional execution logs in predicting coding success.
Teams adopting coding agents may be accelerating development at the cost of doubling cognitive complexity and increasing static-analysis warnings without proper AI configuration.
LLM coding agents can now leverage a literate programming environment that enhances context utilization and mimics human IDE capabilities.
Missing spatial constraints in CAD generation can be effectively filled using past design experiences, significantly enhancing the accuracy of text-to-CAD translation.
DeepRepoQA achieves substantial performance gains in code repository question answering by enabling LLM agents to perform multi-hop reasoning through systematic exploration.
Security-oriented prompts may reduce invalid outputs but paradoxically increase the prevalence of low-severity vulnerabilities in LLM-generated code.
LLMs can dramatically improve taint analysis for Android apps, achieving an F1-score of 0.96 compared to traditional tools' 0.55.
LLMs struggle with raw electromagnetic signal analysis, scoring as low as 21.2% on complex system design tasks, highlighting a critical gap in their reasoning capabilities.
IncSFS achieves a staggering 9.60x speedup over traditional pointer analysis methods, making it a game-changer for large-scale C/C++ projects.
Directly editing 3D meshes in Blender through a visual-centric, agentic approach could revolutionize how artists interact with complex geometry.
Transforming a single RGB image into a fully interactive 3D scene could redefine how we create simulation assets for Embodied AI.
IterCAD achieves a remarkable improvement in CAD code quality by enabling iterative refinement, correcting errors that traditional one-shot methods overlook.
CodeHID achieves a paradigm shift in code retrieval by organizing semantically similar snippets into a meaningful hierarchical index, leading to superior retrieval performance.
The "handoff tax" reveals that switching between models can significantly degrade performance while increasing costs, challenging the assumption that stronger models always yield better outcomes in coding tasks.
RubSE transforms UI-to-code generation by using structured rubrics to stabilize self-evolution, leading to more reliable visual repairs and higher performance ceilings.
LLMs can optimize database queries on GPUs, achieving over 2.5x speedup by leveraging advanced execution strategies and kernel fusion.
LLM feedback transforms passive learning into an active, engaging process, leading to richer student explanations and improved learning outcomes in programming.
SeriCrypt automates cryptographic message construction, uncovering security flaws and achieving unprecedented code coverage in fuzzing compared to traditional methods.
Paritok-4B compresses coding agent context to just 25.7% of its original size while retaining 86.5% of solve quality, outperforming existing models.
Ockhamareto achieves a staggering 49.9% mutation score while using 44% fewer tests than the best existing method, revolutionizing unit-test generation efficiency.
MGQL bridges the gap between GQL's informal specification and mechanized implementation, paving the way for reliable graph query processing.
MAGE transforms how coding agents operate by externalizing engineering intent and establishing governance, paving the way for more reliable software development.
Access to code generation via AI may rise, but control over software development is likely to become more exclusive, favoring those with deep expertise.
Faulty-code-driven test synthesis boosts code generation performance by 3% in LLMs, tackling reward hacking and validation issues head-on.
ReproAgent achieves unprecedented accuracy in translating research papers into executable code, outperforming existing methods by leveraging a dual-channel contract system.
LLMs struggle with semantic fidelity, achieving only a 30% alignment with scientific paper specifications, revealing a critical gap in AI-driven research reproducibility.
LLMs misclassify non-equivalent code as equivalent, with performance deteriorating on more complex tasks and showing unexpected language-specific biases.
SPECMINE reveals the intricate relationship between AI-generated specifications and code implementation, providing a treasure trove of data for understanding Spec-Driven Development.
The integration of machine learning in binary decompilation reveals critical gaps in evaluation standards that could redefine the field's future.
AI-generated code can be made secure and compliant, with some models achieving up to 100% compliance when guided by best practices.
Current evaluations of coding agents may mismeasure performance by conflating action, task, and step levels, revealing a critical flaw in how we assess agent execution.
Agentic AI could revolutionize software development, but it also introduces severe IT security risks that demand immediate attention.