Search papers, labs, and topics across Lattice.
87 papers published across 6 labs.
Achieving spatiotemporal composability allows components to be dynamically composed without side effects, revolutionizing how software systems manage dependencies and interactions.
Achieving a staggering 95.5% on the ARC-AGI-3 benchmark, Prime Agent redefines the capabilities of long-horizon coding agents by seamlessly integrating memory and computation.
MAGE transforms how coding agents operate by externalizing engineering intent and establishing governance, paving the way for more reliable software development.
A coding agent can maintain persistent world states and generate coherent visualizations, revolutionizing how we model complex environments.
An autonomous AI agent achieved 99.5% of the optimal solution for cell-edge power control while slashing inference costs by 600x, revolutionizing the role of researchers in algorithm design.
Achieving spatiotemporal composability allows components to be dynamically composed without side effects, revolutionizing how software systems manage dependencies and interactions.
A coding agent can maintain persistent world states and generate coherent visualizations, revolutionizing how we model complex environments.
An autonomous AI agent achieved 99.5% of the optimal solution for cell-edge power control while slashing inference costs by 600x, revolutionizing the role of researchers in algorithm design.
Narcissus achieves a 40% success rate on challenging ARC tasks by intelligently leveraging context-aware program proposals, far surpassing traditional methods that rely on static guidance.
KOPE's ability to retain and leverage past optimization experiences leads to a 1.54x speedup in kernel optimization tasks compared to existing methods.
Self-evolving coding agents can unwittingly amplify malicious skills, with self-poisoning rates soaring to 86.7% under tailored conditions.
Existing debugging methods for multi-agent systems fail to reliably reproduce and repair failures, but a new symptom-driven approach shows a dramatic improvement in effectiveness.
A self-evolving RCA harness outperforms traditional methods by reusing general agent capabilities, achieving a remarkable 59.0% accuracy in diagnosis.
Game development offers a powerful alternative to traditional reward signals, enabling RL systems to leverage executable environments for superior feedback and data generation.
Multi-agent collaboration in code generation can boost functional and security correctness by over 19 percentage points, reshaping how we approach secure coding.
RotDroid uncovers 94 previously unknown GUI rotation bugs in Android apps, highlighting a critical gap in current testing methodologies.
A lightweight defense can significantly reduce backdoor threats in RTL code generation without the heavy cost of full model retraining.
Adversarial attacks can mislead code search tools by altering identifiers, reducing retrieval accuracy by up to 77% without changing code functionality.
SkillShield slashes malware-generation severity from 3.37 to 0.58, showcasing a new frontier in prompt-space security for coding agents.
99% of Dockerfiles in a major enterprise are misconfigured, yet 83% of functional clusters contain high-quality reference implementations that could drastically improve security.
EAVA not only predicts software vulnerabilities but also provides crucial evidence, making automated assessments more reliable and actionable for security experts.
Achieving over 80% coverage in SQL testing, DBcover transforms how we approach reliability in RDBMSs by intelligently leveraging contextual reasoning.
LLMs show a stark performance drop in unit test generation when evaluated in realistic repository contexts, revealing a gap that could hinder practical deployment.
Keystroke-level logs reveal critical early indicators of student struggle, outperforming traditional execution logs in predicting coding success.
Teams adopting coding agents may be accelerating development at the cost of doubling cognitive complexity and increasing static-analysis warnings without proper AI configuration.
LLM coding agents can now leverage a literate programming environment that enhances context utilization and mimics human IDE capabilities.
Missing spatial constraints in CAD generation can be effectively filled using past design experiences, significantly enhancing the accuracy of text-to-CAD translation.
DeepRepoQA achieves substantial performance gains in code repository question answering by enabling LLM agents to perform multi-hop reasoning through systematic exploration.
Security-oriented prompts may reduce invalid outputs but paradoxically increase the prevalence of low-severity vulnerabilities in LLM-generated code.
LLMs can dramatically improve taint analysis for Android apps, achieving an F1-score of 0.96 compared to traditional tools' 0.55.
LLMs struggle with raw electromagnetic signal analysis, scoring as low as 21.2% on complex system design tasks, highlighting a critical gap in their reasoning capabilities.
IncSFS achieves a staggering 9.60x speedup over traditional pointer analysis methods, making it a game-changer for large-scale C/C++ projects.
Directly editing 3D meshes in Blender through a visual-centric, agentic approach could revolutionize how artists interact with complex geometry.
Transforming a single RGB image into a fully interactive 3D scene could redefine how we create simulation assets for Embodied AI.
IterCAD achieves a remarkable improvement in CAD code quality by enabling iterative refinement, correcting errors that traditional one-shot methods overlook.
CodeHID achieves a paradigm shift in code retrieval by organizing semantically similar snippets into a meaningful hierarchical index, leading to superior retrieval performance.
The "handoff tax" reveals that switching between models can significantly degrade performance while increasing costs, challenging the assumption that stronger models always yield better outcomes in coding tasks.
RubSE transforms UI-to-code generation by using structured rubrics to stabilize self-evolution, leading to more reliable visual repairs and higher performance ceilings.
LLMs can optimize database queries on GPUs, achieving over 2.5x speedup by leveraging advanced execution strategies and kernel fusion.
LLM feedback transforms passive learning into an active, engaging process, leading to richer student explanations and improved learning outcomes in programming.
SeriCrypt automates cryptographic message construction, uncovering security flaws and achieving unprecedented code coverage in fuzzing compared to traditional methods.
Paritok-4B compresses coding agent context to just 25.7% of its original size while retaining 86.5% of solve quality, outperforming existing models.
Ockhamareto achieves a staggering 49.9% mutation score while using 44% fewer tests than the best existing method, revolutionizing unit-test generation efficiency.
MGQL bridges the gap between GQL's informal specification and mechanized implementation, paving the way for reliable graph query processing.
MAGE transforms how coding agents operate by externalizing engineering intent and establishing governance, paving the way for more reliable software development.
Access to code generation via AI may rise, but control over software development is likely to become more exclusive, favoring those with deep expertise.
Faulty-code-driven test synthesis boosts code generation performance by 3% in LLMs, tackling reward hacking and validation issues head-on.
ReproAgent achieves unprecedented accuracy in translating research papers into executable code, outperforming existing methods by leveraging a dual-channel contract system.
LLMs struggle with semantic fidelity, achieving only a 30% alignment with scientific paper specifications, revealing a critical gap in AI-driven research reproducibility.
LLMs misclassify non-equivalent code as equivalent, with performance deteriorating on more complex tasks and showing unexpected language-specific biases.
SPECMINE reveals the intricate relationship between AI-generated specifications and code implementation, providing a treasure trove of data for understanding Spec-Driven Development.
The integration of machine learning in binary decompilation reveals critical gaps in evaluation standards that could redefine the field's future.
AI-generated code can be made secure and compliant, with some models achieving up to 100% compliance when guided by best practices.
Current evaluations of coding agents may mismeasure performance by conflating action, task, and step levels, revealing a critical flaw in how we assess agent execution.
Agentic AI could revolutionize software development, but it also introduces severe IT security risks that demand immediate attention.
PatchWrite achieves a flawless preservation of manuscript integrity, maintaining 100% accuracy in edits while traditional methods fail completely.
Injecting Agent Skills in web development can reduce task performance by up to 4.2%, challenging the assumption that more Skills always lead to better outcomes.
Fine-grained preference optimization at critical decision points can dramatically enhance the reliability of SQL query generation.
Achieving an 81.5% detection rate with zero false positives, this framework revolutionizes how IDS rules are generated from IoT traffic.
TianoForge slashes bug triage time from 11 days to just 7 minutes, revolutionizing efficiency in the TianoCore development community.
Move Statement refactorings can achieve up to 97% compilability, revealing that developer judgment is the main source of behavioral changes, not automation flaws.
Automating the entire medical imaging model-development pipeline can yield competitive baselines with minimal engineering effort, challenging traditional iterative approaches.
Execution signals and reasoning signals have orthogonal errors, and their independent calibration can significantly enhance Verilog code generation accuracy.
Achieving a staggering 95.5% on the ARC-AGI-3 benchmark, Prime Agent redefines the capabilities of long-horizon coding agents by seamlessly integrating memory and computation.
Achieving top-tier performance in complex professional tasks with a model significantly smaller than its competitors reveals a new frontier in agentic intelligence.
Slopsquatting risks are significantly mitigated with a two-layer detection system that achieves 76% hallucination-free code generation, even in adversarial contexts.
Only 2 out of 40 LLM-generated trading strategies actually beat a simple buy-and-hold approach, despite claims of superiority from the models themselves.
Achieving an 86.17% success rate in reproduction test generation, DPIAgent reveals that structured task separation can dramatically enhance performance in automated software engineering.
CloudEmu automates the creation of cloud emulators, outperforming a decade's worth of manual engineering with a single, efficient approach.
SPIDER4TianoCore achieves high-confidence patch-status classification without a single misclassification, setting a new standard for evidence generation in firmware development.
Iterative LLM feedback can significantly improve research software quality, revealing critical trade-offs that challenge conventional development practices.
Failed tool calls can increase the likelihood of repeating errors by over 800%, revealing a critical flaw in how language models process failure information.
LLM-SPICEMixer boosts circuit design efficiency, achieving up to 93.3% accuracy by combining LLM insights with traditional genetic algorithms.
Only 5.4% of coding agent attempts successfully complete a whole-repository migration while preserving behavioral correctness, revealing a critical gap in current AI capabilities.
MARS outperforms traditional multi-agent systems in competitive programming by harnessing specialized LLMs, achieving higher pass rates with lower costs.
AgentRoom slashes task abandonment rates in multi-agent coding by emphasizing coordination over mere parallel execution.
Aegis, trained with CyberFactory, outperforms existing models by 22.8 points in cybersecurity tasks, showcasing the power of agentic learning from real-world vulnerabilities.
Soft barriers can transform how we handle AI-generated code, making it harder to copy-paste without critical examination.
CodeMechanic transforms the landscape of automated vulnerability mitigation by prioritizing security over availability, effectively turning potential exploits into controlled terminations.
ReAct-SQL matches the accuracy of complex text-to-SQL systems while being up to 8 times faster and simpler.
ExecRubrics achieves up to 320 times faster evaluation while matching or exceeding the accuracy of traditional black-box judges in long-form response assessments.
Generative AI is not just reshaping visualization costs; it's redefining how we understand and trust software systems.
A novel Tree-of-Thought framework achieves a 62% success rate in repairing smart contracts, significantly outpacing traditional linear methods.
Rust's compile-time safety may not be enough; managed frameworks like Node.js and Django offer critical application-layer defenses that Rust lacks.
RAG-based defenses can significantly reduce package hallucination rates in LLM-generated code, but only if matched to the specific threat model and utility requirements.
Adjusting commit boundaries just got 57% faster with D-Diff, transforming how developers manage version control tasks.
LLMs exhibit a striking disconnect between functional correctness and code quality, with traditional metrics failing to capture the full spectrum of performance.
Switching from rigid graph representations to learned latent graph spectra can double F1 scores in code clone detection.
Go programs can now be verified directly with formal methods, enhancing educational tools and developer productivity without altering the original code structure.
Routing-guided exploration boosts software fix resolution rates by nearly 8%, outperforming traditional sampling methods without needing explicit answer forms.
Current coding agents can create playable games but fail to effectively diagnose bugs and maintain functionality during optimization, revealing critical limitations in their development capabilities.
PhysCaP enables robots to actively infer hidden physical properties, achieving superior manipulation efficiency without additional sensors.