Search papers, labs, and topics across Lattice.
100 papers published across 9 labs.
einx transforms tensor programming by providing a universal, readable notation that eliminates shape errors and simplifies complex operations.
Even the top-performing LLM struggles with cross-file reasoning, achieving only 69.1% accuracy on a new benchmark designed to reflect real-world software development challenges.
CADENA reconstructs 3D models feature by feature, achieving unprecedented accuracy in CAD reverse engineering.
Verified symmetry breaking in LeanCSP can reduce solver search efforts by up to 20 million times, revolutionizing trust in constraint programming results.
Iterative reformulation using LLMs can dramatically accelerate solver efficiency by leveraging diverse context retention strategies.
CADENA reconstructs 3D models feature by feature, achieving unprecedented accuracy in CAD reverse engineering.
einx transforms tensor programming by providing a universal, readable notation that eliminates shape errors and simplifies complex operations.
Verified symmetry breaking in LeanCSP can reduce solver search efforts by up to 20 million times, revolutionizing trust in constraint programming results.
Iterative reformulation using LLMs can dramatically accelerate solver efficiency by leveraging diverse context retention strategies.
Students leverage AI not for automation but as a critical support tool, enhancing their grasp of requirements quality in software engineering.
Typed local edits in BlueprintRepair are not only the most efficient method for fixing proof failures but also achieve near-complete coverage within a constrained token budget.
Selective correction boosts Text-to-SPARQL accuracy to 98.33% while cutting inference time by nearly half.
Syntropy achieves up to 99.5% validity in generating deadlock-free protocol refinements, revolutionizing how we ensure correctness in distributed systems.
LLMs can resolve Java merge conflicts with 100% precision, significantly outperforming traditional tools by handling a wider array of cases.
Confidence gating in co-decoding can lead to a substantial 12.6% increase in secure code generation, revealing the critical role of expert model confidence in safety.
Structural evaluations of LLM-generated microservices reveal that apparent differences in prompting strategies may mask underlying methodological biases rather than true architectural quality.
SetFit outshines traditional ML and advanced language models, achieving an F1-score of 0.80 in identifying security-related bug reports.
CHARGE uncovers new vulnerabilities in hardware designs while automating the generation of security properties, significantly reducing manual effort.
LLMs can outperform traditional trading strategies in parent-order execution, revealing a new dimension of AI's role in finance.
Change2Task recovers 29.2% more verified coding tasks than traditional methods, streamlining the training of coding agents.
Language model agents struggle with oncall RCA, achieving only 25.3% accuracy on realistic tasks, revealing a critical readiness gap for production environments.
Few-shot prompting boosts LLM performance in generating microservice architectures, achieving an impressive F1 score of 0.97 for service identification.
The Locksmith Loop can achieve nearly complete test coverage for legacy code migrations, ensuring that the migrated Java code functions identically to its COBOL predecessor.
Outperforming GPT-5.4, IndustryForge-27B achieves a remarkable 33.65 percentage point lift in CAD-specific tasks, setting a new standard for multimodal models in industrial applications.
Proof assistants could be the key to unlocking robust security certifications across diverse domains like cryptography and secure compilation.
Vibe modeling could be the key to bridging the gap between natural language prompts and trustworthy software generation, enhancing both understanding and validation.
LimICE not only solves more loop invariant problems but does so significantly faster, redefining efficiency benchmarks in program verification.
ARES achieves up to a 27% reduction in optimization costs by intelligently adapting reasoning effort based on progress, outperforming fixed-effort approaches.
LLMs struggle with code deletion, often opting for workarounds that compromise code maintainability, revealing a significant training gap in their editing capabilities.
Fixed-position models struggle with any-order inference, but new masked diffusion techniques unlock flexible generation capabilities that enhance performance in coding and reasoning tasks.
Achieving up to 6.65× speedup in NPU inference by automating Ascend C operator generation could revolutionize performance optimization in low-corpus environments.
Leveraging user history can cut clarification requests in coding assistants by identifying and resolving recurring ambiguities, leading to more efficient coding sessions.
EvoPINN autonomously discovers new algorithms for physics-informed neural networks, achieving significant performance improvements while ensuring scientific validity.
Coding agents may boost productivity, but they risk diminishing developers' understanding and long-term coding skills.
Willow's innovative type-and-effect system reveals hidden timing dependencies in reactive programs, enabling static detection of performance-degrading render cascades.
VITAL-RAG boosts code retrieval efficiency by over 60% while slashing token usage, reshaping how coding agents manage context.
MRCoder slashes token consumption by up to 50% while boosting code generation accuracy, redefining efficiency in repository-level code generation.
MultiFixer repairs 420 bugs, including complex multi-hunk cases, establishing a new benchmark in Automated Program Repair.
Developers trust human-written code tour descriptions significantly more than those generated by LLMs, revealing critical gaps in AI-generated documentation. WHY_IT MATTERS: This research could reshape the design of AI-assisted debugging tools, ensuring they align better with developer preferences and trust dynamics.
A novel dataset of untangled commits sourced from PRs reveals that traditional heuristic methods may overlook critical distinctions that affect machine learning performance.
MalGuard reveals that capturing cohesive program behaviors through operational roles can dramatically enhance malware detection accuracy in organizational settings.
A minimal LLM-based analyzer can rediscover 68% of AI-discovered CVEs, while frontier models fail to detect any, exposing critical gaps in current vulnerability detection methods.
A composite attack can breach a supposedly robust self-check defense, achieving up to 67% success where individual methods fail.
Lower bounds for 1-query unitary synthesis reveal surprising complexities in synthesizing seemingly simple quantum operations, challenging existing assumptions in quantum circuit design.
Semantic code retrieval can outperform traditional methods by 36% in precision, especially in name-obfuscated tasks where lexical approaches fail.
Compiled harnesses in SIGIL boost execution of mandated steps to 86%, outperforming prose skills by a staggering 30%.
Coding agents struggle with non-functional improvements, scoring as low as 1.3 on structural changes compared to human developers' 1.5.
RLPF transforms how code generation models are trained by prioritizing runtime efficiency alongside correctness, leading to a dramatic increase in both runnable solutions and execution speed.
PROGRESS exposes 58% of existing bugs in large-scale Java systems, far surpassing traditional regression testing methods that fail to detect any.
Developers are turning to external safeguards due to deep-seated mistrust in LIDEs, revealing systemic vulnerabilities that could compromise security and privacy.
CodeSpec transforms feature development by ensuring that LLM-based code agents produce reliable and verifiable functional chains, achieving up to 70.7% pass rates on complex benchmarks.
A new optimization-based approach to inductive invariant synthesis achieves an 86% increase in benchmark solvability compared to traditional methods.
Explanation quality is a critical yet overlooked dimension of LLM agent performance, with many agents generating misleading explanations that can lead to incorrect code assessments.
Baseline choice can radically alter the diagnosis of numerical deviations, revealing that compiler-induced discrepancies don't always equate to accuracy loss.
Fine-tuning a small language model with MindForge's innovative training environments boosts its performance to rival much larger models in software engineering tasks.
SpecFirst boosts program synthesis success rates by up to 21.3% by prioritizing behavioral specification before coding.
Historical data can be dynamically adapted to achieve an impressive 1.732x speedup in compiler optimization, reshaping how we approach autotuning.
Attention misallocation in LLMs is a critical factor behind their inconsistent performance in Automated Program Repair, with successful outcomes linked to broader attention across bug report components.
SHarD can embed and distribute robust security controls for AI coding agents with a single command, achieving perfect security efficacy without compromising performance.
Achieving 59.6% accuracy with LoRA rank 16 while training less than 1% of parameters showcases the potential for extreme efficiency in fine-tuning language models.
SkillGate reduces the risk of malicious skill files in AI coding agents by achieving an impressive F1 score of 0.817 while slashing LLM input requirements by 77%.
Misaligned training pairs in code review datasets can undermine the effectiveness of LLMs, revealing that dataset cleaning alone won't solve the problem of generating actionable feedback.
Bottom-up enumeration in miniKanren can now achieve superior performance in program synthesis by leveraging pruning and memoization techniques.
Execution time can be transformed into a learnable reward, leading to substantial improvements in code optimization performance in RL settings.
Uncovering 23 previously unknown frontend bugs in deep learning compilers could significantly enhance the reliability of AI frameworks like PyTorch.
Specula uncovers deep bugs in system code that traditional methods often miss, revolutionizing formal specification generation.
Incorporating live service-call graphs into LLM-generated patches boosts correctness rates for Kubernetes security fixes from 11.1% to 78.0%, revealing a crucial oversight in current remediation approaches.
Code-mixing fingerprints can achieve robust ownership verification in LLMs without sacrificing performance, avoiding accidental activations that plague traditional methods.
Achieving 98.1% accuracy in IUPAC name generation could revolutionize how chemists and researchers communicate molecular structures.
A structured contract can boost HLS testbench pass rates by over 6% while integrating hardware feedback for real-time design improvements.
Industry-specific prompts may appear to enhance code security, but the real determinant of vulnerability rates is the choice of model itself.
Code models may leverage tests more for semantic guidance than as executable specifications, with surprising implications for their performance consistency.
ARCHER enables self-hosted models to achieve 97.8% of frontier-API accuracy at just a quarter of the cost, revolutionizing compliance checking in building regulations.
NFR traceability is not just harder than FR traceability; it reveals a critical gap in how security-related requirements are implemented in code.
Even the top-performing LLM struggles with cross-file reasoning, achieving only 69.1% accuracy on a new benchmark designed to reflect real-world software development challenges.
KQFuzz uncovers bugs in quantum libraries with 18.44% better coverage than existing methods, revealing critical flaws that developers can quickly address.
Nonlinear cost structures in probabilistic programs can be analyzed for moments with a new method that balances accuracy and computational efficiency.
Blind resampling reduces token usage by up to 5.5 times and outperforms self-repair strategies in small code models, challenging conventional wisdom about model feedback.
A unified type-safety framework that simplifies complex type systems into a single logic, ensuring well-typed programs never abort.
Julia can achieve impressive parallel scaling on HPC tasks, rivaling traditional Fortran implementations, but faces unique challenges that need addressing.
Achieving 100% semantic accuracy in OSC command generation while eliminating wrong-send errors could redefine reliability standards in live performance technology.
UML diagrams dominate LLM applications in software engineering, but critical gaps in behavioral modeling and evaluation practices could hinder progress.
Foundational mechanized proofs can now be produced at scale using LLMs, transforming the economics of formal verification in programming languages.
TraceCoder reveals the hidden narrative behind code generation, linking 30% of snippets to specific repair events and transforming how we audit AI-generated code.
Coding agents can now access repository context 8.7x faster, dramatically improving efficiency in navigating evolving codebases.
LLM-generated shuttling compilers can cut development time from months to days while achieving superior performance compared to traditional hand-crafted solutions.
Trust in AI-assisted code reviews can be paradoxically undermined by too much explanation, as developers may question recommendations more when provided with detailed reasoning.
Runtime procedure injection can boost model performance by nearly 7 points, revealing a critical gap between abstraction and execution in scientific computing.
Sketch2DES turns complex queuing network diagrams into verifiable simulation models, making advanced simulation accessible even to non-programmers.
FlowLog allows users to seamlessly switch between one-shot and incremental evaluations of static analyses, achieving updates in milliseconds while tuning performance on-the-fly.
Capability classifiers enable precise control over resource tracking in Scala 3, addressing previously inexpressible constraints that hindered library functionality.
KernelScript can catch cross-boundary bugs at compile time that traditional methods miss, enhancing the reliability of eBPF applications.
Achieving performance parity with conventional miniKanren unification, this new algorithm redefines efficiency in rational term unification for persistent settings.
Most developers edit AI-generated code within 15 minutes, often discarding the original completions entirely, highlighting a critical gap in LLM training data.
Uncertainty in retrieval can be harnessed to boost code generation accuracy by over 20% when properly managed.
Repeated revisions in coding agents can lead to a significant drop in reliability, highlighting the need for structured evidence-bound contracts in code repair processes.
LLM-based compression techniques can achieve up to 82% better compression on source code compared to traditional methods, revealing hidden efficiencies in code representation.
Typos and camel-case abbreviations no longer hinder code completion accuracy in Pharo, thanks to innovative extensions that enhance user experience without sacrificing performance.
Locality in rank-metric codes can now achieve efficient recovery of any support element, challenging previous assumptions about dependence on basis choices.
RedPhuzz detects all vulnerabilities in web applications while being 73% faster than its predecessor, Phuzz, which missed 16 critical vulnerabilities.
Adversarial comments can bypass LLM-based vulnerability detectors with over 90% success, exposing a critical vulnerability in AI-driven security tools.
Skills for large language model agents can be engineered like software, ensuring better performance and reliability through structured design principles.
LLMs can generate operational Docker configurations, but they often miss crucial deployment intents, revealing a significant gap in their utility for production environments.
NL2Test transforms raw execution traffic into reliable API regression tests, achieving an 82.4% exact-match rate while slashing manual testing efforts.
Structurally-challenging functions in JavaScript are not random; they cluster in specific files and emerge from diverse patterns, challenging conventional testing assumptions.