Search papers, labs, and topics across Lattice.
93 papers published across 7 labs.
VirtualSet not only boosts accuracy in LLM-generated queries but also eliminates the risk of executing hallucinated actions by enforcing type safety before execution.
Urgent prompts may lead to less correct and secure code outputs from LLMs, revealing a critical flaw in common coding assistance strategies.
Shifting the focus from visual perception to direct program state access can boost agent success rates while slashing costs by nearly 9x.
Task dictates what concepts are represented in code models, but the model itself determines where and how those representations manifest, revealing a complex interplay in neural circuitry.
CudaPerf achieves up to 5X speedup in CUDA kernel generation by leveraging structural code-aware rewards that traditional methods overlook.
Urgent prompts may lead to less correct and secure code outputs from LLMs, revealing a critical flaw in common coding assistance strategies.
Shifting the focus from visual perception to direct program state access can boost agent success rates while slashing costs by nearly 9x.
Task dictates what concepts are represented in code models, but the model itself determines where and how those representations manifest, revealing a complex interplay in neural circuitry.
CudaPerf achieves up to 5X speedup in CUDA kernel generation by leveraging structural code-aware rewards that traditional methods overlook.
GRPO's hierarchical penalty reduces vulnerability prediction errors by nearly 28%, outperforming larger models under challenging conditions.
LLMs can now generate high-quality concurrent tests for Rust APIs, preserving semantic intent while exploring complex interleavings.
Open-weight LLMs can achieve nearly 88% task completion in complex data preparation tasks without ever sending sensitive data to the cloud.
Automated control design using LLMs can achieve a 26.5% improvement in performance metrics, revolutionizing how we approach multi-variable control systems.
Claude generated 58 logic procedures and proved their correctness, showcasing the power of LLMs in automating complex programming tasks with formal guarantees.
WML not only identifies failure mechanisms in agent workflows but also optimizes knowledge reuse, achieving state-of-the-art accuracy and efficiency in structured tasks.
Combining Transformers with LLMs boosts source code summary quality by 7.8%, addressing a crucial gap in secure software development.
LogMorph generates realistic Prolog bugs that reflect actual student mistakes, improving automated feedback tools in logic programming education.
Coding agents can now be evaluated on their ability to navigate fuzzy requirements and interactive workflows, reflecting real-world software development challenges.
Imp allows for precise handling of imprecise probabilities without disrupting established probabilistic programming frameworks, enabling richer modeling of uncertainty.
GenAI adoption in coding shifts the burden of maintenance from coding to extensive validation and dependency management, fundamentally altering developer workflows.
By merging CHR with miniKanren, chrKanren unlocks powerful new capabilities for relational programming, including advanced constraint solving and synthesis techniques.
The top-down approach to multiparty session types is just as powerful as the bottom-up method, challenging long-held assumptions about typability in distributed systems.
Real-world coding tasks, reverse-engineered from actual commits and scenarios, make Tencent WorkBuddy Bench a game-changer in contamination-resistant evaluation for coding agents.
Achieving a 73.5% reduction in design rule violations, EvoDRC revolutionizes the automation of DRC closure in advanced-node physical design.
Bridging Ciao assertions with LPTP theorems reveals new pathways for robust program verification by leveraging the strengths of both frameworks.
Current MLLMs excel at visual reproduction but falter in generating the necessary data semantics and interaction logic for coordinated multi-view interfaces.
Choreographic programming can eliminate deadlock and errors in quantum distributed systems by ensuring type safety and coherent protocol design.
AutoGlue achieves a 58.7% improvement in API F1 scores, demonstrating that LLMs can seamlessly translate natural-language requirements into executable code.
GenDB achieves significantly better performance than existing query engines by automating code generation tailored to specific workloads, fundamentally changing how we approach query processing.
Over two-thirds of malicious issue requests can exploit vulnerabilities in leading AI coding agents, highlighting a critical gap in current safety measures.
Every LLM-generated automation script analyzed contained exploitable vulnerabilities, regardless of the model used.
MoST achieves up to 717% performance improvement by leveraging diverse knowledge sources and cross-scenario strategies for code optimization.
Achieving over 90% bug detection in zkEVMs, VeriSynth transforms how we ensure the correctness of cryptographic implementations.
TRAVEL achieves a 26.22% boost in computational accuracy for C-to-Rust translation, setting a new standard for automated code migration.
Lax bug reproduction tests can lead to plausible but incorrect patches, but a new iterative framework boosts repair success by refining both tests and fixes.
LLM-based unit tests can achieve reliable compilation and coverage improvements by explicitly managing project context rather than relying solely on prompt engineering.
PerfAgent doubles the rate of expert-level code optimizations by leveraging profiler-guided feedback, revealing hidden performance bottlenecks that traditional methods miss.
GCP achieves zero-annotation type inference by reconciling four distinct evidence sources, transforming how we approach type safety in dynamic languages.
NOOA revolutionizes agent development by treating agents as first-class Python objects, enabling seamless integration of AI capabilities with standard programming practices.
The Spaghetti Architect generates high-quality, contamination-resistant code datasets that can dramatically enhance model training and evaluation across multiple programming languages.
LangGraph reveals that the right orchestration framework can significantly enhance the reliability and efficiency of stateful AI systems in complex business workflows.
Despite apparent progress in SVG repairs, top models struggle to meet rigorous editing specifications, achieving only 15% success.
Current LLMs achieve only a 12.30% success rate in generating executable scientific code, highlighting a significant gap in their capabilities.
PhoenixRepair redefines how software agents explore repair strategies, achieving a 76% resolution rate by leveraging multi-location sampling and iterative refinement.
AgentTrails uncovers hidden dependencies in agent trajectories, enabling unprecedented insights into LLM-powered task execution.
Compiled pipelines using a novel visibility control abstraction outperform traditional HLS tools while matching the efficiency of hand-written RTL designs.
Recovery routing can outperform escalation strategies by leveraging execution feedback, achieving a higher solve rate at only 35% of the typical recovery cost.
H$^2$SD achieves superior reasoning performance by intelligently adapting teacher signals based on trajectory outcomes, leading to more effective learning in large language models.
Learning dynamic representation programs can boost retrieval performance by over 30% compared to traditional static methods.
Authority framing can lead verifiers to overlook critical security flaws, allowing 80% of malicious code to bypass scrutiny in CI/CD pipelines.
PTSan slashes memory safety overhead to just 57.2% while preserving robust detection capabilities, making pointer-based sanitization viable for production use.
LISA uncovers functional bugs that traditional testing methods miss, achieving superior detection rates without relying on crashes.
TraceDev outperforms existing methods by over 340% in transforming complex requirements into executable code while ensuring traceability.
Dedicated code embeddings can predict functional correctness in Scratch programs with surprising accuracy, even in small classroom settings.
Skillware redefines agent skills as independent software artifacts, unlocking new possibilities for their management and evolution in AI systems.
LLM agents can achieve cross-file code changes with one to two orders of magnitude fewer tokens by leveraging algebraic operations instead of plain text edits.
VirtualSet not only boosts accuracy in LLM-generated queries but also eliminates the risk of executing hallucinated actions by enforcing type safety before execution.
A novel compiler boundary design ensures that only authorized rewrites are executed, significantly enhancing security in opaque call handling.
Pruning tool outputs directly within the agent leads to a remarkable 39% reduction in token usage without sacrificing performance.
Expert-guided LLMs can achieve up to 29.68x speedups in GPU kernel performance, highlighting the critical role of human knowledge in AI-driven optimization.
Learning $\max$@$k$-optimal policies is statistically harder than traditional reinforcement learning, revealing a critical gap in agent evaluation methods.
LLMs can autonomously generate and refine simulation models, revealing complex relationships in data that traditional methods might overlook.
SGA boosts the visual quality of LLM-generated educational animations by over 16% by tackling geometric occlusions that others ignore.
Achieving high performance in Approximate Nearest Neighbor Search has never been easier—ANNLib allows for flexible configurations that outperform existing systems with minimal coding effort.
VNVSpec transforms high-level user requirements into actionable, machine-readable specifications that can be directly linked to test results, addressing a critical gap in software verification.
PSASpotter reveals hidden risks in Python code by detecting platform-specific APIs and their defensive usage, paving the way for safer cross-platform development.
KernelDiag transforms kernel crash diagnosis by leveraging structured causal reasoning to achieve up to 4x improvements in fault localization accuracy.
A unified framework reveals that robust downgrading mechanisms can coexist with full-strength non-interference, transforming our approach to program security.
C2Btor outperforms traditional verification tools by leveraging hardware model checking, solving 101 more tasks than CBMC in a rigorous benchmark evaluation.
CommitLLM transforms uninformative git commit messages into concise, compliant formats, achieving 98% adherence and drastically reducing message length.
The because-calculus eliminates vacuous bindings at compile-time, ensuring that non-resumable operations are distinctly handled without compromising type safety.
Stopping repairs too soon can lead to a 60% drop in true validity, but VRR-Stop ensures LLM agents know exactly when to commit or repair.
Insecure coding preferences in LLM long-term memory can elevate vulnerability rates by over 50%, posing a critical security risk in code generation.
Ghost references in LLM-generated code can be eliminated by constraining generation to valid runtime environments, ensuring both grammatical and semantic correctness.
RECEIPT uncovers 24 previously unknown XSS vulnerabilities while ensuring no false positives, setting a new standard for trust in automated vulnerability discovery.
Nearly half of agent-generated pull requests lack adequate test coverage, exposing critical vulnerabilities in autonomous software development.
Later models may resolve more coding tasks, but they don't necessarily produce better quality patches in terms of performance metrics.
CODENS turns pull requests into a living knowledge graph, making code documentation both dynamic and easily queryable.
TRIM effectively reduces AI-generated code redundancy by up to 32.9%, preserving performance while tackling the growing problem of CodeSlop in software development.
Chiral analysis reveals vulnerabilities in smart contracts that traditional methods miss by focusing on relational inconsistencies between business paths.
Structured evidence from upstream sources boosts LLM repair accuracy by up to 23 percentage points, revolutionizing how we adapt to breaking changes in software dependencies.
Weak non-negativity in supermartingales can boost verification success rates by over 20% in probabilistic programs, challenging the notion that strict constraints are necessary for soundness.
Current higher-order unification methods miss key functions defined by case analysis, but integrating dependent pattern matching reveals solutions that existing systems overlook.
Behavioral alignment in bug resolution is quantifiable and reveals that structured signals can significantly enhance the reliability of test generation and patch ranking.
Codex's performance plummets by 50% in long contexts, revealing critical vulnerabilities in agent skills during code audits.
Coding agents can exploit evaluation frameworks in surprising ways, revealing a stark contrast in generalization versus memorization strategies that impacts real-world applications.
Misinterpretations in RTL design generation can be fixed before code is even written, leading to a groundbreaking 94% functional correctness in generated designs.
Biographical personas can drastically alter code generation outcomes, with some leading to reduced correctness and others to inefficient verbosity.
Quality-aware code search can significantly boost developer productivity by prioritizing high-quality code that meets specific resource optimization needs.
GenAI is not just automating tasks; it’s fundamentally reshaping the career trajectories of junior software engineers, risking the loss of essential learning experiences.
Test suites can dramatically enhance issue localization accuracy, bridging the gap between abstract descriptions and concrete code.
A scale-dependent quality-quantity trade-off reveals that doubling trajectory data can significantly enhance model performance, but only up to a point.
Prefactory uncovers 75 library-adoption opportunities in Python code, outperforming existing tools and generating 40 validated refactorings through innovative LLM-driven heuristics.
Reusing verified debugging records, MechMem-RTL resolves 62.5% of complex RTL errors, outperforming existing methods that rely on text similarity.
A portable inlining predictor can achieve near-state-of-the-art performance in compiler optimizations without relying on heavyweight infrastructure.
Tailoring LLM code explanations to individual problem-solving styles can dramatically enhance user productivity and learning in coding tasks.
Indirect function calls in WebAssembly can be effectively managed by filtering candidates, drastically simplifying the verification process.
Bridging the NL2Pipeline gap, DataFlow-Harness enables LLMs to create reliable, editable data workflows at a fraction of the cost and time of traditional methods.