Search papers, labs, and topics across Lattice.
100 papers published across 4 labs.
Interactive type highlighting can drastically reduce debugging time by clarifying the derivation of static types and dynamic casts in real-time.
Bridging the NL2Pipeline gap, DataFlow-Harness enables LLMs to create reliable, editable data workflows at a fraction of the cost and time of traditional methods.
LQCDMaster automates complex LQCD workflows, achieving expert-level precision while slashing computation time from hours to minutes.
Unlocking the potential of unused fine-tuning data can enhance reasoning models' performance in complex domains without sacrificing their core capabilities.
Reward-driven optimization in continuous-time RL can significantly enhance the fine-tuning of discrete diffusion models, even with non-differentiable reward signals.
Bridging the NL2Pipeline gap, DataFlow-Harness enables LLMs to create reliable, editable data workflows at a fraction of the cost and time of traditional methods.
LQCDMaster automates complex LQCD workflows, achieving expert-level precision while slashing computation time from hours to minutes.
Unlocking the potential of unused fine-tuning data can enhance reasoning models' performance in complex domains without sacrificing their core capabilities.
Reward-driven optimization in continuous-time RL can significantly enhance the fine-tuning of discrete diffusion models, even with non-differentiable reward signals.
LLMs can effectively assist in building complex solvers from academic literature, but human oversight remains crucial for validation and performance optimization.
Execution-verified programmatic distillation allows smaller models to outperform larger ones in financial reasoning tasks, achieving remarkable accuracy through reliable numerical computation.
Automating the synthesis of leakage contracts could revolutionize CPU security by eliminating the need for extensive manual effort in their development.
LEAF transforms Rust program analysis by providing a dynamic framework that captures rich semantic information at runtime, enabling advanced analysis tasks.
Specialist agents can outperform generalist LLMs by up to 20 percentage points in execution accuracy while slashing costs and errors, making them essential for reliable software development.
AI-driven interviews achieved a 90.9% positive experience rating, but raised concerns about depth and privacy that could redefine ESE methodologies.
Trusting evidence over agent claims led to zero false-DONE outcomes and a significant reduction in hidden-fail amplification in autonomous coding workflows.
English isn't always the best choice for generating high-quality code—language bias significantly impacts LLM performance across programming tasks.
Artifact-centered evaluations reveal that LLM agents can achieve an 88.6% success rate in structural engineering tasks, but still struggle with invalid inputs and model consistency.
Coding agents can significantly improve payment integration performance with targeted skills, achieving up to 91.37% success in complex scenarios.
Structural priors can boost semantic vulnerability recall to 100% in synthetic settings but collapse to just 48.9% on real-world data, revealing a critical cross-domain challenge in LLM performance.
Eye-tracking data reveals that our model predicts programmer attention with unprecedented accuracy, outperforming existing AI models by significant margins.
FirmPilot doubles the web-service reachability of firmware rehosting, transforming how we analyze IoT devices in emulated environments.
SYNAPSE achieves a remarkable balance of accessibility and engagement, scoring 76.4 on usability while effectively teaching secure software development to both neurodivergent and neurotypical learners.
Pattern-guided exploration can slash FPGA design search spaces by over 80% while preserving optimal performance.
Navigating the chaotic landscape of quantum computing just got easier with a new framework that balances technology pull and normative goals.
Nondeterministic choices in choreographic programming can now be accurately mechanised, ensuring robust concurrency without sacrificing expressiveness.
Capturing design pattern variability could revolutionize mobile app generation, ensuring both customization and architectural integrity.
Embedding cotangent fibers in a common ambient type reveals a surprising equivalence between dependent and simply typed semantics in reverse-mode automatic differentiation.
Attackers can weaponize ordinary project documentation to compromise AI coding agents, exposing a critical security gap in how these systems handle package installations.
Interactive type highlighting can drastically reduce debugging time by clarifying the derivation of static types and dynamic casts in real-time.
DREA not only boosts vulnerability detection accuracy but also slashes API costs by up to 48%, revealing a hidden flaw in LLM reasoning that could reshape security assessments.
GFlowRL achieves unprecedented stability and performance in large language models by eliminating the problematic learned partition function, setting a new standard for GFlowNet-style reinforcement learning.
Despite the promise of agentic coding tools, most GitHub projects see minimal adoption, with only a few exceeding the industry standard for PRs per participant.
AI agents can autonomously verify security software with astonishing efficiency, but their trustworthiness is limited by the strength of their feedback mechanisms.
ProfMalPlus not only achieves a 98.1% F1-score but also uncovers hundreds of previously undetected malicious packages, demonstrating a significant leap in supply-chain security for open-source software.
Git's version control could revolutionize memory management in coding agents, achieving 60x better retrieval performance than traditional methods.
VisualRepair resolves 196 software issues by intelligently focusing on relevant visual regions, outperforming existing methods and showcasing the power of multimodal understanding in automated repair.
AQLM outperforms full-precision baselines in code generation, while QuIP# falters under complex prompts, revealing critical trade-offs in quantization strategies.
LLMs may excel in code generation, but they frequently falter on complex tasks and basic errors, revealing critical reliability gaps in automated coding solutions.
NexForge transforms the landscape of LLM training by synthesizing 43.2K tasks, propelling model performance to unprecedented levels without the need for domain-specific infrastructure.
Differentiating probabilistic programs reveals that cotangents must navigate both deterministic and probabilistic structures, challenging traditional views on automatic differentiation.
ATLAS transforms LLMs from unreliable code generators to trusted partners in analog design, successfully producing SAR ADCs that meet rigorous simulation standards.
Structured feedback can boost LLM agent success rates by up to 44 percentage points, revealing the critical role of admissible alternatives in the repair process.
Generative compilation enables AI models to receive compiler feedback during code generation, drastically reducing errors and improving code quality in real-time.
Achieving a 76.42% compilation success rate, Chat2Scenic revolutionizes scenario generation for autonomous driving by effectively bridging regulatory language and executable scripts.
Autoregressive drift in transformer models leads to a dramatic drop in exact equivalence for complex quantum circuits, revealing a fundamental limitation in current synthesis methods.
LLM-based evaluators can outperform traditional metrics in method name prediction, but the real breakthrough comes from a novel approach that enhances name quality through summarization and refinement.
CodeOwl can automatically generate effective tiered programming problems, achieving a 98.7% success rate in complexity progression, but educators want more control over curriculum integration.
Antiproof uncovers hundreds of previously unknown vulnerabilities, including critical zero-days that could compromise LLM training and inference systems.
Alerus bridges the gap between formal verification and practical probabilistic programming in Rust, enabling the verification of complex sampling algorithms that were previously unmanageable.
LLM-generated bug reports often hinge on implicit assumptions, and this framework reveals how to validate their correctness through a novel witness-generation approach.
VIZDETOUR reveals that subtle rendering bugs can be detected by exploiting the equivalence of API calls, leading to the discovery of 47 previously unknown issues in popular visualization libraries.
AI Literacy and Development are now essential competencies that universities must integrate into software engineering curricula to meet industry demands.
Mystra achieves 95.5% recall on dynamic taint analysis with zero false positives while significantly reducing runtime overhead compared to existing tools.
LLMs not only differ significantly from human developers in security practices but also struggle with effectively repairing vulnerabilities, raising concerns about their reliability in security-sensitive applications.
Monty achieves up to 20 points higher precision in generating executable assertions from natural language compared to naive LLM approaches, revolutionizing software verification processes.
Cycle intersections in Cayley graphs can efficiently solve the TopSpin puzzle, revealing new pathways in permutation puzzle research.
Task-aware execution can cut costs by 85% while maintaining a 100% success rate in complex workflows.
Mid-training with function-aware fill-in-the-middle boosts coding agent performance while preventing capability erosion in non-agentic tasks.
Code-MUE reveals a striking -0.98 correlation between uncertainty and functional correctness in Code LLMs, offering a new lens for assessing model reliability.
Nearly 40% of agent-generated pull requests harbor security vulnerabilities, with human collaborators unintentionally introducing the majority of critical leaks.
Shifting from external dependencies to local code synthesis can preserve functionality while slashing the attack surface by 93%.
Error information in small frozen code LLMs may not enhance self-repair capabilities as previously assumed, challenging existing notions in the field.
Line-anchored feedback can slash token generation costs by up to 58% while simultaneously improving correctness in AI code editing tasks.
Many software engineering studies rely on fewer than 12 interviews, raising questions about the rigor of qualitative research in the field.
Customized inference engines can now be generated automatically from user-defined runtime constraints, drastically simplifying the development process.
Multi-perspective reasoning in CT-Repair leads to a 99-bug improvement over the best individual analysis strategy, showcasing the power of structured evidence in automated program repair.
Cadre repairs Dockerfile drift with a 35.22% success rate, outperforming existing methods by leveraging context-aware dependency modeling.
FLEX achieves a remarkable 95.7% automatic discharge rate of CHCs, redefining trust and expressiveness in program verification.
The integration of AI into parallel programming could revolutionize how we develop and optimize multi-core and distributed systems.
Effect handlers not only simplify the development of program logics but also yield stronger reasoning rules than traditional methods, revolutionizing how we approach program effects.
Quantum programs can be analyzed for expected runtime without needing upper bounds, transforming how we approach verification in this domain.
Executable JavaScript can match the fidelity of formal specifications, but understanding complex protocols remains a critical hurdle for large language models.
SemaDiff achieves 100% precision in detecting semantic-changing commits, revolutionizing how we assess software modifications.
AI-assisted development tools can slash delivery times by up to 69% while enhancing design fidelity in front-end workflows.
AI-driven analysis can drastically cut down curriculum revision time, potentially transforming graduation rates for Software Engineering students.
Behavior localization is revolutionized, enabling developers to seamlessly connect high-level modification requests to specific code locations in complex AI harnesses.
Agent-involved code reviews speed up decision-making but compromise on quality, challenging the assumption that faster reviews equate to better outcomes.
Every type query can be answered with a minimal program slice, transforming how developers understand and debug type information in their code.
Centaur's predictions align more closely with human responses in program comprehension than traditional models, revealing the power of cognitive insights in software engineering.
LLMs can now certify bug reports with machine-checked proofs, drastically cutting down on false alarms in software development.
Reducing output tokens doesn't guarantee lower costs; in fact, it can lead to higher bills and lower task success rates.
Mechanism-guided synthetic bugs can dramatically enhance LLM-based unit test generation, leading to superior real-bug detection.
ThinkLog achieves a 20.55% accuracy in log statement generation, marking a significant leap over previous methods while slashing inference costs in half.
Tool grounding in LLMs can boost IaC repair accuracy from 26.6% to 78.4%, revealing a path to more reliable automated cloud configuration management.
Efficient generation of test cases is fundamentally limited by complexity classes, revealing that not all languages in P can be efficiently generated.
Transforming GUI test code into app-specific voice assistants could revolutionize how developers create voice interactions, slashing costs and improving functionality.
Despite improving functional correctness, AI code assistants like Copilot leave developers blind to security vulnerabilities in their API usage.
FOCAL outperforms traditional methods by significantly improving failure detection on unseen projects while providing rich behavioral explanations.
Existing test migration methods falter in OpenHarmony, but a new approach boosts success rates from 26% to 81% by leveraging system-specific insights.
RepTran achieves a remarkable 74.7% repair rate for Transformer models, significantly outperforming existing repair methods.
Role-aware summaries can boost bug localization effectiveness by 40% while being 10.4 to 20.9 times more efficient than raw source code.
ToFu achieves superior token efficiency and cost-effectiveness in agentic coding, all while empowering researchers with a transparent, modifiable framework.
FlowArk cuts data-flow analysis costs by over a quarter while boosting task completion rates, making it a game-changer for Android app security auditing.
A staggering 47.57% of rewarded outputs in RLVR systems correspond to genuine bugs, challenging the reliability of current reward suites.
ProgramTab outperforms all LLM-based baselines in table reasoning by transforming unstructured data into actionable insights through programmatic preprocessing.
Structured project packaging with FORAP reduces instructor workload and enhances student engagement in computing education, making PjBL more accessible than ever.
Curriculum fine-tuning can significantly enhance the success rate of LLMs in neural architecture synthesis, but distinct failure modes require different repair strategies.
HierCAD achieves unprecedented fidelity in CAD generation by aligning structural reasoning with geometric parameters, setting a new benchmark for text-to-CAD systems.
A staggering 68.8% of LLM-generated code snippets contain security vulnerabilities, revealing a complex web of interrelated flaws that challenge traditional isolation strategies in software security.
ReviewDSE achieves a 1.78% reduction in wirelength while exposing and repairing design flaws that traditional methods miss.
AHA reveals a reusable vulnerability core across production LLM agents, significantly enhancing the efficiency of red-teaming efforts.
ACQUIRE turns knowledge gaps into actionable insights, boosting software repair accuracy by over 4% while keeping costs low.
LLMs can generate backend services that pass initial tests, but struggle significantly under comprehensive evaluation, revealing a critical gap in their capabilities.
Achieving logarithmic complexity for certain assertion checks in quantum programs could revolutionize how we approach debugging in quantum computing.