Search papers, labs, and topics across Lattice.

Leading Asian AI research university. Active across NLP, computer vision, and multimodal learning.
100
0
0
This work proposes OPTED (on-policy fine-tuning for end-to-end driving) which decouples reinforcement learning from the post-training of the end-to-end policy: a privileged teacher is trained using RL on vectorized inputs (HD-map and bounding boxes) which provides supervision to the pre-trained student during closed-loop post-training.
This work presents a preprocessing technique that improves the efficiency of ZkUnsat without introducing additional leakage, and normalizes the proof so that each derived clause is justified by a resolution chain of fixed public length k, which eliminates chain-length leakage and reduces prover memory usage.
Benchmark-average rankings mask massive sample-level disagreement among vision pruning techniques: dynamically routing inputs across existing pruning methods yields a 26.9% relative accuracy jump over any single fixed strategy.
Scaling generalized zero-shot recognition to massive vocabularies does not require re-architecting backbones: treating seen-versus-unseen gating as a downstream Monte Carlo calibration problem boosts unseen accuracy by over 20% in just 15 dimensions.
Interactive communication yields zero minimax advantage in 1-bit distributed mean estimation: purely non-adaptive queries match the adaptive rate for any heavy-tailed moment condition $k > 1$.
Achieving an 11% boost in energy efficiency for edge AI applications by smartly balancing throughput and latency through hierarchical operator parallelism.
Transforming scientific papers into editable vector figures, \textsc{FigTree} makes figure generation and refinement as seamless as writing.
Iterative problem-solving reveals that LLMs, while less accurate, can generate insightful solutions that illuminate failure modes in student coding attempts.
As human oversight wanes, LRMs could autonomously evolve, but this shift introduces significant risks like reward hacking and feedback drift.
RailSyn achieves up to 4.9 points improvement in RFOD detection accuracy by intelligently generating synthetic data that addresses specific task deficiencies.
Reducio slashes memory requirements for serverless cloud deployments while enhancing privacy, all without the need for costly infrastructure overhauls.
PriMD revolutionizes multimodal emotion recognition by enabling robust performance even when critical data modalities are missing.
Missing spatial constraints in CAD generation can be effectively filled using past design experiences, significantly enhancing the accuracy of text-to-CAD translation.
Quantum kernels can significantly boost fraud detection accuracy by aligning feature-map geometry with interaction-sensitive decision boundaries, outperforming traditional methods.
Achieving up to a 72.81脳 speedup in long-sequence inference could redefine efficiency standards for Diffusion Language Models.
Agents using ParallelWorld can efficiently evaluate multiple future trajectories, leading to superior decision-making in complex environments.
RecVerse outperforms traditional simulators by maintaining a nuanced understanding of user intent and memory, leading to more realistic shopping behaviors.
SPACE recalibrates multivariate forecasting by leveraging current sample clouds, resulting in improved coverage efficiency that outperforms traditional methods.
Internal visual modality entropy can serve as a reliable self-evaluation metric for action generation in heterogeneous VLA architectures, outperforming traditional methods.
MSCNet not only reconstructs missing MRI sequences but also enhances lesion fidelity, outperforming traditional methods in clinical settings.
Cheap auxiliary signals can provide unbiased conditional performance estimates for language models, revealing hidden strengths and weaknesses that gold labels alone miss.
Current AI-generated videos that mislead viewers are also the most challenging for existing detection systems to identify, revealing a critical vulnerability in misinformation defenses.
Making the teacher's privileged context learnable end-to-end enables agents to evolve more efficiently, outperforming traditional methods with less than 30% of their rollout budget.
SafeCA slashes jailbreak success rates by 20% while adding virtually no latency, revolutionizing defenses for text-to-video models.
StreamFlow achieves a remarkable 67.73% accuracy in streaming video understanding while cutting latency and memory usage by over 50%.
Forgetting in hyperbolic continual learning is driven by semantic drift and hierarchical distortion, revealing critical insights for preserving multimodal representations.
Achieving state-of-the-art performance in both 4D reconstruction and point tracking, Uni4R leverages continuous velocity fields to model dynamics at any timestamp, breaking free from traditional limitations.
ST-Omni-R1 not only excels in sound-event recognition but also sets a new standard for spatial audio reasoning, achieving nearly double the accuracy of existing models.
CED reveals that VLMs can be trained to prioritize evidence-based reasoning over language shortcuts, leading to more reliable visual understanding.
Models that ignore context may seem robust, but they can fail spectacularly when the context is actually trustworthy.
F$^2$Agent achieves over 20% better annualized returns than existing models by dynamically capturing inter-modality dependencies and resisting market noise.
World rehearsal enables LLM agents to internalize environment dynamics, achieving superior performance without costly external interactions.
Task-conditioned wrist modeling boosts robot manipulation accuracy, revealing how wrist interactions can be anticipated for better performance.
PrivDPO achieves robust LLM alignment while maintaining privacy, outperforming traditional methods in balancing privacy and utility.
Modality Balance can be harnessed as a powerful form of privileged information, leading to substantial gains in reasoning performance for multimodal models.
The first publicly available dataset for early pregnancy fetal ultrasound screening could revolutionize automated diagnostics and standardization in prenatal care.
RAG-Stack uncovers quality-performance trade-offs in RAG systems, achieving up to 153% more effective configurations than current methods.
Hierarchical tokenization transforms bounding-box representation, leading to substantial improvements in localization accuracy and model performance across various benchmarks.
SPEAR achieves a staggering 99.5% increase in click recall by aligning query rewrites with user intent, transforming e-commerce search effectiveness.
Hypergraph-based failure attribution can boost LLM reasoning accuracy by efficiently pinpointing the root causes of errors, outperforming traditional methods.
Personalization in LLM agents is more complex than previously thought, with existing benchmarks failing to capture the dynamic nature of user preferences and their impact on task execution.
Search effort doesn't guarantee better answers; instead, the quality of retrieved evidence is the true driver of accuracy in long-horizon search agents.
Personalized 3D avatars and LLM-driven therapists can significantly enhance emotional bonding in digital mental health interventions.
SPIRAL aligns visual representations with textual semantics, achieving near-native performance in Vision-Text Compression without external supervision.
DART achieves a remarkable 75% reduction in inference cache size while enhancing associative recall in long-context sequence modeling.
Low-frequency bias in EEG models can be corrected, leading to a significant boost in performance across a wide range of tasks.
Fetch-then-Explore enables search agents to retain and revisit selected pages, significantly enhancing their accuracy and efficiency in information retrieval.
Redundant neuron encoding can thwart white-box attacks, ensuring safety without sacrificing model performance.
Switching document formats can lead to accuracy drops of over 53%, revealing a hidden vulnerability in LLM workflows that demands urgent attention.
ConMem cuts input tokens by 88.2% while boosting QA accuracy to 76.0%, revolutionizing how inspection logs are utilized in risk assessment.
Even the best vision-language models struggle with reliable evaluation of computer-using agents, but OS-Shepherd models offer a low-cost solution that matches their performance.
Incomplete modalities can lead to significant distortions in multimodal recommendations, but CaIRec effectively calibrates item representations to enhance performance and coherence.
CoRAS can significantly reduce the number of measurements needed for accurate image reconstruction while ensuring error thresholds are met, outperforming fixed-rate methods.
FleetScape reveals that spatial interaction can significantly enhance drone fleet supervision, but larger fleets may compromise situational awareness.
SkillRise achieves up to 8.5 percentage points better performance than leading methods by effectively reusing transferable skills across related tasks.
Action diversity is most valuable when a VLA model is likely to fail, leading to a 17.3% boost in success rates without unnecessary perturbations when success is probable.
Predicting human decisions requires understanding not just the physical world, but also the mental states that drive behavior鈥擬WM makes this explicit.
Explanation quality is a critical yet overlooked dimension of LLM agent performance, with many agents generating misleading explanations that can lead to incorrect code assessments.
Functionalization can lead to chemically distinct changes in electronic structure, revealing critical insights into lithium-metal electrolyte behavior.
Proprietary MLLMs may achieve high diagnostic accuracy, but they still struggle with reliable clinical reasoning, revealing significant gaps in their practical utility.
LaP-Forensics reveals that leveraging reconstruction-based evidence can significantly enhance deepfake detection accuracy against state-of-the-art generative models.
Existing AI-generated image detectors falter dramatically, with accuracy plummeting from 91-96% to as low as 54-66% when faced with realistic manipulations.
Coding agents struggle to find the right context, with logged trajectories missing gold files in 27-35% of cases, revealing critical gaps in retrieval effectiveness.
Task transformation enables LLMs to achieve self-improvement in open-ended tasks without the biases and costs associated with human evaluators.
PrefReward reveals how an explicit user preference matrix can drastically enhance personalization in text generation, outperforming traditional methods in both quality and interpretability.
IRIS can detect model substitutions and routing dilutions in LLM gateways with unprecedented accuracy using only the output text, challenging the reliability of commercial AI services.
Hallucinations in chemical reasoning models coexist with correct answers, revealing a complex relationship that challenges our understanding of model reliability.
Extreme heat is not just a risk factor; it amplifies fault predictions over time, impacting nearly 30% of EV charging posts and reshaping maintenance strategies for climate resilience.
Achieving state-of-the-art rendering quality with over 5.7 times fewer Gaussians, ATSplat redefines efficiency in 3D scene representation.
Unveiling the hidden NUMA architecture of GPUs could revolutionize how we optimize memory efficiency in high-performance computing.
NSM primes outperform traditional appraisal methods in controlling emotional responses in LLMs, revealing a more effective explanatory framework for AI emotion.
Automating real-to-sim conversion with vision-language agents could revolutionize how we simulate robotic interactions, making it faster and cheaper than ever before.
SGN enables effective data generation for shifted target domains without the need for retraining, transforming how we approach data augmentation in machine learning.
Test suites can dramatically enhance issue localization accuracy, bridging the gap between abstract descriptions and concrete code.
Landmark bias can lead to significant inaccuracies in geo-localization, but HoloGeo effectively mitigates this issue through evidence-driven reasoning, outperforming existing models.
Small visual perturbations can cause World-Action Models to execute harmful actions while still predicting a plausible future, revealing a critical vulnerability in their design.
Kaleidoscope reveals that a structured, context-aware evaluation process can significantly enhance the reliability of automated scoring in AI applications.
Quantum topological data encoding reveals that quantum representations can unlock deeper insights from complex datasets, surpassing classical methods in capturing topological nuances.
NodeImport reveals that strategically filtering nodes based on importance can dramatically enhance GNN performance in imbalanced settings.
Induced anger can lock LLMs into poor decision patterns by reducing their sensitivity to penalties, unlike human decision-making.
RecRec reveals that decoupling reasoning from prediction can significantly enhance sequential recommendation performance, breaking free from the constraints of fixed-dimensional states.
EnCF outperforms traditional filters in complex observation scenarios, revealing a new frontier in data assimilation techniques.
LLM-generated bug reports often hinge on implicit assumptions, and this framework reveals how to validate their correctness through a novel witness-generation approach.
EasyOPD unifies on-policy distillation methods, enabling seamless integration and superior performance across diverse tasks in large language models.
TIGER achieves remarkable speedup in multimodal generation by intelligently routing visual tokens based on textual context, outpacing traditional methods.
TC-MAF achieves unprecedented anomaly detection performance by effectively fusing RGB and 3D evidence, setting a new benchmark in multimodal industrial applications.
Surpassing traditional Monte Carlo methods, this approach offers a stable and efficient alternative for learning neural set functions, dramatically cutting down computational costs.
Modern LLM performance hinges on dependency structures rather than individual instruction latencies, revealing a critical insight for GPU optimization.
CycleGRPO achieves simultaneous region understanding and localization in MLLMs without any reliance on textual ground truths, revolutionizing multimodal task integration.
Conveying depth as text rather than images can significantly boost spatial reasoning in vision-language models, challenging conventional approaches.
LLM judges may misinterpret peer review quality, favoring superficial traits over genuine analytical depth, raising questions about their reliability in academic assessments.
Over 1,100 submissions reveal groundbreaking advancements in sports video understanding, with new methods pushing the boundaries of action prediction and localization.
nMAS can cut gastric biopsy report review time from over 83 hours to just 1.4 hours, unlocking substantial efficiency gains in clinical settings.
Navigation models trained in Image2Sim's synthetic environments outperform traditional methods, achieving zero-shot transfer to real-world settings.