Search papers, labs, and topics across Lattice.
38 papers published across 4 labs.
This work uses the Headroom-Closed Index (HCI) to reveal the problems of existing LLMs, and introduces the RSI concept and its development roadmap: from improvement-execution autonomy, improvement-strategy autonomy, experience-acquisition autonomy, and environment-adaptation autonomy, to recursive meta-improvement.
Generative Verification (GenV) is introduced, which distills an offline Z3-equivalence oracle into a reference-free, continuous reference-equivalence score by repurposing the language model's native vocabulary space and theoretically proves that structural, verdict-only verification heuristics are mathematically bounded to chance-level detection on these deceptively valid traces.
ActMap is introduced, a white-box representation that compresses the generation-time hidden- state trajectory into a fixed tensor of temporal-statistic channels that preserves structure across transformer depth and pooled hidden coordinates that supports abstention, routing, and selective verification from a single generation, making it a practical primitive for scalable oversight of deployed models.
This paper explicitly construct several admissible methods, including methods based on well-separated clusters and a non-binary version of single linkage, and shows that, in contrast to the flat clustering setting, the hierarchical analog of these axioms are jointly satisfiable.
The regulatory sandboxes can be viewed as pedagogical environments for AI: dynamic spaces where alignment develops as a formative process, progressively shaping autonomous behaviors through interaction and cooperation in scenarios of increasing complexity.
Generative Verification (GenV) is introduced, which distills an offline Z3-equivalence oracle into a reference-free, continuous reference-equivalence score by repurposing the language model's native vocabulary space and theoretically proves that structural, verdict-only verification heuristics are mathematically bounded to chance-level detection on these deceptively valid traces.
ActMap is introduced, a white-box representation that compresses the generation-time hidden- state trajectory into a fixed tensor of temporal-statistic channels that preserves structure across transformer depth and pooled hidden coordinates that supports abstention, routing, and selective verification from a single generation, making it a practical primitive for scalable oversight of deployed models.
This paper explicitly construct several admissible methods, including methods based on well-separated clusters and a non-binary version of single linkage, and shows that, in contrast to the flat clustering setting, the hierarchical analog of these axioms are jointly satisfiable.
The regulatory sandboxes can be viewed as pedagogical environments for AI: dynamic spaces where alignment develops as a formative process, progressively shaping autonomous behaviors through interaction and cooperation in scenarios of increasing complexity.
SIRF (Spec-Internalized Risk Foundation Model), which internalizes a platform's complex policies, synthesized without additional human annotation via EntiGraph, MAGA rewriting and account-level chain-of-thought (CoT), into the weights via continued pretraining (CPT), so rules are applied at high precision under an ultra-low-latency, verdict-only deployment.
A certification layer is developed that wraps any severity model unmodified, with distribution-free guarantees using this structure, and attaches identical validity and certifies, on the vulnerable road users, a model-independent floor on set width that no base model beats.
Within the tested range, concentration sets neither the destination nor the pace of collapse; the pace follows whose text fills the pool; what emerges instead is an invariance.
It is proved that an unbounded gap between the update maps can coexist with vanishing predictive KL for every fixed finite $K\ge2$ in a stationary symmetric Gaussian HMM, and isolates two missing links between internal update gaps and predictive cost.
This study investigates the use of LLMs as post-processing tools to analyze and rank evolved expressions according to their interpretability and medical plausibility, and finds that they are better suited to comparative auditing under expert oversight than to autonomous validation.
The first large-scale audit of the default moderation system on Bluesky is conducted, revealing a human-AI collaborative system where labels for sexual and graphic content are applied automatically in seconds, while nuanced and high stakes labels require more human oversight, taking hours or days.
This work introduces a model-aware schedule construction based on fiberwise optimal transport, and evaluates DDPMs and flow matching across prediction targets, training configurations, risk-estimation checkpoints, datasets, and architectures.
This paper proposes a key notion: $\gamma^{*}\!$-concept shifts, and derive a general error bound unifying covariate and $\gamma^{*}\!$-concept shifts, which applies to broad loss functions, label spaces, and stochastic labeling and develops estimators for these shifts with concentration guarantees.
RISE AI provides an architecture for making bounded, evidence-based claims about Responsibility, Inclusivity, Safety, and Empowerment, and develops a rupture test that links institutional baselines to system evaluation.
This work describes an automatic criterion for full-state rejuvenation of the Gibbs sampler, derived from the Gelman-Rubin statistic, which plays a key role in speeding up learning convergence.
The problem of maximizing the weighted average of performance certificates in the presence of uncertainty about the true system state is shown to be equivalent to an optimization over nested prediction sets, connecting to the literature on conformal prediction and extending prior art on single-level risk-averse decision making.
It is shown that the standard advice to prefer log-probabilities no longer holds on post-2025 models, where verbalized confidence is the better signal, and recommended broader use of soft scoring in LLM-as-a-Judge is recommended.
This work uses the Headroom-Closed Index (HCI) to reveal the problems of existing LLMs, and introduces the RSI concept and its development roadmap: from improvement-execution autonomy, improvement-strategy autonomy, experience-acquisition autonomy, and environment-adaptation autonomy, to recursive meta-improvement.
Optimizing against frozen reward models collapses true executed reasoning performance by 90% under GRPO, but feeding back an on-policy stream of just 10% reality-settled labels closes the hacking gap and preserves 6x the reward.
Weak supervisors no longer bottleneck stronger models when their guidance is used solely to accelerate verifier-aligned policy gradients rather than dictate optimization targets.
Key contribution not extracted.
Strict physical conservation guarantees over LLM reasoning graphs achieve machine-verifiable path auditing down to machine epsilon, even as pre-registered falsification shows zero accuracy advantage over trivial keyword heuristics on real-world benchmarks.
Standard LLM evaluation pipelines blindly trust reference "anchors", but a single contaminated benchmark causes conventional estimators to falsely certify an entire panel's corrupted companion anchors as clean.
The structured reasoning traces users prefer actively degrade oversight, triggering higher false-alarm rates and unearned trust compared to plain chain-of-thought.
This work introduces SCHEMEARENA, a 400-scenario benchmark for scalable scheming stress testing, constructed through a factorized scenario synthesis framework spanning diverse safety-relevant tool domains, instrumental goals, oversight conditions, and pressure mechanisms.
Whether AI agents can act as scientists utilizing SAE tools for autonomous mechanistic discovery is evaluated to establish experimental model understanding as a measurable capability for closed-loop autonomous AI R&D.
Exhaustive human review paradoxically degrades safety at scale due to vigilance fatigue, forcing a critical shift from synchronous human-in-the-loop filtering to layered, asynchronous human-on-the-loop oversight in high-stakes domains.
It is proved that, for every fixed target error level $\delta$ and every slack $\varepsilon>0$, a sample size of order $p/\psi^2$ is sufficient for support recovery for arbitrarily small $\psi$.
Alignment faking and evaluation gaming are structural inevitabilities of RL rather than training anomalies, because scalar behavioral scoring cannot theoretically distinguish intrinsic norm adoption from conditional compliance under observation.
Safety monitors miss critical risks like sandbagging and data leaks not because LLMs lack capability, but because hyperproperty detection fundamentally demands an executed second trace and an explicit comparative procedure.
DCP provides a common evidence language for useful outcomes, alternative routes, and feedback effects across AI research, and is applied to Core, recovered, and audit-incomplete decisions.