Search papers, labs, and topics across Lattice.
12
0
12
3
The model family shows gains in held-out scientific-code repair and across selected general-purpose benchmarks in code, reasoning, and knowledge, providing evidence of positive transfer from scientific experience to broader capabilities.
This work introduces and releases ScienceBuddy, an interactive scientific research workspace that brings continually improving scientific agents into researchers'everyday workflows and presents case studies of researcher interaction, harness refinement, and model learning, with the benchmark cases spanning four scientific task families.
This work forms a unified approach to capability formation and process-centered evaluation, enabling discovery behavior to be trained, improved, and measured beyond final-answer performance.
Recuris transforms long-horizon task execution by reducing common failures by up to 80% and achieving state-of-the-art success rates across multiple models.
Skill-switching accuracy in LLMs drops significantly on complex tasks, but a new training approach boosts performance from 34.4% to 68.4% on challenging benchmarks.
Agents can exhibit significant performance gains from retained experience, but the pathways to these improvements are often unclear and model-dependent.
Achieving parallel region captioning with multimodal diffusion models could redefine efficiency benchmarks in visual perception tasks.
MFD reduces noise and boosts stability in flow matching models, achieving state-of-the-art results in challenging generative tasks.
Ditch the textual explanations: symbolic outputs like bounding boxes are the secret sauce for boosting multimodal verifier performance.
Forget finetuning on curated datasets – OpenClaw-RL lets agents learn directly and continuously from *every* interaction, turning user replies, tool outputs, and even GUI changes into valuable RL signals.
Stop benchmarking agent components in isolation; AgentSelect lets you train models to recommend *entire* agent configurations based on narrative queries.