Search papers, labs, and topics across Lattice.
This study evaluates Qiushi Engine v0.8, configured with a DeepSeek deepseek-v4pro-preview backend, across 40 complex research tasks in AstaBench E2E-Bench-Hard that span hypothesis design, code generation, execution, and manuscript delivery. Establishing baseline performance on rigorous, fully autonomous scientific workflows is critical as the community transitions from static code generation to multi-step experimental discovery. The system satisfied 82.1% of all rubric criteria and achieved a 10% complete-task success rate鈥攁 3.3脳 improvement over official baselines at an average cost of $15.21 per task鈥攚hile revealing that missing ablations and repeated experimental runs remain primary bottlenecks for autonomous agents.
Tripling the previous state-of-the-art on full-lifecycle scientific discovery benchmarks, Qiushi Engine satisfies over 82% of end-to-end research rubrics for just $15 per project.
This report analyzes Qiushi Engine v0.8 across all 40 test tasks in AstaBench E2E-Bench-Hard, a benchmark that requires autonomous agents to carry a research question through experimental design, code implementation, actual execution, result analysis, and report delivery. Qiushi Engine is model-configurable; this evaluation selected DeepSeek deepseek-v4pro-preview as the model backend. The official AstaBench leaderboard records a score of 0.816 and an average benchmark cost of USD 15.209 per task, while the full-precision local recomputation is $81.59 \pm 1.87$. Four tasks satisfied every rubric item, yielding a full-task completion rate of 4/40 = 10% -- 7 percentage points above, and about 3.3 times, the approximately 3% best rate reported for AstaBench's official agents. Across 507 required rubric items, 416 were satisfied (82.1%). Official scoring archives and 40 Meta-Trace records show sustained production and verification of reports, code, and experimental artifacts; the principal gaps lie in repeated runs, external dependencies, specified metrics, and ablation studies. The report explains the benchmark, system workflow, aggregate results, representative cases, and limits of interpretation.