Search papers, labs, and topics across Lattice.
This pilot study investigates the optimal decomposition of tasks in LLM-agent systems for cross-border VAT determination, comparing configurations from a single strong agent to multiple narrow agents. By fixing the activity surface and varying the assignment of subtasks, the research evaluates performance across 4,400 runs, revealing that intermediate configurations yield higher accuracy than the single agent but do not meet the predetermined performance thresholds. The findings suggest that while task decomposition can enhance accuracy, it does not guarantee a Pareto improvement over a single agent, highlighting the complexities in designing effective LLM-agent systems.
Intermediate task decomposition in LLM-agent systems can improve accuracy but fails to consistently outperform a single agent in VAT determination tasks.
Recent LLM-agent systems make conflicting design bets: decompose work across many narrow agents, or use one strong tool-using agent. This pilot studies that choice on bounded cross-border VAT determination with reverse charge, where every case has an oracle label and each intermediate decision is independently scoreable. We hold the activity surface fixed (subtasks, tools, I/O schemas, validation checks, orchestrator, base model, and merge policy) and vary only the assignment of subtasks to workers across four orchestrated configurations, from one wide worker to five narrow ones, against S0, a tuned no-orchestrator single agent, with a deterministic rule engine as oracle. The program spans 4,400 runs: a 40-case, five-repeat main sweep, matched-token arms separating prompt-budget from agent-count effects, and three failure-injection arms, all judged against pre-registered falsification criteria. The two intermediate configurations lead on accuracy (0.830, against endpoints at 0.720 and 0.770) but miss the pre-stated bar against the fine endpoint, so the intermediate-optimum hypothesis remains unsupported at pilot scale. The single agent does not Pareto-dominate the orchestrated set. The matched-token criterion fires: the budget-matched single agent lands 6.5 points below the leader, but the interval includes zero, so any advantage is consistent with a prompt-budget explanation. Under injection, availability faults are absorbed at every granularity, with wide-scope restart over-recovering its baseline by +0.160, while one schema-conforming hallucinated record degrades every configuration and inverts the ordering, hitting fragmented configurations hardest. The contribution is a bounded, preregistered pilot heuristic for right-sizing decomposition (place one partition boundary at the dependency-layer midpoint), released with oracle, dataset, harness, raw traces, and analysis pipeline.