Search papers, labs, and topics across Lattice.
5
0
6
0
VA-Bench is introduced to evaluate the complete observe-reason-act-revise loop and tests whether general-purpose MLLMs can turn visual demonstrations and actively acquired evidence into successful embodied action.
This work presents TeleAntiFraud 2.0, constructed with the Mixed-Tree Anti-Fraud Generation Pipeline and evaluated under a monthly frozen evaluation protocol, establishing near-domain construction and collapse-aware reporting as core requirements for evaluating audio-based telecom-fraud models under realistic confusable conditions.
FRAUDSkill is proposed, a structured frozen-weight adaptation framework that leaves the underlying audio-language model unchanged while optimizing an external layer of skill programs, route-specific policies, and decision rules and combines structured output control with validation-guided multi-path inference to ensure protocol-compliant predictions.
RiskChainBench introduces RiskChainBench, pairing 3,600 synthetic token-text restoration inputs from 600 source sessions with 600 corresponding human-labeled local web environments that produces a frozen, evidence-cited risk report without message-side semantics or domain-reputation cues.