Search papers, labs, and topics across Lattice.
4
0
7
11
Reliable detection of malicious agent skills hinges on a comprehensive benchmark that reveals stark differences in threat composition across sources.
Real-world coding tasks, reverse-engineered from actual commits and scenarios, make Tencent WorkBuddy Bench a game-changer in contamination-resistant evaluation for coding agents.
CRISP achieves superior long-horizon point cloud forecasting and versatile downstream task performance by leveraging a unique forecasting-based pretraining approach with camera-radar fusion.
Prompt complexity is a critical dimension that significantly influences maintenance effort, challenging traditional views that prioritize code-level metrics alone.