Mar 16, 2026arXiv:2603.15295

Datasets for Verb Alternations across Languages: BLM Templates and Data Augmentation Strategies

AI Summary

This paper introduces new datasets for evaluating LLMs' ability to capture cross-sentence verb alternations in English, German, Italian, and Hebrew. The datasets are based on the Blackbird Language Matrices (BLMs) task, a linguistic puzzle requiring models to identify sentences that complete syntactic and semantic patterns. The authors explore different template complexities and data augmentation strategies, providing baseline results that highlight the datasets' diagnostic potential.

Key Contribution

LLMs struggle with systematic cross-sentence knowledge of verb alternations, a weakness exposed by new Blackbird Language Matrices (BLMs) datasets in English, German, Italian, and Hebrew.

Abstract

Large language models (LLMs) have shown remarkable performance across various sentence-based linguistic phenomena, yet their ability to capture cross-sentence paradigmatic patterns, such as verb alternations, remains underexplored. In this work, we present curated paradigm-based datasets for four languages, designed to probe systematic cross-sentence knowledge of verb alternations (change-of-state and object-drop constructions in English, German and Italian, and Hebrew binyanim). The datasets comprise thousands of the Blackbird Language Matrices (BLMs) problems. The BLM task -- an RPM/ARC-like task devised specifically for language -- is a controlled linguistic puzzle where models must select the sentence that completes a pattern according to syntactic and semantic rules. We introduce three types of templates varying in complexity and apply linguistically-informed data augmentation strategies across synthetic and natural data. We provide simple baseline performance results across English, Italian, German, and Hebrew, that demonstrate the diagnostic usefulness of the datasets.

Data Curation & Synthetic Data Eval Frameworks & Benchmarks Natural Language Processing

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Datasets for Verb Alternations across Languages: BLM Templates and Data Augmentation Strategies

Related Papers