Datasets for Verb Alternations across Languages: BLM Templates and Data Augmentation Strategies
| Type of publication: | Conference paper |
| Citation: | Samo_LREC2026_2026 |
| Booktitle: | Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026) |
| Year: | 2026 |
| DOI: | 10.63317/4t48qjruy2ce |
| Abstract: | Large language models (LLMs) have shown remarkable performance across various sentence-based linguistic phenomena, yet their ability to capture cross-sentence paradigmatic patterns, such as verb alternations, remains underexplored. In this work, we present curated paradigm-based datasets for four languages, designed to probe systematic cross-sentence knowledge of verb alternations (change-of-state and object-drop constructions in English, German and Italian, and Hebrew binyanim). The datasets comprise thousands of the Blackbird Language Matrices (BLMs) problems. The BLM task – an RPM/ARC-like task devised specifically for language – is a controlled linguistic puzzle where models must select the sentence that completes a pattern according to syntactic and semantic rules. We introduce three types of templates varying in complexity and apply linguistically-informed data augmentation strategies across synthetic and natural data. We provide simple baseline performance results across English, Italian, German, and Hebrew, that demonstrate the diagnostic usefulness of the datasets. |
| Additional Research Programs: |
AI for Everyone |
| Keywords: | |
| Projects: |
Idiap |
| Authors: | |
| Added by: | [UNK] |
| Total mark: | 0 |
|
Attachments
|
|
|
Notes
|
|
|
|
|