Search papers, labs, and topics across Lattice.
This study investigates the extent to which multilingual language models transfer factual knowledge across languages during continued pretraining, using a novel intervention-based framework that removes specific facts from Persian training data. By constructing the SIFT resource, which includes 500 triples stratified by cultural origin, the authors reveal that the transfer of English-acquired facts to Persian is largely limited, particularly under stringent fact removal conditions. The findings highlight that sentence-level co-occurrence removal does not effectively eliminate fact signals, and reliance on easier negative candidate sets can misrepresent the actual transfer capabilities of these models.
Multilingual models struggle to transfer factual knowledge across languages, with most English-acquired facts failing to make the leap to Persian after targeted data interventions.
Do multilingual language models transfer factual knowledge across languages during continued pretraining, or do they mostly recall facts learned directly from the target-language data? To answer this question more reliably, we propose an intervention-based framework: starting from an English-pretrained model, we continue pretraining on Persian data from which specific facts have been systematically removed at varying levels of granularity. We construct SIFT, a resource of 500 triples across 20 relations, stratified by the cultural origin of each fact's subject into general (globally prominent) and Persian-related entities, designed for both systematic fact removal from training data and evaluation, with natively written Persian cloze templates. Our results show that fact transfer is very limited: under the strictest removal condition, a large majority of English-acquired facts fail to transfer into Persian. We further show that sentence-level co-occurrence removal is insufficient to eliminate fact signal, and that easier (randomly selected) negative candidate sets substantially inflate apparent transfer by rewarding shallow associative heuristics, while performance on a harder candidate set that allows for less reliance on heuristics is much lower. Finally, we show that source-language entity frequency has a large influence, with Persian-related facts, which are orders of magnitude rarer in the English corpus, hardly transferring.