- Research Article
- 10.1007/s11185-026-09330-4
The words фигушки and хренушки as specific means of expressing negation in colloquial Russian
- Feb 26, 2026
- Russian Linguistics
- Elena Nikishina
Publications from 2021 to 2026
Showing 10 of 51 papers
The words фигушки and хренушки as specific means of expressing negation in colloquial Russian
Detecting LLM-Generated Text with Trigram–Cosine Stylometric Delta: An Unsupervised and Interpretable Approach
Background: Contemporary methods for detecting synthetic text, including model-specific detectors and transformer-based classifiers, often rely on intensive training or on features tied to particular language models, which restricts their generalizability to unfamiliar LLMs and diverse domains. Purpose: This study aims to advance text attribution research by introducing a stylometry-based approach that utilizes trigram-based cosine delta as a lightweight and interpretable metric for distinguishing LLM-generated texts from human-written texts, irrespective of the underlying generation strategy. Method: A corpus of Russian diary entries was compiled, encompassing both authentic human-written texts and synthetic counterparts generated through few-shot prompting and finetuned LoRA models. To evaluate the effectiveness of the proposed approach, multiple stylometric-delta variations were examined, integrating uni-, bi-, and trigram features with Manhattan and cosine distance metrics. Results: The evaluation demonstrated that the trigram–cosine delta consistently achieved the highest performance across experimental conditions, reaching an Adjusted Rand Index of approximately 0.70. This markedly surpassed both the finetuned RuModernBERT baseline (ARI ≈ 0.28) and the classic unigram-based delta (ARI ≈ 0.53). Importantly, the method proved effective not only within the Russian diary corpus but also when applied to the RuATD benchmark, where it successfully separated human-authored and machine-generated texts and produced coherent clustering of related model families. Conclusion: The findings confirm that trigram–cosine stylometric delta offers a robust, interpretable, and computationally efficient strategy for detecting LLM-generated texts across diverse generation strategies, including few-shot prompting and finetuning. By capturing discourse-level stylistic cohesion, the method advances beyond surface fluency and provides a scalable, unsupervised alternative to classifier-based detectors. While current validation is limited to Russian diaries and selected generation models, the approach demonstrates clear potential for broader application across domains, languages, and emerging state-of-the-art LLMs.
Read moreEmancipation of affixes in modern literary speech
The article examines the emancipation of affixes, i. e. their separation from the morphemic composition of a word and their isolated use in context, using poetic and prosaic texts of the 19th–21st centuries. The article aims to identify the main stages of this process characteristic of Russian literary speech and to trace its history. The study uses descriptive and structural-semantic methods as well as word-formation analysis of lexical units. It is noted that emancipation in literary speech can be observed in different types of morphemes, primarily prefixes and suffixes. We have identified the types of affixes that are most frequently emancipated and lexicalised in texts. The paper concludes that morpheme emancipation in Russian literary speech has a long history: its examples are found in texts as early as in the first half of the 19th century. In modern literary speech, the use of emancipated affixes is becoming more extensive. We have identified several stages of morpheme emancipation, the last being morpheme lexicalisation, i. e., the final transformation of an affix into an independent lexical unit. It is shown that emancipated affixes can be combined in a text and form the basis of its composition, perform a text-forming function, and serve as a literary work title.
Read moreMutando Mutanda. II: Notes on Emendations
The article, which continues the problems of our 2021 paper of the same name, analyzes examples that testify to the unreliability and even danger of old non-linguistic editions of monuments of Old Russian writing, primarily due to arbitrary emendations and conjectures of the text, often introduced by publishers without any reservations. Among such conjectures, the author analyzes the imaginary ancient fixation of the word kondakar in the record to a Stiherarium of the 12th century, which hides a unique use of the previously unknown conjunction ponda ‘however’, allegedly the oldest examples of the nom. pl. masc. on -a in the Chronicle of Avraamka of the 15th century, which entered scientific circulation due to the incorrect disclosure of superimposed letters by the first publishers, as well as three cases of the hypercorrect “restoration” of the preposition vъ in primordial constructions with locative and accusative without prepositions in the canonical edition of the Laurentian Chronicle of 1377. The article ends with a conclusion about the need for new linguistic editions of the main written monuments and the importance of turning to the manuscript tradition.
Read moreCorpus linguistics nowadays
The article proposes a general presentation of corpus linguistics, its history, its methodology and its influence on current views on how a language must be studied – the process which is usually referred to as “corpus revolution”.
Read moreWilliam Croft’s approach to constructions in comparison with the system of constructions in the Russian Constructicon
In this article, we discuss the approach to description of constructions adopted in William Croft’s monograph “Morphosyntax. Constructions of the World’s languages” (2022) and compare it to the approach to description and classification that is used in the Russian Constructicon. We conclude that Croft’s system and the Russian Constructicon show several substantial differences. First, Croft’s approach is based on the notion of Comparative Concepts and on an a priori classification of grammatical domains. By contrast, in the Russian Constructicon, a bottom-up approach is taken: it presupposes collecting the most representative inventory of constructions and, then, creating a system for classifying them. Second, the Russian Constructicon represents properties of constructions as a system of tags, where a construction can bear multiple tags, while Croft does not discuss cases of multiple tagging or grammatical class intersection. Finally, Croft’s system focuses on the core of grammar and includes mainly those values that are grammaticalized, while in the Russian Constructicon, attention is given not only to grammatical constructions, but also to constructions that can be termed ‘quasi-grammatical’ or ‘lexicalized’ — they have narrow semantics and combinational properties (here belong, for instance, iterative / frequentative constructions, such as to i delo ‘frequently’, and constructions with the terminative / resultative meaning, such as svoё otguljal ‘[he] is done with having fun’).
Read moreTextual studies of the era of big data and neural networks
This article analyses the emergence of new technologies for working with big data which can be highly helpful to philologists, studying the diachronic development. First of all, that applies to the study and publication of Old Russian manuscripts with traditional liturgical texts used for church service. These manuscripts existed in a huge number of folios, and in the process of copying were subjected to considerable textual unification. That makes it very difficult to study them by the laborious methods of traditional textual criticism. Now, when the full text of the monuments can be automatically processed, the creation of the Linguistic intellectual environment (LIE) has been lounged. This tool will provide new opportunities for the study of Slavonic liturgical texts from different historical periods. As a result, we will create a corpus of liturgical texts of the 11th–17th centuries, obtained using a program for automatic text recognition of manuscripts, with annotation und search module. The user of the LIE will be able to receive a complete list of variant readings for each fragment of a liturgical book of the widest range of manuscripts. In fact, we are talking of a new type of publication of traditional liturgical texts, when the user can set the parameters for the edition in accordance with his research interests.
Read moreComplete reduction of unstressed vowels in the standard russian language and its reflection in dictionaries
The article is devoted to the complete reduction of unstressed vowels in the standard Russian language. The article shows that the cases of dieresis of unstressed vowels in the dictionaries of various types are not always represented consistently. For example, in the academic “Dictionary of the Modern Standard Russian language” variants with a graphic reflection of the vowel dieresis are noted in prítolka, páportnik, zhávronok, púgvitsa, but “Great Academic Dictionary of the Russian Language” fixes only the variants púgvitsa, púgvichnikъ. In orthoepic sources, the corresponding pronunciation recommendations are also contradictory: in the “Dictionary of Stress and Pronunciation of Words in the Russian Language” the variants príto[lk]a and próvo[lk]a are marked as incorrect, while in the “Big Orthoepic Dictionary” these variants are recognized as the main ones. The article proposes to distinguish, on the one hand, the lexical or morphological attachment of the vowel dieresis to a specific word or morpheme, and, on the other hand, the actual lexicalization of a pronunciation without a vowel. In the first case, fixing variants with vowel dieresis in orthoepic-type dictionaries seems redundant, since the choice of pronunciation, as a rule, depends on phrasal position and occurs automatically. In the second case, the pronunciation of the words does not necessarily derive from their spelling, and therefore they require orthoepic commentary.
Read moreDEMONSTRATIVE TOT IN THE HISTORY OF THE PRESENT STATE OF THE OTTOMAN EMPIRE BY PIOTR TOLSTOY
The article examines the use of the demonstrative pronoun tot in the language of The History of the Present State of the Ottoman Empire (manuscript BAN, 31.3.22), which was translated by Piotr Tolstoy from Italian at the beginning of the 18 th century. The examples of tot are grouped by syntactic and semantic functions and compared to the Italian original. While some Italian-Russian parallels for other demonstratives are merely constant (such as quello - onyj , questo - sei , tale , tanto - takoj ), the demonstrative tot can be used to translate any Italian anaphoric demonstrative, which indicates that it was being unmarked as a means of anaphora. The noun + tot group often serves as the equivalent of a noun with a definite article, but in these cases an antecedent can be found in the preceding text and therefore the demonstrative is a means of anaphora; if an antecedent cannot be found, the translator has to use another grammar construction. This implies that the demonstrative tot was not a complete analogue of the definite article. In many cases tot does not have a corresponding demonstrative in the Italian text: the fact that reveals the various grammatical functions of tot . For example, as a determiner, it is widely used in the Russian text to translate different Italian syntactic constructions; some adverbial collocations and compound conjunctions with tot are used to translate Italian adverbs and conjunctions that do not contain demonstratives.
Read moreThe Slavonic Alphabetical Hymn in Medieval Russian Horologia of the 13th–14th Centuries
The article is devoted to a Slavonic alphabetical hymn found in the earliest Slavonic Horologia and preserved in two sources: the well-known Yaroslavl Horologion of the 13th century and the recently discovered liturgical book for cell prayer Sof. 1129 from the 14th century. In the latter source the alphabetical hymn contains two verses starting with the letter Ш, a feature that it shares also with the Alphabetical Prayer of Constantine of Preslav and with the oldest Slavonic alphabetical hymnography, recently identified by G. Popov, found in the Festal Menaion. The question of the evidence of the Glagolitic alphabet found in Cyrillic acrostics is once again discussed in the article, and the opinion of V. Mošin that the Glagolitic acrostics are closely related to the alphabets that figure in On the Letters by Chernorizets Hrabar and in the Munich Abecedarium is confirmed. The subsequent fate in Medieval Rus of the alphabetical hymn from the Horologion is then traced. The hymn, in a significantly reworked form, is widely found in manuscripts starting with the end of the 15th century as a prayer titled “Alphabetical [prayer] of repentance”. At the end of the article, an attempt is made to show the main differences between acrostics based on the Glagolitic and Cyrillic alphabets.
Read more