- Research Article
- 10.1007/s40747-025-02224-w
UniMTES: a unified framework with trait-description-aware for multi-trait essay scoring
- Jan 14, 2026
- Complex & Intelligent Systems
- Jingbo Sun + 3 more +3
Publications from 2021 to 2026
Showing 10 of 21 papers
UniMTES: a unified framework with trait-description-aware for multi-trait essay scoring
Exploring Motivation and Emotional Experience in Observational Research for Individuals at Risk of ALS/FTD Spectrum Disorders.
Numerous observational studies are available to asymptomatic individuals at risk to carry or known carriers of pathogenic variations associated with amyotrophic lateral sclerosis and frontotemporal degeneration (ALS-FTD) spectrum disorders. Little is known about such individuals' motivations for participation or the impact on their emotional well-being. Asymptomatic at-risk adults, with or without genetic status known, were recruited through social media advocacy groups and by National Society of Genetic Counselors ALS-FTD special interest group members. Interviews were conducted through secure videoconferencing. Two coders independently analyzed interview transcripts, followed by thematic content analysis. Twelve participants (9 status-aware and 3 status-unaware) were interviewed, representing experience with 11 observational studies. Some motivations for participation aligned with previous literature, including altruism, health focus, and intellectual interest. Motivations unique to this population stemmed from the hereditary nature of the disease, including fear of future disease onset and the desire to establish a relationship with a specialized clinical care team, reflecting individual, familial, and societal factors. Benefits of participation included meeting these motivational goals, social connection and support, psychological well-being, and practical benefits. Challenges to participation fell into research-related (e.g., struggles with the observational nature of research), disease-related (e.g., anxiety about disease risk), and logistical (e.g., travel and study procedures) categories. Compared with status-unaware participants, status-aware participants more frequently cited individual motivators for research participation and encountered more research-related challenges when their participation did not align with their anticipated personal health benefits. Interviewees found relationships with providers through research to be rewarding but noted confusion between research and clinical care as a significant challenge. Participation in observational research helps address unmet emotional and medical needs for asymptomatic individuals who are at risk of ALS-FTD spectrum disorders. However, some of these needs are beyond the scope of research, highlighting the need for new models of clinical care for at-risk individuals.
Read moreRelation between Depression Dimensions and Speech Acoustic and Emotion‐based Features
BackgroundDepression is a common and highly heterogeneous disorder in older adults, often linked to faster cognitive decline. While standard questionnaires are subjective, speech analysis may offer a more objective method for characterizing depression. This study investigates the relationship between speech features and depression dimensions in participants at Mount Sinai Alzheimer's Disease Research Center (ADRC).MethodParticipants included healthy controls (n = 31) and individuals with mild cognitive impairment (MCI= 22) and Alzheimer's Disease (AD, n = 16). They described three pictures with neutral, negative and positive themes. Speech features were analyzed using automated pipelines and emotion‐based variables and acoustic features were used for analysis. Depression dimensions (dysphoria, apathy, hopelessness, and memory complaints) were assessed based on Geriatric Depression Scale‐15 (GDS‐15). Mixed model regression was used to assess the relationship of depression dimensions and the emotional nature of the pictures.ResultThe 73 participants (41% male, average age 80.14±8.01) showed that dysphoria and apathy had opposite associations with emotion‐based measures. Dysphoria was associated with higher valence (more positive emotion), while apathy to lower valence. Subjective memory complaint was also associated with lower valence words.Further analysis revealed that apathy was associated with lower pitch and slower speech when describing negative pictures, and dysphoria with a wider pitch range and faster speech for negative pictures. Patients with memory complaints used a narrower pitch range in both positive and negative tasks.ConclusionDysphoria was associated with heightened emotional reactivity, while apathy showed decreased reactivity. Apathy and memory complaints shared similar speech features. Our preliminary results support the use of speech features in distinguishing different depression dimensions even at the subclinical level, offering promising opportunities for use of technology to enhance both understanding and diagnosis of depression.
Read moreA Systematic Review and Bayesian Meta-Analysis of Acoustic Measures of Prosody in Parkinson's Disease.
Linguistic prosody is affected in Parkinson's disease (PD), which implicates the basal ganglia's role in the production of prosody. However, there is no recent systematic synthesis of the available acoustic evidence of prosodic impairment in PD. This study aimed to identify the acoustic features of linguistic prosody that are consistently affected in PD. The authors systematically reviewed articles that reported acoustic features of prosodic production in PD. Articles focused on fundamental frequency (F0) and its variability, intensity and its variability, speech and articulation rate, and pause duration and ratio. From a total of 648 records identified, 36 met criteria for inclusion and exclusion. For each acoustic measurement and task, data from people with PD (PwPD) were compared with those from controls to extract effect sizes. Pooled effect sizes were estimated using robust Bayesian hierarchical regression models. PD was associated with decreased F0 variability and increased pause duration. There was limited evidence of reduced intensity variability and speech rate in PwPD. No evidence was found to suggest that PD affects articulation rate or pause ratio. The primary acoustic parameters of prosody affected by PD are F0 variability and pause duration. The identification of these acoustic parameters has important clinical implications for the selection of PD management strategies. The association of F0 variability and pause duration with PD suggests that the neural circuits controlling these parameters are at least partly shared and might include the basal ganglia. While the current study focused on the phonetic realization of prosodic cues, future studies should examine whether and how PD affects prosody at higher levels of processing. https://doi.org/10.23641/asha.25892923.
Read moreAutomated lexical analysis of story recall in healthy controls
Abstract BackgroundImpaired episodic memory is one of the earliest and most prominent symptoms of Alzheimer’s disease (AD). Examining how patients produce words and phrases in story recall tasks is useful for tracking and understanding progression of their disease. Traditional approaches rely on manual assessments of recalled stories which can be limited, time‐consuming, and not applicable in large‐scale studies. Here, we implement an automated, computational approach to measure verbal episodic memory in a digitally recorded story recall task performed by young healthy speakers using natural language processing models.MethodWe analyzed digitized speech samples of pre‐recorded immediate and delayed Craft Story recall tasks performed by healthy speakers (n = 67, mean age = 20.31 years, SD = 2.17, 45 (67%) females). All transcripts were transcribed by trained annotators and automatically scored for number of verbatim and paraphrase recall, total recall score ( = verbatim+paraphrase), and semantic distance from the original story (degree of semantic similarity between the original craft story and the recalled story based on the Word2vec module in python). We also calculated total word count and number of unique words, and rated all recalled words for word familiarity, concreteness, and semantic ambiguity based on published norms. Number of pauses and total durations of speech and silent pauses were also measured.ResultWe found larger semantic distance between successive words in the delayed recall task compared with immediate recall within speaker (p<.001). Delayed recall also elicited lower total recall scores than immediate recall (p = .03). Comparing the effect sizes of semantic distance vs. total recall scores, our model comparison showed a smaller prediction error for semantic distance (AIC = 95.45) than total recall score (AIC = 736.18). In addition, speakers in the delayed recall task performed more paraphrase recall, had higher total word count, number of unique words, more ambiguous words, and produced more speech (all p‐values<.001).ConclusionOur findings suggest that automated speech analysis can provide informative and cost‐effective measures from a recorded story recall task. The study provides an automated and quantitative way to better measure memory abilities by including semantic distance. Future work will test procedures in patients with AD or other neurodegenerative diseases.
Read moreLanguage-Specific Constraints on Conversation: Evidence from Danish and Norwegian.
Establishing and maintaining mutual understanding in everyday conversations is crucial. To do so, people employ a variety of conversational devices, such as backchannels, repair, and linguistic entrainment. Here, we explore whether the use of conversational devices might be influenced by cross-linguistic differences in the speakers' native language, comparing two matched languages-Danish and Norwegian-differing primarily in their sound structure, with Danish being more opaque, that is, less acoustically distinguished. Across systematically manipulated conversational contexts, we find that processes supporting mutual understanding in conversations vary with external constraints: across different contexts and, crucially, across languages. In accord with our predictions, linguistic entrainment was overall higher in Danish than in Norwegian, while backchannels and repairs presented a more nuanced pattern. These findings are compatible with the hypothesis that native speakers of Danish may compensate for its opaque sound structure by adopting a top-down strategy of building more conversational redundancy through entrainment, which also might reduce the need for repairs. These results suggest that linguistic differences might be met by systematic changes in language processing and use. This paves the way for further cross-linguistic investigations and critical assessment of the interplay between cultural and linguistic factors on the one hand and conversational dynamics on the other.
Read moreSex differences in the temporal dynamics of autistic children’s natural conversations
BackgroundAutistic girls are underdiagnosed compared to autistic boys, even when they experience similar clinical impact. Research suggests that girls present with distinct symptom profiles across a variety of domains, such as language, which may contribute to their underdiagnosis. In this study, we examine sex differences in the temporal dynamics of natural conversations between naïve adult confederates and school-aged children with or without autism, with the goal of improving our understanding of conversational behavior in autistic girls and ultimately improving identification.MethodsForty-five school-aged children with autism (29 boys and 16 girls) and 47 non-autistic/neurotypical (NT) children (23 boys and 24 girls) engaged in a 5-min “get-to-know-you” conversation with a young adult confederate that was unaware of children’s diagnostic status. Groups were matched on IQ estimates. Recordings were time-aligned and orthographically transcribed by trained annotators. Several speech and pause measures were calculated. Groups were compared using analysis of covariance models, controlling for age.ResultsAutistic girls used significantly more words than autistic boys, and produced longer speech segments than all other groups. Autistic boys spoke more slowly than NT children, whereas autistic girls did not differ from NT children in total word counts or speaking rate. Autistic boys interrupted confederates’ speech less often and produced longer between-turn pauses (i.e., responded more slowly when it was their turn) compared to other children. Within-turn pause duration did not differ by group.LimitationsOur sample included verbally fluent children and adolescents aged 6–15 years, so our study results may not replicate in samples of younger children, adults, and individuals who are not verbally fluent. The results of this relatively small study, while compelling, should be interpreted with caution and replicated in a larger sample.ConclusionThis study investigated the temporal dynamics of everyday conversations and demonstrated that autistic girls and boys have distinct natural language profiles. Specifying differences in verbal communication lays the groundwork for the development of sensitive screening and diagnostic tools to more accurately identify autistic girls, and could inform future personalized interventions that improve short- and long-term social communication outcomes for all autistic children.
Read moreUsing Forced Alignment for Phonetics Research
Longitudinal changes of automated speech markers in MCI and mild AD
Abstract BackgroundMild cognitive impairment (MCI) and mild Alzheimer’s disease (mAD) are characterized by decline in memory and other cognitive domains, including language. Previous studies have shown that MCI and mAD patients produce impaired speech. In this study, we investigated longitudinal changes in lexical and acoustic speech features from natural speech of MCI and mAD patients using highly reliable automated methods.MethodWe analyzed 116 digitized 1‐minute picture descriptions produced by 24 MCI (65.8±9.2y, 9 females, mean mini‐mental state exam (MMSE) score at the first recording (T1)=25.4±1.1) and 19 mAD patients (66.9±8.8y, 5 females, mean MMSE at T1=21.9±1.1). The groups did not differ in demographic characteristics, disease duration, and mean inter‐sample interval. Using automated pipelines, we tagged the part‐of‐speech categories of all words, rated words for word frequency, familiarity, concreteness, and semantic ambiguity, and counted partial words and total number of words. Speech samples were segmented into speech and silent pause segments to calculate mean speech segment duration, pause rate, percent of speech out of total time, and articulation rate. Linear mixed‐effects models examined changes over time for each measure, controlling for disease duration at T1. The difference between the first and the last recordings (mean=24±10.9 months) in each measure was related to changes in MMSE, covarying for disease duration at T1.ResultPatients’ adjectives (beta=‐0.04, p=0.006) and determiners (beta=‐0.06, p=0.01) decreased over time, whereas pronouns increased (beta=0.06, p=0.014). Patients produced more frequent (beta=0.004, p=0.043) and more familiar (beta=0.001, p=0.046) words over time. Patients articulated more slowly (beta=‐0.02, p=0.001) and produced more partial words (beta=0.04, p=0.004) and fewer total words (beta=‐0.9, p=0.018) over time. Patients produced shorter speech segments (beta=0.008, p=0.021), paused more frequently (beta=0.32, p=0.001), and produced less speech (beta=‐0.35, p<0.001) over time. All measures showed no group interaction. Patients whose MMSE scores decreased spoke more slowly (beta=2.25, p=0.04) and produced higher‐frequency words (beta=‐6.4, p=0.043) and more partial words (beta=‐0.72, p=0.028) compared with patients whose MMSE scores did not change.ConclusionMCI and mAD patients’ language changed over time, and this was significantly related to patients’ cognitive decline. These findings support the use of automated speech analyses in studying longitudinal changes in neurodegeneration.
Read moreJoint Coreference Resolution for Zeros and non-Zeros in Arabic
Most existing proposals about anaphoric zero pronoun (AZP) resolution regard full mention coreference and AZP resolution as two independent tasks, even though the two tasks are clearly related. The main issues that need tackling to develop a joint model for zero and non-zero mentions are the difference between the two types of arguments (zero pronouns, being null, provide no nominal information) and the lack of annotated datasets of a suitable size in which both types of arguments are annotated for languages other than Chinese and Japanese. In this paper, we introduce two architectures for jointly resolving AZPs and non-AZPs, and evaluate them on Arabic, a language for which, as far as we know, there has been no prior work on joint resolution. Doing this also required creating a new version of the Arabic subset of the standard coreference resolution dataset used for the CoNLL-2012 shared task (Pradhan et al.,2012) in which both zeros and non-zeros are included in a single dataset.
Read more