- Research Article
- 10.1016/j.neucom.2026.133330
Improving clustering quality evaluation in noisy Gaussian mixtures
- Jun 01, 2026
- Neurocomputing
- Renato Cordeiro De Amorim + 1 more +1
Publications from 2021 to 2026
Showing 10 of 219 papers
Improving clustering quality evaluation in noisy Gaussian mixtures
Leveraging Differentiable Climate-Economy Models for Hybrid Modeling and Inverse Problems
Robust carbon cycle science and effective carbon market governance depend on accurate monitoring, transparent modelling and credible representation of climate–economic feedbacks. Integrated Assessment Models (IAMs) such as RICE provide a long-standing framework for linking carbon emissions, climate dynamics and economic development and are widely used to inform mitigation pathways, carbon pricing and international climate policy. However, traditional IAMs rely on hand-calibrated parameters, simplified damage functions and fixed ethical assumptions, limiting their ability to integrate observational data, quantify uncertainty and support evidence-based carbon management. We build on recent advances in machine learning for climate policy and introduce RICE-N-JAX, a fully differentiable implementation of the multi-region RICE-N model (Zhang et al., 2025). RICE-N extends classical IAMs with multi-agent reinforcement learning to model strategic interactions and international climate negotiations. Our JAX-based reimplementation makes the entire climate–economic simulation fast and differentiable, including carbon emissions, climate response, production, trade, mitigation decisions and negotiation dynamics. Differentiability enables a new class of hybrid, data-driven climate–economic models. Our current research focuses on two key directions. First, we develop non-parametric hybrid damage functions in which the traditional analytical damage formulation is replaced by neural or spline-based surrogates trained on empirical and scenario data. This allows the damage–temperature relationship to be learned directly from data. Second, we perform inverse modelling of ethical and behavioural parameters, such as regional risk aversion, time preferences and mitigation bias, by calibrating the model against emissions, GDP and temperature trajectories from the Shared Socioeconomic Pathways (SSPs). This enables the recovery of latent normative assumptions embedded in scenario narratives and provides a data-informed basis for policy analysis. Finally, differentiability supports gradient-based calibration, uncertainty quantification, and sensitivity analysis of carbon price trajectories, mitigation pathways, and long-term climate impacts. We demonstrate a proof-of-concept end-to-end calibration of climate damage functions and show how parameter uncertainty propagates into future economic and emissions outcomes. By bridging process-based climate–economic theory with hybrid, knowledge-guided machine learning, RICE-N-JAX provides a foundation for fast and data-driven carbon-cycle modelling. The framework supports policy-relevant applications ranging from carbon pricing and climate clubs to carbon market design, illustrating how hybrid ML can strengthen the scientific basis of carbon management and climate mitigation.
Read moreIdentifying and Analyzing Performance-Critical Tokens in Large Language Models
In-context learning (ICL) has emerged as an effective solution for few-shot learning with large language models (LLMs). However, how LLMs leverage demonstrations to specify a task and learn a corresponding computational function through ICL is underexplored. Drawing from the way humans learn from content-label mappings in demonstrations, we categorize the tokens in an ICL prompt into content, stopword, and template tokens. Our goal is to identify the types of tokens whose representations directly influence LLM's performance, a property we refer to as being performance-critical. By ablating representations from the attention of the test example, we find that the representations of informative content tokens have less influence on performance compared to template and stopword tokens, which contrasts with the human attention to informative words. We give evidence that the representations of performance-critical tokens aggregate information from the content tokens. Moreover, we demonstrate experimentally that lexical meaning, repetition, and structural cues are the main distinguishing characteristics of these tokens. Our work sheds light on how LLMs learn to perform tasks from demonstrations and deepens our understanding of the roles different types of tokens play in LLMs.
Read moreWhat to Ask Next? Probing the Imaginative Reasoning of LLMs with TurtleSoup Puzzles
We investigate the capacity of Large Language Models (LLMs) for imaginative reasoning—the proactive construction, testing, and revision of hypotheses in information-sparse environments. Existing benchmarks, often static or focused on social deduction, fail to capture the dynamic, exploratory nature of this reasoning process. To address this gap, we introduce a comprehensive research framework based on the classic "Turtle Soup" game, integrating a benchmark, an agent, and an evaluation protocol. We present TurtleSoup-Bench, the first large-scale, bilingual, interactive benchmark for imaginative reasoning, comprising 800 turtle soup stories sourced from both the Internet and expert authors. We also propose Mosaic-Agent, a novel agent designed to assess LLMs' performance in this setting. To evaluate reasoning quality, we develop a multi-dimensional protocol measuring logical consistency, detail completion, and conclusion alignment. Experiments with leading LLMs reveal clear capability limits, common failure patterns, and a significant performance gap compared to humans. Our work offers new insights into LLMs' imaginative reasoning and establishes a foundation for future research on exploratory agent behavior.
Read moreEvolutionarily conserved neural dynamics across mice, monkeys, and humans.
On evolutionary timescales, brain circuits adapt to support survival in each species' ecological niche. While some anatomical aspects of neural circuitry are conserved across species with distant evolutionary origins, each species also exhibits specific circuit adaptations that enable its behavioral repertoire. It remains unclear whether homologous brain regions leverage analogous neural computations as different species perform common behaviors such as reaching and manipulating objects. Here, we directly assessed conservation of neural computations using intracortical recordings from mouse, monkey, and human motor cortex-a homologous region across many mammals-during motor behaviors crucial for survival. We hypothesized that, despite their phylogenetic distance, rodents and primates produce movements through conserved neural computations implemented by motor cortical population dynamics. Remarkably, we found that movement-related neural dynamics were highly conserved across species, while variations in behavioral output were uniquely captured in neural trajectory geometries. Strikingly, neural dynamics during movement across species were more conserved than those across brain regions in the same human and between motor preparation and execution in the same monkeys. Lastly, through manipulation of neural network models trained to perform reaching movements, we reinforce that conservation of neural dynamics across species likely stems from shared circuit constraints. We thus assert that evolution maintains neural computations across phylogeny even as behavioral repertoires expand.
Read moreA Chemical-Genetic Interaction Matrix Reveals Drug Mechanism and Genetic Architecture
To probe drug mechanism of action (MOA) and interrogate the genetic architecture of human cells, we carried out isogenic genome-wide CRISPR/Cas9 knockout screens against 310 diverse drugs, bioactive compounds, and stress conditions. Stringent statistical correction for gene knockout fitness defects yielded a large-scale matrix of >12,000 high confidence chemical-genetic interactions (CGIs). This dataset revealed many previously unappreciated off-target effects for well-characterized compounds and novel MOAs for uncharacterized compounds. The CGI matrix uncovered dense genetic modules that yielded new biological insights into phospholipidosis, mitotic regulation, metabolism, the DNA damage response, and mTOR signaling. The dataset allowed identification of multi-drug sensitization and resistance mechanisms, inference of gene function, elaboration of cross-process connectivity, evaluation of the cell type specificity of CGIs, prediction of chemical synergism, and extensive annotation of understudied genes. This resource provides a map of the genetic landscape in human cells and a framework to help guide drug discovery.
Read moreConformal Selection for Efficient and Accurate Compound Screening in Drug Discovery.
Reliable compound screening is fundamental to drug discovery, yet the process remains undermined by lack of robust risk controls of false compound selection or omission in current methods. To address these challenges, we introduced conformal selection as an enhanced approach to optimize the compound screening process with balanced risks and benefits. Leveraging conformal inference, our approach constructs p-values for each candidate molecule to quantify statistical evidence for selection. The final selection of molecules is determined by comparing these p-values against thresholds derived from multiple testing principles. Our approach offers rigorous control over the false discovery/omission rate, ensuring validity independent of data set size and requiring minimal assumptions. By avoiding the estimation of prediction errors required in previous approaches, our method achieves higher power, thereby improving the ability to identify promising candidates. We validate these advantages through numerical simulations on real-world data sets.
Read moreLarge Language Model Applications in the Algebra Domain: A Systematic Review
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
Semantic and syntactic bootstrapping posit that children use their prior knowledge of one linguistic domain, say syntactic relations, to help later acquire another, such as the meanings of new words. Empirical results supporting both theories may tempt us to believe that these are different independent learning strategies. Here, we argue for a unified approach, where instead they are both contingent on a more general learning strategy for language acquisition: joint learning. Using a series of neural visually-grounded grammar induction models, we demonstrate that both syntactic and semantic bootstrapping effects are strongest when syntax and semantics are learnt simultaneously via joint learning. This more general learning strategy results in better grammar induction, realistic lexical category learning, and better interpretations of novel sentence and verb meanings. Joint learning makes language acquisition easier for learners by mutually constraining the hypotheses spaces for both syntax and semantics. Studying the dynamics of joint inference over many input sources and modalities represents an important new direction for language modeling and learning research in both cognitive sciences and AI, as it may help us explain how language can be acquired in more constrained learning settings.
Read moreTrustVis: A Multi-Dimensional Trustworthiness Evaluation Framework for Large Language Models
As Large Language Models (LLMs) continue to revolutionize Natural Language Processing (NLP) applications, critical concerns about their trustworthiness persist, particularly in safety and robustness. To address these challenges, we introduce TrustVis, an automated evaluation framework that provides a comprehensive assessment of LLM trustworthiness. A key feature of our framework is its interactive user interface, designed to offer intuitive visualizations of trustworthiness metrics. By integrating well-known perturbation methods like AutoDAN and employing majority voting across various evaluation methods, TrustVis not only provides reliable results but also makes complex evaluation processes accessible to users. Preliminary case studies on models like Vicuna-7b, Llama2-7b, and GPT-3.5 demonstrate the effectiveness of our framework in identifying safety and robustness vulnerabilities, while the interactive interface allows users to explore results in detail, empowering targeted model improvements. Video Link: https://youtu.be/k1TrBqNVg8g
Read more