• Home
  • Search
  • Evaluating multimodal commercial and open-source large language models for dynamical astronomy: a benchmark study of resonant behavior classification.
  • https://doi.org/10.1038/s41598-026-45926-yCopy DOI Icon

Evaluating multimodal commercial and open-source large language models for dynamical astronomy: a benchmark study of resonant behavior classification.

Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

We present a systematic evaluation of modern multimodal large language models (LLMs) for the classification of mean-motion and secular resonances from images of resonant arguments. Four benchmark datasets (RB-TEST, RB-PILOT, RB-SMALL, RB-FULL) were constructed to cover clear, ambiguous, and transient cases, with both binary and three-class outputs. Using standardized prompts (a full prompt for large models and a simplified variant for small models that cannot process complex instructions), we tested flagship commercial models, large open-source models, and small locally runnable models. Commercial LLMs reach [Formula: see text] on simple cases and up to [Formula: see text] on the three-class RB-SMALL dataset, while the best open-source models also reach [Formula: see text] on unambiguous cases and [Formula: see text] on the complex ones. On the full binary benchmark, open-source models approach commercial performance ([Formula: see text]-[Formula: see text]). Most errors occur in transient and resonance-sticking regimes. The results show that LLMs can perform resonance classification at levels comparable to those of classical or machine-learning methods without training or fine-tuning, and that even small open-source models achieve practically useful accuracy. The released benchmarks establish a reproducible standard for evaluating LLMs on dynamical astronomy tasks.

Similar Papers
  • PDF
  • Research Article
  • Citations26

Large language models for error detection in radiology reports: a comparative analysis between closed-source and privacy-compliant open-source models

  • Feb 20, 2025
  • European Radiology
  • Babak Salam +12
  • Research Article
  • Citations45

Custom Large Language Models Improve Accuracy: Comparing Retrieval Augmented Generation and Artificial Intelligence Agents to Non-Custom Models for Evidence-Based Medicine

  • Nov 07, 2024
  • Arthroscopy: The Journal of Arthroscopic and Related Surgery
  • Joshua J Woo +7
  • PDF
  • Conference Article
  • Citations2

Test Large Language Models on Driving Theory Knowledge and Skills for Connected Autonomous Vehicles

  • Nov 18, 2024
  • Zuoyin Tang +5
  • Research Article
  • Citations1

AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights

  • Oct 15, 2025
  • Proceedings of the AAAI/ACM Conference on AI Ethics and Society
  • Jiannan Xu +2
  • PDF
  • Research Article
  • Citations8

Transforming education: tackling the two sigma problem with AI in journal clubs – a proof of concept

  • May 08, 2025
  • BDJ Open
  • Fahad Umer +4
  • Research Article

A systematic literature review of large language models in phishing attack generation and detection

  • Jul 01, 2026
  • Array
  • Dinushan Sivaneswaran +5
  • Conference Article
  • Citations22

The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks

  • Dec 02, 2024
  • Xiaoyi Chen +9
  • Research Article

Performance of LLMs on VITA test: potential for AI-assisted tax returns for low income taxpayers

  • Jul 08, 2025
  • Artificial Intelligence and Law
  • Sina Gogani-Khiabani +3
  • Conference Article

Are LLMs Good Annotators for Discourse-level Event Relation Extraction?

  • Jan 01, 2024
  • Kangda Wei +2
  • Preprint Article

Quality of Automatic Speech Recognition -- Polish Language case study -- from Wav2Vec to Scribe ElevenLabs

  • Feb 25, 2026
  • arXiv (Cornell University)
  • Pietroń, Marcin +8
  • Research Article

Humans and Large Language Models in Clinical Decision Support: A Study with Medical Calculators.

  • Jan 01, 2024
  • AMIA ... Annual Symposium proceedings. AMIA Symposium
  • Nicholas C Wan +10
  • Research Article

Diagnostic accuracy of large language models in the classification of superior labial frenulum attachments.

  • Dec 10, 2025
  • Odontology
  • Mehmet Gümüş Kanmaz +1
  • Conference Article
  • Citations3

On the use of Large Language Models to Detect Brazilian Politics Fake News

  • Nov 17, 2024
  • Marcos P S Gôlo +6
  • Research Article

Lightweight open-source large language models versus cTAKES for information extraction from discharge summaries: tobacco smoking status test case

  • Jan 13, 2026
  • JAMIA Open
  • David M Dávila-García +2
  • Research Article
  • Citations1

Automating Air Pollution Map Analysis with Multi-Modal AI and Visual Context Engineering

  • Dec 19, 2025
  • Atmosphere
  • Szymon Cogiel +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.