• Home
  • Search
  • OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution
  • Cite Icon1
  • https://doi.org/10.1145/3728871Copy DOI Icon

OmniGIRL: A Multilingual and Multimodal Benchmark for GitHub Issue Resolution

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

The GitHub issue resolution task aims to resolve issues reported in repositories automatically. With advances in large language models (LLMs), this task has gained increasing attention, and several benchmarks are proposed to evaluate the issue resolution ability of LLMs. However, existing benchmarks have three main limitations. First, current benchmarks focus on a single programming language, limiting the evaluation of issues from repositories across different languages. Second, they usually cover a narrow range of domains, which may fail to represent the diversity of real-world issues. Third, existing benchmarks rely solely on textual information in issue descriptions, overlooking multimodal information such as images in issues. In this paper, we propose OmniGIRL, a GitHub Issue ResoLution benchmark that is multilingual, multimodal, and multi-domain. OmniGIRL includes 959 task instances, which are collected from repositories across four programming languages (i.e., Python, JavaScript, TypeScript, and Java) and eight different domains. Our evaluation shows that current LLMs show limited performances on OmniGIRL. Notably, the best-performing model, GPT-4o, resolves only 8.6% of the issues. Besides, we find that current LLMs struggle to resolve issues requiring understanding images. The best performance is achieved by Claude-3.5-Sonnet, which resolves only 10.5% of the issues with image information. Finally, we analyze the reasons behind current LLMs’ failure on OmniGIRL, providing insights for future improvements.

Similar Papers
  • Research Article

Research and selection of Large Learning Models for automation of ABAP-code migration

  • Sep 24, 2025
  • Management of Development of Complex Systems
  • Oleg Pozdnyakov +1
  • Research Article

Methodology for Creating a Benchmark to Evaluate LLM Performance on Numerals

  • Dec 04, 2025
  • Информатика и автоматизация
  • Sergey Karpovich +2
  • Conference Article
  • Citations129

Benchmarking Large Language Models for Automated Verilog RTL Code Generation

  • Apr 01, 2023
  • Shailja Thakur +7
  • Research Article
  • Citations6

ProtTeX: Structure-In-Context Reasoning and Editing of Proteins with Large Language Models.

  • Jun 25, 2025
  • Journal of chemical information and modeling
  • Zicheng Ma +9
  • Research Article
  • Citations1

Optimization of traditional methods for determining the similarity of project names and purchases using large language models

  • Apr 01, 2024
  • Litera
  • Aleksei Aleksandrovich Golikov +2
  • Research Article

Assessing the Latent Automated Program Repair Capabilities of Large Language Models using Round-Trip Translation

  • Oct 16, 2025
  • ACM Transactions on Software Engineering and Methodology
  • Fernando Vallecillos Ruiz +3
  • Research Article
  • Citations2

Towards Secure Code Generation With LLMs: A Study on Common Weakness Enumeration

  • Dec 01, 2025
  • IEEE Transactions on Software Engineering
  • Jianguo Zhao +6
  • Research Article
  • Citations12

FELLAS: Enhancing Federated Sequential Recommendation with LLM as External Services

  • Sep 12, 2025
  • ACM Transactions on Information Systems
  • Wei Yuan +5
  • Book Chapter
  • Citations1

Image Information Prompt: Tips for Learning Large Language Models

  • Jan 31, 2024
  • Yin Zhang
  • Supplementary Content

Automating Code Generation for a New Ecosystem: Establishing Baselines with Large Language Model Based Code Generation for ArkTS and HarmonyOS

  • Sep 04, 2025
  • Research Square
  • Mehmet Cem Aytekin +2
  • Conference Article
  • Citations16

HOLMES: Hyper-Relational Knowledge Graphs for Multi-hop Question Answering using LLMs

  • Jan 01, 2024
  • Pranoy Panda +4
  • Research Article

Analyzing LLM-generated code according to four ISO/IEC 5055:2021 categories

  • Jan 01, 2025
  • IEEE Access
  • Rasmus Krebs +1
  • Research Article
  • Citations1

Porting Software Libraries to OpenHarmony: Transitioning from TypeScript or JavaScript to ArkTS

  • Jun 22, 2025
  • Proceedings of the ACM on Software Engineering
  • Bo Zhou +6
  • Research Article
  • Citations11

Denoising Alignment with Large Language Model for Recommendation

  • Jan 24, 2025
  • ACM Transactions on Information Systems
  • Yingtao Peng +7
  • Research Article
  • Citations1

CUPCase: Clinically Uncommon Patient Cases and Diagnoses Dataset

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Oriel Perets +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.