• Home
  • Search
  • JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
  • https://doi.org/10.1609/aaai.v40i7.37443Copy DOI Icon

JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics

  • Mar 14, 2026
  • Simindokht Jahangard +4 more
Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Recent advances in Vision-Language Models (VLMs) and large language models (LLMs) have greatly enhanced visual reasoning, a key capability for embodied AI agents like robots. However, existing visual reasoning benchmarks often suffer from several limitations: they lack a clear definition of reasoning complexity, offer have no control to generate questions over varying difficulty and task customization, and fail to provide structured, step-by-step reasoning annotations (workflows). To bridge these gaps, we formalize reasoning complexity, introduce an adaptive query engine that generates customizable questions of varying complexity with detailed intermediate annotations, and extend the JRDB dataset with human-object interaction and geometric relationship annotations to create JRDB-Reasoning, a benchmark tailored for visual reasoning in human-crowded environments. Our engine and benchmark enable fine-grained evaluation of visual reasoning frameworks and dynamic assessment of visual-language models across reasoning levels.

Similar Papers
  • Conference Article

Multimodal Analysis of Google Bard: Experiments in Visual Reasoning

  • Nov 25, 2023
  • David Noever +1
  • Research Article

Comparative Evaluation of ChatGPT-4o, Gemini 2.5 Pro and Grok-4 in Answering Orthodontics Questions from the Dentistry Specialty Examination

  • Jan 01, 2026
  • Lokman Hekim Health Sciences
  • Samet Özden
  • Research Article
  • Citations1

GPT-4o and OpenAI o1 Performance on the 2024 Spanish Competitive Medical Specialty Access Examination: Cross-Sectional Quantitative Evaluation Study

  • Jan 12, 2026
  • JMIR Medical Education
  • Pau Benito +9
  • Conference Article

Hybrid-NL2SVA: Integrating RAG and Finetuning for LLM-based NL2SVA

  • Sep 08, 2025
  • Weihua Xiao +3
  • Research Article
  • Citations8

Evaluating ChatGPT's performance across radiology subspecialties: A meta-analysis of board-style examination accuracy and variability.

  • Sep 01, 2025
  • Clinical imaging
  • Dan Nguyen +2
  • Research Article
  • Citations2

Preliminary assessment of large language models' performance in answering questions on developmental dysplasia of the hip.

  • Apr 15, 2025
  • Journal of children's orthopaedics
  • Shiwei Li +2
  • Research Article
  • Citations62

A rapid review on current and potential uses of large language models in nursing

  • Mar 13, 2024
  • International journal of nursing studies
  • Mollie Hobensack +13
  • Research Article
  • Citations3

Off-the-Shelf Large Language Models for Causality Assessment of Individual Case Safety Reports: A Proof-of-Concept with COVID-19 Vaccines

  • Jan 01, 2025
  • Drug Safety
  • Andrea Abate +8
  • Research Article
  • Citations44

Assessment of fine-tuned large language models for real-world chemistry and material science applications†

  • Jan 01, 2025
  • Chemical Science
  • Joren Van Herck +49
  • Research Article
  • Citations19

Assessment of Large Language Models in Cataract Care Information Provision: A Quantitative Comparison

  • Nov 08, 2024
  • Ophthalmology and Therapy
  • Zichang Su +5
  • Research Article

Feasibility and exploratory assessment of large language models for pediatric dentistry queries: a comparative study.

  • Apr 24, 2026
  • Frontiers in oral health
  • Sanjeev B Khanagar +7
  • PDF
  • Research Article
  • Citations18

Enhancing Robot Task Planning and Execution through Multi-Layer Large Language Models.

  • Mar 06, 2024
  • Sensors
  • Zhirong Luan +6
  • Conference Article
  • Citations1

An improved human-object interaction detection method based on short-term memory selection network

  • Nov 10, 2020
  • Chang Wang +1
  • Supplementary Content

Affordances, Challenges, and Opportunities of ChatGPT in Mathematics Education: A Scoping Review

  • Aug 24, 2025
  • Phuong Bui +3
  • Research Article
  • Citations20

Mini-Gemini: Mining the Potential of Multi-Modality Vision Language Models.

  • Mar 01, 2026
  • IEEE transactions on pattern analysis and machine intelligence
  • Yanwei Li +7
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.