• Home
  • Search
  • TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts
  • Cite Icon8
  • https://doi.org/10.18653/v1/2024.emnlp-main.667Copy DOI Icon

TheoremLlama: Transforming General-Purpose LLMs into Lean4 Experts

  • Jan 1, 2024
  • Ruida Wang +6 more
Show More
  • Abstract
  • Literature Map
  • Citations
  • Similar Papers
Abstract

Proving mathematical theorems using computer-verifiable Formal Languages (FL) like Lean significantly impacts mathematical reasoning.One approach to formal theorem proving involves generating complete proofs using Large Language Models (LLMs) based on Natural Language (NL) proofs.However, due to the scarcity of aligned NL and FL theorem-proving data, most modern LLMs exhibit suboptimal performance.This scarcity results in a paucity of methodologies for training LLMs and techniques to fully utilize their capabilities in composing formal proofs.To address these challenges, this paper proposes TheoremLlama, an end-to-end framework that trains a general-purpose LLM to be a Lean4 expert.TheoremLlama includes NL-FL dataset generation and bootstrapping method to obtain aligned dataset, curriculum learning and block training techniques to train the model, and iterative proof writing method to write Lean4 proofs that work together synergistically.Using the dataset generation method in TheoremLlama, we provide Open Bootstrapped Theorems (OBT), an NL-FL aligned and bootstrapped dataset.Our novel NL-FL bootstrapping method, where NL proofs are integrated into Lean4 code for training datasets, leverages the NL reasoning ability of LLMs for formal reasoning.The TheoremLlama framework achieves cumulative accuracies of 36.48% and 33.61% on MiniF2F-Valid and Test datasets respectively, surpassing the GPT-4 baseline of 22.95% and 25.41%.Our code, model checkpoints, and the generated dataset is published in GitHub

Similar Papers
  • Research Article

Evaluating gpt-4 for zero-shot classification of bleeding and clotting events: Can large language models serve as second reviewers?

  • Nov 03, 2025
  • Blood
  • Samantha Rizzo +5
  • Research Article
  • Citations51

Large language models for biomedicine: foundations, opportunities, challenges, and best practices.

  • Apr 24, 2024
  • Journal of the American Medical Informatics Association : JAMIA
  • Satya S Sahoo +8
  • Research Article
  • Citations1

Logical and Physical Optimizations for SQL Query Execution over Large Language Models

  • Jun 17, 2025
  • Proceedings of the ACM on Management of Data
  • Dario Satriani +5
  • Research Article
  • Citations67

The life cycle of large language models in education: A framework for understanding sources of bias

  • Jul 12, 2024
  • British Journal of Educational Technology
  • Jinsook Lee +4
  • Research Article

#2924 Comparison of large language models and traditional natural language processing techniques in predicting arteriovenous fistula failure

  • May 23, 2024
  • Nephrology Dialysis Transplantation
  • Suman Lama +6
  • Research Article

Can LLMs effectively provide game-theoretic-based scenarios for cybersecurity?

  • Dec 11, 2025
  • Frontiers in Computer Science
  • Daniele Proverbio +5
  • Conference Article
  • Citations5

Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation

  • Oct 08, 2024
  • David Austin +3
  • Conference Article

Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural Integration

  • Oct 27, 2025
  • Yicheng Pan +8
  • Research Article
  • Citations20

DialogueLLM: Context and emotion knowledge-tuned large language models for emotion recognition in conversations.

  • Dec 01, 2025
  • Neural networks : the official journal of the International Neural Network Society
  • Yazhou Zhang +6
  • Research Article
  • Citations1

Large Language Model Powered Symbolic Execution

  • Oct 09, 2025
  • Proceedings of the ACM on Programming Languages
  • Yihe Li +2
  • PDF
  • Research Article
  • Citations13

Advancements and Applications of Large Language Models in Natural Language Processing: A Comprehensive Review

  • Nov 26, 2024
  • Applied and Computational Engineering
  • Mengchao Ren
  • Research Article
  • Citations1

Mathematical Problem Solving in Arabic: Assessing Large Language Models

  • Jan 01, 2024
  • Procedia Computer Science
  • Abeer Mahgoub +2
  • Research Article
  • Citations2

Enhancing Decision Making Through the Integration of Large Language Models and Operations Research Optimization

  • Apr 11, 2025
  • Proceedings of the AAAI Conference on Artificial Intelligence
  • Segev Wasserkrug +6
  • Research Article
  • Citations127

Interactive computer-aided diagnosis on medical image using large language models

  • Sep 17, 2024
  • Communications Engineering
  • Sheng Wang +5
  • Research Article
  • Citations2

Programming Chatbots Using Natural Language: Generating Cervical Spine MRI Impressions.

  • Sep 14, 2024
  • Cureus
  • Ramin Javan +6
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.