• Home
  • Search
  • When blockchain meets AI: Optimal mining strategy achieved by machine learning
  • Cite Icon64
  • https://doi.org/10.1002/int.22375Copy DOI Icon

When blockchain meets AI: Optimal mining strategy achieved by machine learning

Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

This study applies reinforcement learning (RL) from the AI machine learning field to derive an optimal Bitcoin-like blockchain mining strategy. A salient feature of the RL learning framework is that an optimal (or near-optimal) strategy can be obtained without knowing the details of the blockchain network model. Previously, the most profitable mining strategy was believed to be honest mining encoded in the default blockchain protocol. It was shown later that it is possible to gain more mining rewards by deviating from honest mining. In particular, the mining problem can be formulated as a Markov Decision Process (MDP) which can be solved to give the optimal mining strategy. However, solving the mining MDP requires knowing the values of various parameters that characterize the blockchain network model. In real blockchain networks, these parameter values are not easy to obtain and may change over time. This hinders the use of the MDP model-based solution. In this study, we employ RL to dynamically learn a mining strategy with performance approaching that of the optimal mining strategy. Since the mining MDP problem has a nonlinear objective function (rather than linear functions of standard MDP problems), we design a new multidimensional RL algorithm to solve the problem. Experimental results indicate that, without knowing the parameter values of the mining MDP model, our multidimensional RL mining algorithm can still achieve optimal performance over time-varying blockchain networks.

Similar Papers
  • Book Chapter
  • Citations10

Consensus Algorithm Analysis in Blockchain: PoW and Raft

  • Oct 22, 2021
  • Taotao Wang +2
  • PDF
  • Research Article
  • Citations1

To Exit or Not to Exit: Cost-Effective Early-Exit Architecture Based on Markov Decision Process

  • Jul 19, 2024
  • Mathematics
  • Kyu-Sik Kim +1
  • Research Article
  • Citations13

Reinforcement Learning for Clinical Applications.

  • Feb 08, 2023
  • Clinical Journal of the American Society of Nephrology
  • Kia Khezeli +5
  • Research Article
  • Citations14

Peer-to-peer electricity transaction decision of user-side smart energy system based on SARSA reinforcement learning method

  • Jan 01, 2020
  • CSEE Journal of Power and Energy Systems
  • Dan Wang +5
  • Research Article

Goal Agnostic Learning and Planning without Reward Functions

  • Jan 01, 2023
  • Advances in Artificial Intelligence and Machine Learning
  • Christopher Robinson +1
  • Research Article
  • Citations28

Exploring reinforcement learning in process control: a comprehensive survey

  • Feb 28, 2025
  • International Journal of Systems Science
  • N Rajasekhar +2
  • Components
  • Citations7

Reward-predictive representations generalize across tasks in reinforcement learning

  • Oct 15, 2020
  • Lucas Lehnert +3
  • Research Article
  • Citations26

Probably Approximately Correct (PAC) exploration in reinforcement learning

  • Jan 01, 2007
  • Rutgers University Community Repository (Rutgers University)
  • Alexander L Strehl
  • Conference Article
  • Citations75

State aggregation in Markov decision processes

  • Dec 10, 2002
  • Zhiyuan Ren +1
  • Research Article
  • Citations20

Rebalancing the car-sharing system with reinforcement learning

  • Apr 20, 2020
  • World Wide Web
  • Changwei Ren +4
  • Research Article
  • Citations3

전력손실 최소화를 위한 심층 강화학습 기반 배전계통 재구성

  • Nov 30, 2020
  • The transactions of The Korean Institute of Electrical Engineers
  • Se-Heon Lim +2
  • Book Chapter
  • Citations17

Solving Markov Decision Processes via Simulation

  • Sep 18, 2014
  • Abhijit Gosavi
  • Research Article
  • Citations124

Ant-TD: Ant colony optimization plus temporal difference reinforcement learning for multi-label feature selection

  • Apr 27, 2021
  • Swarm and Evolutionary Computation
  • Mohsen Paniri +2
  • PDF
  • Research Article
  • Citations23

A Countermeasure Against Random Pulse Jamming in Time Domain Based on Reinforcement Learning

  • Jan 01, 2020
  • IEEE Access
  • Quan Zhou +2
  • Research Article
  • Citations306

Advanced planning for autonomous vehicles using reinforcement learning and deep inverse reinforcement learning

  • Jan 15, 2019
  • Robotics and Autonomous Systems
  • Changxi You +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.