• Home
  • Search
  • Refining Evaluation Functions for Game 2048 by Extended Temporal Difference Learning
  • https://doi.org/10.1109/tg.2025.3606517Copy DOI Icon

Refining Evaluation Functions for Game 2048 by Extended Temporal Difference Learning

Show More
  • Abstract
  • Literature Map
  • References
  • Similar Papers
Abstract

The game 2048 has attracted millions of people with its simple yet challenging gameplay, leading to the development of numerous computer players. Most successful computer players for 2048 use evaluation functions trained through temporal difference learning (TD learning) or its variants. While TD learning is highly effective and can improve evaluation functions quickly, the performance of these functions often plateaus after a certain number of timesteps. Therefore, it is important to refine those evaluation functions to further enhance the performance of computer players. In this paper, we extend the conventional TD learning approach and propose two refinement algorithms for 2048. Firstly, we conducted detailed experiments to refine the best open-source neural network, and achieved significant performance improvements, increasing the average score from <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$2.49 \times 10^{5}$</tex-math></inline-formula> to <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$3.37 \times 10^{5}$</tex-math></inline-formula> in greedy play (1-ply lookahead) and from <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$4.87 \times 10^{5}$</tex-math></inline-formula> to <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$5.45 \times 10^{5}$</tex-math></inline-formula> with 3-ply expectimax search. We also applied our refinement method to the state-of-the-art N-tuple network, improving the average score from <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$5.85\times 10^{5}$</tex-math></inline-formula> to <inline-formula xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink"><tex-math notation="LaTeX">$6.10\times 10^{5}$</tex-math></inline-formula> with 6-ply expectimax search and the tile-downgrading trick.

Similar Papers
  • Book Chapter
  • Citations28

Reinforcement Learning with N-tuples on the Game Connect-4

  • Jan 01, 2012
  • Markus Thill +2
  • Conference Article
  • Citations1

Automatic Generation of Evaluation Features for Computer Game Players

  • Jan 01, 2007
  • Makoto Miwa +2
  • PDF
  • Research Article

Exploring the link between temporal difference learning and spike-timing-dependent plasticity

  • Jul 13, 2009
  • BMC Neuroscience
  • Cătălin V Rusu +1
  • Research Article

Digitising a Turn-Based Strategy Board Game: Implementation of Spellcaster Using the Mini-Max Algorithm

  • Nov 24, 2024
  • International Journal of Computer Science and Information Technology
  • Yukai Wang
  • Conference Article

Virtual Mittweida - Creating a game-based approach to teach artificial intelligence for games

  • Jan 01, 2025
  • AHFE international
  • Alexander Thomas Kühn +3
  • Conference Article

Dyna-like reinforcement learning based on accumulative and average rewards

  • Oct 01, 2010
  • Kao-Shing Hwang +1
  • Conference Article
  • Citations1

Design of Amazon Chess Evaluation Function Based on Reinforcement Learning

  • Jun 01, 2019
  • Bo Jianbo +4
  • Conference Article

A Soft Sensing Method Based on the Temporal Difference Learning Algorithm

  • Jan 01, 2006
  • Tao Ye +2
  • Research Article
  • Citations269

Effects of a social regulation-based online learning framework on students’ learning achievements and behaviors in mathematics

  • Oct 05, 2020
  • Computers &amp; Education
  • Gwo-Jen Hwang +2
  • Research Article
  • Citations4

Analyzing Movements of Tennis Players by Dynamic Image Processing

  • Jan 01, 2007
  • IEEJ Transactions on Electronics, Information and Systems
  • Shunpei Yazaki +1
  • Dissertation

Bayesian partially observable reinforcement learning

  • Jan 01, 2023
  • Sammie Katt
  • Research Article
  • Citations28

Contextual Learning Approach and Performance Assessment in Mathematics Learning

  • May 15, 2016
  • International Research Journal of Management, IT &amp; Social Sciences
  • I Wayan Eka Mahendra
  • PDF
  • Research Article
  • Citations2

CONTEXTUAL LEARNING APPROACH AND PERFORMANCE ASSESSMENT IN MATHEMATICS LEARNING

  • Jun 01, 2015
  • JISAE: JOURNAL OF INDONESIAN STUDENT ASSESMENT AND EVALUATION
  • I Wayan Eka Mahendra
  • Research Article

Stereotypical love: a cluster analysis of self-presentation strategies in tinder profile pictures.

  • Sep 29, 2025
  • The journal of sexual medicine
  • Alejandro García-Alamán +2
  • Research Article
  • Citations124

Ant-TD: Ant colony optimization plus temporal difference reinforcement learning for multi-label feature selection

  • Apr 27, 2021
  • Swarm and Evolutionary Computation
  • Mohsen Paniri +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.