• Cite Icon1
  • https://doi.org/10.1002/9781394272976.ch11Copy DOI Icon

Building a Language Model from Scratch

  • Dec 13, 2024
  • Edward Dongbo Cui
Show More
  • Abstract
  • Literature Map
  • References
  • Citations
  • Similar Papers
Abstract

This chapter presents the evolution of language modeling using the transformer architecture to illustrate how machine learning models, specifically, large language models like LLaMA (Large Language Model Meta AI), came to be. The sinusoidal position information is hard-coded. However, it is also possible to learn to represent the position information. An approach to learning to represent positional information is to treat positions as a categorical feature. Then, through back-propagation, the network learns to efficiently encode positional information. Normalization can help stabilize neural network training by reducing the likelihood of exploding or vanishing gradients and improving model generalization. In the original transformer paper, the authors used position encodings added on top of the input embedding vectors of each word token to allow the attention model to disambiguate words from different positions along the sequence.

Similar Papers
  • Conference Article
  • Citations41

Transliteration Based Approaches to Improve Code-Switched Speech Recognition Performance

  • Dec 01, 2018
  • Jesse Emond +4
  • Research Article
  • Citations18

A fast and memory-efficient N-gram language model lookup method for large vocabulary continuous speech recognition

  • Dec 19, 2005
  • Computer Speech & Language
  • Xiaolong Li +1
  • Research Article

Blending Language Models and Domain-Specific Languages in Computer Science Education. A Case Study on API RESTFul

  • Oct 03, 2025
  • International Journal of Interactive Multimedia and Artificial Intelligence
  • Francisco Jurado +3
  • Research Article
  • Citations5

A method to build a super small but practically accurate language model for handheld devices

  • Nov 01, 2003
  • Journal of Computer Science and Technology
  • Genqing Wu +1
  • Conference Article
  • Citations1

Syllable-based Myanmar language model for speech recognition

  • Jun 01, 2015
  • Wunna Soe +1
  • Conference Article
  • Citations3

Improvement of Embedded Human-Machine Interfaces Combining Language, Hypothesis and Error Models

  • Aug 01, 2011
  • Juan-Carlos Perez-Cortes +3
  • Preprint Article

Modelling Implicit Bias in Gender–Career Associations: A systematic comparison of language models

  • May 22, 2025
  • Alexander Porshnev +5
  • PDF
  • Conference Article
  • Citations62

Grammar Induction with Neural Language Models: An Unusual Replication

  • Jan 01, 2018
  • Phu Mon Htut +2
  • Research Article
  • Citations21

Japanese large-vocabulary continuous-speech recognition using a newspaper corpus and broadcast news

  • Jun 01, 1999
  • Speech Communication
  • Katsutoshi Ohtsuki +6
  • Book Chapter
  • Citations1

Planning with Logical Graph-Based Language Model for Instruction Generation

  • Oct 16, 2024
  • Frontiers in artificial intelligence and applications
  • Fan Zhang +2
  • Research Article
  • Citations28

Reward modeling for mitigating toxicity in transformer-based language models

  • Jul 20, 2022
  • Applied Intelligence
  • Farshid Faal +2
  • Research Article
  • Citations48

A study of n-gram and decision tree letter language modeling methods

  • Jun 01, 1998
  • Speech Communication
  • Gerasimos Potamianos +1
  • Conference Article
  • Citations4

Whose Emotions and Moral Sentiments do Language Models Reflect?

  • Jan 01, 2024
  • Zihao He +3
  • Conference Article
  • Citations11

Hybrid word/Part-of-Arabic-Word Language Models for arabic text document recognition

  • Aug 01, 2015
  • Mohamed Faouzi Benzeghiba +2
  • PDF
  • Research Article
  • Citations965

How Can We Know What Language Models Know?

  • Dec 01, 2020
  • Transactions of the Association for Computational Linguistics
  • Zhengbao Jiang +3
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.