• Home
  • Search
  • Do Transformer Interpretability Methods Transfer to RNNs?
  • https://doi.org/10.1609/aaai.v39i26.34969Copy DOI Icon

Do Transformer Interpretability Methods Transfer to RNNs?

Show More
  • Abstract
  • Literature Map
  • Similar Papers
Abstract

Recent advances in recurrent neural network architectures, such as Mamba and RWKV, have enabled RNNs to match or exceed the performance of equal-size transformers in terms of language modeling perplexity and downstream evaluations, suggesting that future systems may be built on completely new architectures. In this paper, we examine if selected interpretability methods originally designed for transformer language models will transfer to these up-and-coming recurrent architectures. Specifically, we focus on steering model outputs via contrastive activation addition, on eliciting latent predictions via the tuned lens, and eliciting latent knowledge from models fine-tuned to produce false outputs under certain conditions. Our results show that most of these techniques are effective when applied to RNNs, and we show that it is possible to improve some of them by taking advantage of RNNs' compressed state.

Similar Papers
  • Conference Article
  • Citations8

Learning based compact thermal modeling for energy-efficient smart building management

  • Nov 01, 2015
  • Hengyang Zhao +6
  • Conference Article

Retraction Notice: Evaluating Recurrent Neural Network Architectures for Natural Language Understanding

  • Jun 24, 2024
  • Davendra Kumar Doda +2
  • PDF
  • Research Article
  • Citations1

Lyapunov-guided representation of recurrent neural network performance

  • Jul 02, 2024
  • Neural Computing and Applications
  • Ryan Vogt +2
  • Research Article
  • Citations13

AI Content Generation Technology based on Open AI Language Model

  • Dec 18, 2023
  • Journal of Artificial Intelligence and Capsule Networks
  • Sangita Pokhrel +1
  • Preprint Article

Estimating Gross Primary Production via Recurrent Neural Networks: A comparative analysis

  • Nov 27, 2024
  • David Montero +8
  • Conference Article
  • Citations3

Which Neural Network Architecture matches Human Behavior in Artificial Grammar Learning?

  • Jan 01, 2019
  • Andrea Alamia +3
  • Research Article
  • Citations173

Representation of Linguistic Form and Function in Recurrent Neural Networks

  • Dec 01, 2017
  • Computational Linguistics
  • Ákos Kádár +2
  • Research Article
  • Citations76

Analysis of surface ozone using a recurrent neural network

  • Feb 11, 2015
  • Science of The Total Environment
  • Fabio Biancofiore +8
  • Book Chapter

Implementation of Text Classification Model Based on Recurrent Neural Networks

  • Jan 01, 2020
  • Ming-Shi Wang +1
  • Book Chapter
  • Citations22

Efficient evolution of asymmetric recurrent neural networks using a PDGP-inspired two-dimensional representation

  • Jan 01, 1998
  • João Carlos Figueira Pujol +1
  • Conference Article
  • Citations31

Gated Recurrent Neural Tensor Network

  • Jul 01, 2016
  • Andros Tjandra +4
  • Research Article

Continuous Human activity recognition using pure Recurrent Neural Network (RNN) architecture and IOT

  • Dec 01, 2025
  • Mid-West University Journal of Engineering & Innovation
  • Dhirendra Kumar Yadav +1
  • Conference Article
  • Citations3

Abnormal Signaling SIP Dialogs Detection based on Deep Learning

  • Apr 01, 2021
  • Diogo Pereira +2
  • Book Chapter
  • Citations162

A Comparison of LSTM and GRU Networks for Learning Symbolic Sequences

  • Jan 01, 2023
  • Roberto Cahuantzi +2
  • Research Article
  • Citations21

A Contextual Recurrent Collaborative Filtering framework for modelling sequences of venue checkins

  • Sep 04, 2019
  • Information Processing & Management
  • Jarana Manotumruksa +2
Cactus Communications logo

Copyright 2026 Cactus Communications. All rights reserved.