---
date: '2024-01-30'
description: well, here we go again.
id: Turing-complete Transformers
modified: 2026-06-05 15:08:22 GMT-04:00
tags:
  - pattern
  - ml
title: Turing-complete Transformers
created: '2024-01-30'
published: '2024-01-30'
pageLayout: default
slug: thoughts/Turing-complete-Transformers
permalink: https://aarnphm.xyz/thoughts/Turing-complete-Transformers.md
generator:
  quartz: v4.6.0
  hostedProvider: Cloudflare
  baseUrl: aarnphm.xyz
full: https://aarnphm.xyz/llms-full.txt
---
<article class="twitter-post" data-twitter-source="https://twitter.com/burny_tech/status/1744100637187461455"><span class="twitter-post-badge" role="img" aria-label="Twitter post" title="Twitter post"><svg xmlns="http://www.w3.org/2000/svg" width="16" height="16" viewBox="64 64 896 896" fill="currentColor" stroke="none" stroke-width="0" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true" focusable="false"><path d="M928 254.3c-30.6 13.2-63.9 22.7-98.2 26.4a170.1 170.1 0 0075-94 336.64 336.64 0 01-108.2 41.2A170.1 170.1 0 00672 174c-94.5 0-170.5 76.6-170.5 170.6 0 13.2 1.6 26.4 4.2 39.1-141.5-7.4-267.7-75-351.6-178.5a169.32 169.32 0 00-23.2 86.1c0 59.2 30.1 111.4 76 142.1a172 172 0 01-77.1-21.7v2.1c0 82.9 58.6 151.6 136.7 167.4a180.6 180.6 0 01-44.9 5.8c-11.1 0-21.6-1.1-32.2-2.6C211 652 273.9 701.1 348.8 702.7c-58.6 45.9-132 72.9-211.7 72.9-14.3 0-27.5-.5-41.2-2.1C171.5 822 261.2 850 357.8 850 671.4 850 843 590.2 843 364.7c0-7.4 0-14.8-.5-22.2 33.2-24.3 62.3-54.4 85.5-88.2z"></path></svg></span><header class="twitter-post-header"><span class="twitter-post-author">@burny_tech</span><time datetime="2024-01-07">2024-01-07</time><a class="twitter-post-source" href="https://twitter.com/burny_tech/status/1744100637187461455" target="_blank" rel="noopener noreferrer" aria-label="Open original post (opens in a new tab)" title="Open original post" data-skip-icons="true"><svg xmlns="http://www.w3.org/2000/svg" width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="1.5" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true" focusable="false"><path d="M7 17 17 7M7 7h10v10"></path></svg></a></header><div class="twitter-post-body">
			
				
					<p>Turing Complete Transformers: Two Transformers Are More Powerful Than One<br>"We prove transformers are not Turing complete, propose a new architecture that is Turing complete, and empirically demonstrate that the new architecture can generalize more effectively than transformers."<br>"This paper presents Find+Replace transformers, a family of multi-transformer architectures that can provably do things no single transformer can, and which outperforms GPT-4 on several challenging tasks. We first establish that traditional transformers and similar architectures are not Turing Complete, while Find+Replace transformers are. Using this fact, we show how arbitrary programs can be compiled into Find+Replace transformers, potentially aiding interpretability research. We also demonstrate the superior performance of Find+Replace transformers over GPT-4 on a set of composition challenge problems. This work aims to provide a theoretical basis for multi-transformer architectures, and to encourage their further exploration."</p>
<img src="https://pbs.twimg.com/media/GDRKQ04W0AAuRDh.jpg?name=orig" alt="" loading="lazy" decoding="async">
				
			
			
		</div></article>

The idea is to combine two small [[thoughts/Transformers|transformers]] rather than one [[thoughts/large models]]

More specialised on given tasks, and prove to be Turing-complete?

![[posts/images/shogoth-gpt.webp|Shogoth as GPTs]]

> Speculatively, people might think GPT-4 without any guardrails _could_ pass the Turing-test. A more important question is “What is the Turing-test equivalent for pseudo-intelligence system?”

John Searle famously said:

<blockquote class="quotes"><p>Turing machine is not to be found in nature. They're to be found in our <em>interpretations</em> of nature.</p><p>John Searle</p></blockquote>

the [paper](https://openreview.net/forum?id=MGWsPGogLH) detailed what is very similar to the setup of [recursive language model](https://alexzhang13.github.io/blog/2025/rlm/), ideally we want to deal with long-horizon tasks more effectively and economically viable.

