<feed xmlns="http://www.w3.org/2005/Atom"> <id>/</id><title>Amagibaba</title><subtitle>I write about Mechanistic Interpretability and Statistics.</subtitle> <updated>2026-02-07T21:11:32-08:00</updated> <author> <name>Amagibaba</name> <uri>/</uri> </author><link rel="self" type="application/atom+xml" href="/feed.xml"/><link rel="alternate" type="text/html" hreflang="en" href="/"/> <generator uri="https://jekyllrb.com/" version="4.3.4">Jekyll</generator> <rights> © 2026 Amagibaba </rights> <icon>/assets/img/favicons/favicon.ico</icon> <logo>/assets/img/favicons/favicon-96x96.png</logo> <entry><title>Complexity Series (2 / 3) - Carry-Over Circuits</title><link href="/posts/carry-over-circuits/" rel="alternate" type="text/html" title="Complexity Series (2 / 3) - Carry-Over Circuits" /><published>2026-02-05T21:00:00-08:00</published> <updated>2026-02-05T21:00:00-08:00</updated> <id>/posts/carry-over-circuits/</id> <content src="/posts/carry-over-circuits/" /> <author> <name>Amagibaba</name> </author> <category term="mechanistic interpretability" /> <summary> This is the second installment of the “Complexity Series,” where I endeavor to argue that there are certain classes of problems (simple arithmetic being one of them) that Transformer architectures will never be able to solve. In my previous post, I outlined that we have thus far discovered a few algorithms that are being implicitly implemented by LLMs to perform arithmetic (regular or modulo):... </summary> </entry> <entry><title>Complexity Series (1 / 3) - LLM Arithmetic &amp;#x2B50;</title><link href="/posts/llm-arithmetic/" rel="alternate" type="text/html" title="Complexity Series (1 / 3) - LLM Arithmetic &amp;#x2B50;" /><published>2025-11-29T21:00:00-08:00</published> <updated>2025-12-19T09:58:20-08:00</updated> <id>/posts/llm-arithmetic/</id> <content src="/posts/llm-arithmetic/" /> <author> <name>Amagibaba</name> </author> <category term="mechanistic interpretability" /> <summary> This is the first installment of the “Complexity Series”, where I endeavor to argue that there are certain classes of problems (simple arithmetic being one of them) that Transformer architectures will never be able to solve. To get started, I want to do a deep dive on known algorithms that Transformers have used to perform simple arithmetic. There are thus far a few algorithms that researchers... </summary> </entry> <entry><title>Multi-Layer Latent Space Visualization</title><link href="/posts/multi-layer-latent-spaces/" rel="alternate" type="text/html" title="Multi-Layer Latent Space Visualization" /><published>2025-10-01T22:00:00-07:00</published> <updated>2026-01-07T22:44:31-08:00</updated> <id>/posts/multi-layer-latent-spaces/</id> <content src="/posts/multi-layer-latent-spaces/" /> <author> <name>Amagibaba</name> </author> <category term="mechanistic interpretability" /> <summary> In MLPs, What Input Gives What Output? If you had a series of MLP (nn.Linear) layers chained together, you may like to answer the question: if I wanted the final layer outputs’ first element (call $h_1$) to be activated (i.e. $&amp;gt; 0$), what would my input have to be? To set the stage, one may ask: who cares? Well, wouldn’t it be great if you could answer: I want a model’s response to my one-... </summary> </entry> <entry><title>ICA vs SAEs &amp;#x2B50;</title><link href="/posts/ica/" rel="alternate" type="text/html" title="ICA vs SAEs &amp;#x2B50;" /><published>2025-05-12T22:00:00-07:00</published> <updated>2025-08-24T23:33:11-07:00</updated> <id>/posts/ica/</id> <content src="/posts/ica/" /> <author> <name>Amagibaba</name> </author> <category term="statistics" /> <summary> Thinking about how Sparse Auto-encoders (SAEs) aim to learn a sparse over-complete basis (where you are trying to triangulate a larger number of sources than you have signals; e.g. you only have 8 microphones in the room, but there are 20 speakers) got me thinking about Independent Component Analysis again. In particular, I wanted to see if I could articulate a mapping between ICA and SAEs. Thi... </summary> </entry> <entry><title>Thoughts on Hidden Structure in MLP Space</title><link href="/posts/tegum-factors/" rel="alternate" type="text/html" title="Thoughts on Hidden Structure in MLP Space" /><published>2025-04-27T22:00:00-07:00</published> <updated>2025-06-15T23:09:17-07:00</updated> <id>/posts/tegum-factors/</id> <content src="/posts/tegum-factors/" /> <author> <name>Amagibaba</name> </author> <category term="mechanistic interpretability" /> <summary> After deep-diving into why SAEs succeed at retrieving superposed features, what their limitations are, and closely inspecting the hidden technical implementations of the sae_lens library, I just wanted to write some quick notes of what I think the MLP space looks like, with respect to ‘true’ features. Anticorrelated (“Mutually Sparse”) Features Prefer to be in the Same Tegum Factor This hypot... </summary> </entry> </feed>
