Tilde Research https://tilderesearch.com/research Research publications from Tilde Research http://www.rssboard.org/rss-specification python-feedgen en Fri, 31 Jul 2026 03:30:23 +0000 Online KL Shampoo https://blog.tilderesearch.com/blog/online-kl-shampoo We introduce Online KL Shampoo (OKLS), a zero-staleness, hardware-aligned, Kronecker product-based optimizer that achieves 1.45x Muon's parameter efficiency while maintaining 98% training throughput. https://blog.tilderesearch.com/blog/online-kl-shampoo Technical Release Tue, 28 Jul 2026 00:00:00 +0000 Parallax: Parameterized Local Linear Attention https://blog.tilderesearch.com/blog/parallax Parallax keeps the local-linear principle but parametrizes the solve directly → a single additive correction on top of Softmax. https://blog.tilderesearch.com/blog/parallax Technical Release Tue, 09 Jun 2026 00:00:00 +0000 Towards Compositional Steepest Descent https://blog.tilderesearch.com/blog/compositional-muon We introduce Compositional Muon (CM), which extends Muon-style steepest descent from individual matrices to composed transformer circuits https://blog.tilderesearch.com/blog/compositional-muon Technical Release Fri, 05 Jun 2026 00:00:00 +0000 Wall Attention: Length Generalization With Diagonal Gates https://blog.tilderesearch.com/blog/wall-attn We introduce Wall Attention, which generalizes diagonal forget gates from linear RNNs to softmax attention, yielding a data-dependent positional encoding that is a full replacement for RoPE. https://blog.tilderesearch.com/blog/wall-attn Technical Release Tue, 02 Jun 2026 00:00:00 +0000 Aurora: Leverage-Aware Updates for Rectangular Matrices https://blog.tilderesearch.com/blog/aurora We formulate the problem of steepest descent under the joint constraint of row-norm uniformity and orthogonality and present Aurora as a solution. https://blog.tilderesearch.com/blog/aurora Technical Release Fri, 08 May 2026 00:00:00 +0000 Nitrobrew: Fast, Lossless Distillation for Free https://blog.tilderesearch.com/blog/nitrobrew Nitrobrew exploits the fact that teacher logits are generated from a much lower-dimensional hidden state, achieving 1.5-3x faster end-to-end throughput for on-policy distillation. https://blog.tilderesearch.com/blog/nitrobrew Technical Release Tue, 28 Apr 2026 00:00:00 +0000 Gram-Space Manifold Muon https://blog.tilderesearch.com/vignettes/gram-space Recently, Bernstein and Thinking Machines introduced a Muon variant that constrains weights to the Stiefel manifold. https://blog.tilderesearch.com/vignettes/gram-space Vignette Mon, 13 Oct 2025 00:00:00 +0000 Complexity and Transformers https://blog.tilderesearch.com/vignettes/complexity Can a simple transformer solve problems that constant-depth circuits cannot? https://blog.tilderesearch.com/vignettes/complexity Vignette Mon, 15 Sep 2025 00:00:00 +0000 Regression is All You Need https://blog.tilderesearch.com/vignettes/regression There are many derivations of attention available. In this vignette we present our favorite from nonparametric regression. https://blog.tilderesearch.com/vignettes/regression Vignette Thu, 28 Aug 2025 00:00:00 +0000 A Telescope, Pointed at the Stars https://blog.tilderesearch.com/blog/telescope-stars We believe mechanistic understanding is the foundation for entirely new architectures, capabilities, and safety tools. https://blog.tilderesearch.com/blog/telescope-stars Announcement Sun, 10 Aug 2025 00:00:00 +0000 MoMoE: Memory-optimized Mixture of Experts https://blog.tilderesearch.com/blog/momoe Mixture-of-Experts (MoE) architectures scale model parameter counts efficiently by activating only a fraction of weights per token. https://blog.tilderesearch.com/blog/momoe Technical Release Fri, 25 Jul 2025 00:00:00 +0000 Sparsity is Cool https://blog.tilderesearch.com/blog/sparse-attn Sparse attention models have recently emerged as strong competitors to base attention. https://blog.tilderesearch.com/blog/sparse-attn Vignette Wed, 25 Jun 2025 00:00:00 +0000 Activault: Scalable, Efficient, and Fast Model Activation Storage https://blog.tilderesearch.com/blog/activault Mixture-of-Experts (MoE) architectures scale model parameter counts efficiently by activating only a fraction of weights per token. https://blog.tilderesearch.com/blog/activault Technical Release Mon, 17 Mar 2025 00:00:00 +0000 Sieve: SAEs Beat Baselines on a Real-World Task (A Code Generation Case Study) https://blog.tilderesearch.com/blog/sieve We demonstrate the first application of SAE-based interventions to a downstream task that outperforms classical baselines with minimal effort/time. https://blog.tilderesearch.com/blog/sieve Technical Release Sun, 15 Dec 2024 00:00:00 +0000 The Rate Distortion Dance of Sparse Autoencoders https://blog.tilderesearch.com/blog/rate-distortion-saes Mixture-of-Experts (MoE) architectures scale model parameter counts efficiently by activating only a fraction of weights per token. https://blog.tilderesearch.com/blog/rate-distortion-saes Vignette Tue, 12 Nov 2024 00:00:00 +0000