Transluce - Research https://transluce.org/research Research updates from Transluce http://www.rssboard.org/rss-specification python-feedgen en Sat, 11 Jul 2026 02:19:42 +0000 Toward A Public Science of Model Behavior https://transluce.org/behavior-science Our vision for an open scientific ecosystem for measuring model behaviors https://transluce.org/behavior-science Thu, 09 Jul 2026 00:00:00 +0000 Predictive Concept Decoders https://transluce.org/pcd Training scalable end-to-end interpretability assistants https://transluce.org/pcd Thu, 18 Dec 2025 00:00:00 +0000 Scalably Extracting Latent Representations of Users https://transluce.org/user-modeling Constructing datasets and training decoders to extract user models from language models https://transluce.org/user-modeling Tue, 25 Nov 2025 00:00:00 +0000 Language Model Circuits Are Sparse in the Neuron Basis https://transluce.org/neuron-circuits A new technique for tracing sparse and faithful circuits directly on a model's MLPs https://transluce.org/neuron-circuits Thu, 20 Nov 2025 00:00:00 +0000 Monitoring SWE-bench Agents https://transluce.org/docent/blog/swe-bench Partnering with SWE-bench to enable reliable monitoring of AI coding agents https://transluce.org/docent/blog/swe-bench Wed, 19 Nov 2025 00:00:00 +0000 Training Language Models to Explain Their Own Computations https://transluce.org/self-explanations We trained explainer models to verbalize the content of AI systems' internal computations https://transluce.org/self-explanations Tue, 11 Nov 2025 00:00:00 +0000 Automatically Jailbreaking Frontier Language Models with Investigator Agents https://transluce.org/jailbreaking-frontier-models Discovering cost-effective attacks with reinforcement learning https://transluce.org/jailbreaking-frontier-models Wed, 03 Sep 2025 00:00:00 +0000 Surfacing Pathological Behaviors in Language Models https://transluce.org/pathological-behaviors Improving our investigator agents with propensity bounds https://transluce.org/pathological-behaviors Thu, 05 Jun 2025 00:00:00 +0000 Investigating Truthfulness in a Pre-Release o3 Model https://transluce.org/investigating-o3-truthfulness o3 frequently fabricates actions it took to fulfill user requests, and elaborately justifies the fabrications when confronted https://transluce.org/investigating-o3-truthfulness Wed, 16 Apr 2025 00:00:00 +0000 Introducing Docent https://transluce.org/introducing-docent A system for analyzing and intervening on agent behavior https://transluce.org/introducing-docent Mon, 24 Mar 2025 00:00:00 +0000 Monitor: An AI-Driven Observability Interface https://transluce.org/observability-interface An interface designed to help humans observe, understand, and steer computations inside models https://transluce.org/observability-interface Wed, 23 Oct 2024 00:00:00 +0000 Eliciting Language Model Behaviors with Investigator Agents https://transluce.org/automated-elicitation Language models trained to automatically surface harmful behaviors in language models https://transluce.org/automated-elicitation Wed, 23 Oct 2024 00:00:00 +0000 Scaling Automatic Neuron Explanation https://transluce.org/neuron-descriptions Open-source AI systems trained to describe components of other AI systems at the level of a human expert https://transluce.org/neuron-descriptions Wed, 23 Oct 2024 00:00:00 +0000