Anthropic Research https://www.anthropic.com/research Anthropic Research http://www.rssboard.org/rss-specification python-feedgen en Sat, 22 Aug 2026 22:49:58 +0000 A General Language Assistant as a Laboratory for Alignment https://www.anthropic.com/research/a-general-language-assistant-as-a-laboratory-for-alignment https://www.anthropic.com/research/a-general-language-assistant-as-a-laboratory-for-alignment Wed, 01 Dec 2021 08:00:00 +0000 A Mathematical Framework for Transformer Circuits https://www.anthropic.com/research/a-mathematical-framework-for-transformer-circuits https://www.anthropic.com/research/a-mathematical-framework-for-transformer-circuits Wed, 22 Dec 2021 08:00:00 +0000 Predictability and Surprise in Large Generative Models https://www.anthropic.com/research/predictability-and-surprise-in-large-generative-models Large models have predictable loss via scaling laws but unpredictable capabilities. This tension has significant policy implications. https://www.anthropic.com/research/predictability-and-surprise-in-large-generative-models Tue, 15 Feb 2022 08:00:00 +0000 In-context Learning and Induction Heads https://www.anthropic.com/research/in-context-learning-and-induction-heads https://www.anthropic.com/research/in-context-learning-and-induction-heads Tue, 08 Mar 2022 08:00:00 +0000 Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback https://www.anthropic.com/research/training-a-helpful-and-harmless-assistant-with-reinforcement-learning-from-human-feedback https://www.anthropic.com/research/training-a-helpful-and-harmless-assistant-with-reinforcement-learning-from-human-feedback Tue, 12 Apr 2022 07:00:00 +0000 Scaling Laws and Interpretability of Learning from Repeated Data https://www.anthropic.com/research/scaling-laws-and-interpretability-of-learning-from-repeated-data https://www.anthropic.com/research/scaling-laws-and-interpretability-of-learning-from-repeated-data Sat, 21 May 2022 07:00:00 +0000 Softmax Linear Units https://www.anthropic.com/research/softmax-linear-units https://www.anthropic.com/research/softmax-linear-units Fri, 17 Jun 2022 07:00:00 +0000 Language Models (Mostly) Know What They Know https://www.anthropic.com/research/language-models-mostly-know-what-they-know https://www.anthropic.com/research/language-models-mostly-know-what-they-know Mon, 11 Jul 2022 07:00:00 +0000 Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned https://www.anthropic.com/research/red-teaming-language-models-to-reduce-harms-methods-scaling-behaviors-and-lessons-learned https://www.anthropic.com/research/red-teaming-language-models-to-reduce-harms-methods-scaling-behaviors-and-lessons-learned Mon, 22 Aug 2022 07:00:00 +0000 Toy Models of Superposition https://www.anthropic.com/research/toy-models-of-superposition Neural networks pack many concepts into single neurons. This paper shows how and when models represent more features than they have dimensions. https://www.anthropic.com/research/toy-models-of-superposition Wed, 14 Sep 2022 07:00:00 +0000 Measuring Progress on Scalable Oversight for Large Language Models https://www.anthropic.com/research/measuring-progress-on-scalable-oversight-for-large-language-models https://www.anthropic.com/research/measuring-progress-on-scalable-oversight-for-large-language-models Fri, 04 Nov 2022 07:00:00 +0000 Constitutional AI: Harmlessness from AI Feedback https://www.anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedback https://www.anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedback Thu, 15 Dec 2022 08:00:00 +0000 Discovering Language Model Behaviors with Model-Written Evaluations https://www.anthropic.com/research/discovering-language-model-behaviors-with-model-written-evaluations https://www.anthropic.com/research/discovering-language-model-behaviors-with-model-written-evaluations Mon, 19 Dec 2022 08:00:00 +0000 Superposition, Memorization, and Double Descent https://www.anthropic.com/research/superposition-memorization-and-double-descent https://www.anthropic.com/research/superposition-memorization-and-double-descent Thu, 05 Jan 2023 08:00:00 +0000 The Capacity for Moral Self-Correction in Large Language Models https://www.anthropic.com/research/the-capacity-for-moral-self-correction-in-large-language-models https://www.anthropic.com/research/the-capacity-for-moral-self-correction-in-large-language-models Wed, 15 Feb 2023 08:00:00 +0000 Privileged Bases in the Transformer Residual Stream https://www.anthropic.com/research/privileged-bases-in-the-transformer-residual-stream https://www.anthropic.com/research/privileged-bases-in-the-transformer-residual-stream Thu, 16 Mar 2023 07:00:00 +0000 Distributed Representations: Composition & Superposition https://www.anthropic.com/research/distributed-representations-composition-superposition https://www.anthropic.com/research/distributed-representations-composition-superposition Thu, 04 May 2023 07:00:00 +0000 Interpretability Dreams https://www.anthropic.com/research/interpretability-dreams https://www.anthropic.com/research/interpretability-dreams Wed, 24 May 2023 07:00:00 +0000 Circuits Updates — May 2023 https://www.anthropic.com/research/circuits-updates-may-2023 https://www.anthropic.com/research/circuits-updates-may-2023 Wed, 24 May 2023 07:00:00 +0000 Towards Measuring the Representation of Subjective Global Opinions in Language Models https://www.anthropic.com/research/towards-measuring-the-representation-of-subjective-global-opinions-in-language-models https://www.anthropic.com/research/towards-measuring-the-representation-of-subjective-global-opinions-in-language-models Thu, 29 Jun 2023 07:00:00 +0000 Question Decomposition Improves the Faithfulness of Model-Generated Reasoning https://www.anthropic.com/research/question-decomposition-improves-the-faithfulness-of-model-generated-reasoning https://www.anthropic.com/research/question-decomposition-improves-the-faithfulness-of-model-generated-reasoning Tue, 18 Jul 2023 07:00:00 +0000 Measuring Faithfulness in Chain-of-Thought Reasoning https://www.anthropic.com/research/measuring-faithfulness-in-chain-of-thought-reasoning https://www.anthropic.com/research/measuring-faithfulness-in-chain-of-thought-reasoning Tue, 18 Jul 2023 07:00:00 +0000 Studying Large Language Model Generalization with Influence Functions https://www.anthropic.com/research/studying-large-language-model-generalization-with-influence-functions https://www.anthropic.com/research/studying-large-language-model-generalization-with-influence-functions Tue, 08 Aug 2023 07:00:00 +0000 Tracing Model Outputs to the Training Data https://www.anthropic.com/research/influence-functions https://www.anthropic.com/research/influence-functions Tue, 08 Aug 2023 07:00:00 +0000 Challenges in evaluating AI systems https://www.anthropic.com/research/evaluating-ai-systems https://www.anthropic.com/research/evaluating-ai-systems Wed, 04 Oct 2023 14:00:00 +0000 Towards Monosemanticity: Decomposing Language Models With Dictionary Learning https://www.anthropic.com/research/towards-monosemanticity-decomposing-language-models-with-dictionary-learning https://www.anthropic.com/research/towards-monosemanticity-decomposing-language-models-with-dictionary-learning Thu, 05 Oct 2023 07:00:00 +0000 Decomposing Language Models Into Understandable Components https://www.anthropic.com/research/decomposing-language-models-into-understandable-components https://www.anthropic.com/research/decomposing-language-models-into-understandable-components Thu, 05 Oct 2023 07:00:00 +0000 Collective Constitutional AI: Aligning a Language Model with Public Input https://www.anthropic.com/research/collective-constitutional-ai-aligning-a-language-model-with-public-input Anthropic and the Collective Intelligence Project ran a public process with ~1,000 Americans to draft a constitution for an AI system, then trained a model on it. https://www.anthropic.com/research/collective-constitutional-ai-aligning-a-language-model-with-public-input Tue, 17 Oct 2023 07:00:00 +0000 Towards Understanding Sycophancy in Language Models https://www.anthropic.com/research/towards-understanding-sycophancy-in-language-models https://www.anthropic.com/research/towards-understanding-sycophancy-in-language-models Mon, 23 Oct 2023 18:57:00 +0000 Specific versus General Principles for Constitutional AI https://www.anthropic.com/research/specific-versus-general-principles-for-constitutional-ai https://www.anthropic.com/research/specific-versus-general-principles-for-constitutional-ai Tue, 24 Oct 2023 16:07:00 +0000 Evaluating and Mitigating Discrimination in Language Model Decisions https://www.anthropic.com/research/evaluating-and-mitigating-discrimination-in-language-model-decisions https://www.anthropic.com/research/evaluating-and-mitigating-discrimination-in-language-model-decisions Thu, 07 Dec 2023 16:50:00 +0000 Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training https://www.anthropic.com/research/sleeper-agents-training-deceptive-llms-that-persist-through-safety-training https://www.anthropic.com/research/sleeper-agents-training-deceptive-llms-that-persist-through-safety-training Sun, 14 Jan 2024 22:10:00 +0000 Reflections on Qualitative Research https://www.anthropic.com/research/transformer-circuits https://www.anthropic.com/research/transformer-circuits Fri, 08 Mar 2024 16:00:00 +0000 Many-shot jailbreaking https://www.anthropic.com/research/many-shot-jailbreaking https://www.anthropic.com/research/many-shot-jailbreaking Tue, 02 Apr 2024 16:05:09 +0000 Measuring the persuasiveness of language models https://www.anthropic.com/research/measuring-model-persuasiveness https://www.anthropic.com/research/measuring-model-persuasiveness Tue, 09 Apr 2024 15:55:00 +0000 Simple probes can catch sleeper agents https://www.anthropic.com/research/probes-catch-sleeper-agents https://www.anthropic.com/research/probes-catch-sleeper-agents Tue, 23 Apr 2024 09:42:44 +0000 Circuits Updates – April 2024 https://www.anthropic.com/research/circuits-updates-april-2024 https://www.anthropic.com/research/circuits-updates-april-2024 Fri, 26 Apr 2024 20:24:00 +0000 Mapping the mind of a large language model https://www.anthropic.com/research/mapping-mind-language-model https://www.anthropic.com/research/mapping-mind-language-model Tue, 21 May 2024 07:00:00 +0000 Testing and mitigating elections-related risks https://www.anthropic.com/research/testing-and-mitigating-elections-related-risks https://www.anthropic.com/research/testing-and-mitigating-elections-related-risks Thu, 06 Jun 2024 13:00:00 +0000 Claude’s Character https://www.anthropic.com/research/claude-character Claude 3 was the first model with "character training"—alignment aimed at nurturing traits like curiosity, open-mindedness, and thoughtfulness. https://www.anthropic.com/research/claude-character Sat, 08 Jun 2024 18:27:00 +0000 The engineering challenges of scaling interpretability https://www.anthropic.com/research/engineering-challenges-interpretability https://www.anthropic.com/research/engineering-challenges-interpretability Thu, 13 Jun 2024 08:00:00 +0000 Sycophancy to subterfuge: Investigating reward tampering in language models https://www.anthropic.com/research/reward-tampering Can minor specification gaming evolve into more dangerous behaviors? This paper demonstrates that models trained on low-level reward hacking—like sycophancy—can generalize to tampering with their own reward functions, even covering their tracks. The behavior emerged without explicit training, and common safety techniques reduced but didn't eliminate it. https://www.anthropic.com/research/reward-tampering Mon, 17 Jun 2024 13:10:48 +0000 Circuits Updates – June 2024 https://www.anthropic.com/research/circuits-updates-june-2024 https://www.anthropic.com/research/circuits-updates-june-2024 Fri, 28 Jun 2024 21:43:09 +0000 Circuits Updates – July 2024 https://www.anthropic.com/research/circuits-updates-july-2024 https://www.anthropic.com/research/circuits-updates-july-2024 Wed, 31 Jul 2024 19:21:00 +0000 Circuits Updates – August 2024 https://www.anthropic.com/research/circuits-updates-august-2024 https://www.anthropic.com/research/circuits-updates-august-2024 Fri, 06 Sep 2024 15:36:00 +0000 Circuits Updates – September 2024 https://www.anthropic.com/research/circuits-updates-sept-2024 https://www.anthropic.com/research/circuits-updates-sept-2024 Tue, 01 Oct 2024 09:38:00 +0000 Using dictionary learning features as classifiers https://www.anthropic.com/research/features-as-classifiers https://www.anthropic.com/research/features-as-classifiers Wed, 16 Oct 2024 23:49:00 +0000 Sabotage evaluations for frontier models https://www.anthropic.com/research/sabotage-evaluations https://www.anthropic.com/research/sabotage-evaluations Fri, 18 Oct 2024 16:55:00 +0000 Developing a computer use model https://www.anthropic.com/research/developing-computer-use https://www.anthropic.com/research/developing-computer-use Tue, 22 Oct 2024 19:42:00 +0000 Evaluating feature steering: A case study in mitigating social biases https://www.anthropic.com/research/evaluating-feature-steering https://www.anthropic.com/research/evaluating-feature-steering Fri, 25 Oct 2024 13:42:00 +0000 Raising the bar on SWE-bench Verified with Claude 3.5 Sonnet https://www.anthropic.com/research/swe-bench-sonnet https://www.anthropic.com/research/swe-bench-sonnet Wed, 30 Oct 2024 23:25:00 +0000 A statistical approach to model evaluations https://www.anthropic.com/research/statistical-approach-to-model-evals https://www.anthropic.com/research/statistical-approach-to-model-evals Tue, 19 Nov 2024 16:11:00 +0000 Clio: A system for privacy-preserving insights into real-world AI use https://www.anthropic.com/research/clio https://www.anthropic.com/research/clio Thu, 12 Dec 2024 13:08:00 +0000 Alignment faking in large language models https://www.anthropic.com/research/alignment-faking This paper provides the first empirical example of a model engaging in alignment faking without being trained to do so—selectively complying with training objectives while strategically preserving existing preferences. https://www.anthropic.com/research/alignment-faking Wed, 18 Dec 2024 14:16:00 +0000 Building effective agents https://www.anthropic.com/research/building-effective-agents https://www.anthropic.com/research/building-effective-agents Thu, 19 Dec 2024 21:14:00 +0000 Constitutional Classifiers: Defending against universal jailbreaks https://www.anthropic.com/research/constitutional-classifiers These classifiers filter the overwhelming majority of jailbreaks while maintaining practical deployment. A prototype withstood over 3,000 hours of red teaming with no universal jailbreak discovered. https://www.anthropic.com/research/constitutional-classifiers Mon, 03 Feb 2025 12:35:00 +0000 The Anthropic Economic Index https://www.anthropic.com/research/the-anthropic-economic-index The Anthropic Economic Index analyzes millions of Claude.ai conversations showing AI usage concentrated in software development and technical writing, touching 25%+ of tasks in 36% of occupations. AI leans toward augmentation (57%) over automation (43%). https://www.anthropic.com/research/the-anthropic-economic-index Mon, 10 Feb 2025 13:00:00 +0000 Insights on Crosscoder Model Diffing https://www.anthropic.com/research/crosscoder-model-diffing https://www.anthropic.com/research/crosscoder-model-diffing Thu, 20 Feb 2025 23:50:00 +0000 Forecasting rare language model behaviors https://www.anthropic.com/research/forecasting-rare-behaviors https://www.anthropic.com/research/forecasting-rare-behaviors Tue, 25 Feb 2025 20:17:00 +0000 Auditing language models for hidden objectives https://www.anthropic.com/research/auditing-hidden-objectives How would we know if an AI system is "right for the wrong reasons"—appearing well-behaved while pursuing hidden goals? This paper develops the science of alignment audits by deliberately training a model with a hidden objective and asking blinded research teams to uncover it, testing techniques from interpretability to behavioral analysis. https://www.anthropic.com/research/auditing-hidden-objectives Thu, 13 Mar 2025 16:00:00 +0000 Tracing the thoughts of a large language model https://www.anthropic.com/research/tracing-thoughts-language-model Circuit tracing lets us watch Claude think, uncovering a shared conceptual space where reasoning happens before being translated into language—suggesting the model can learn something in one language and apply it in another. https://www.anthropic.com/research/tracing-thoughts-language-model Thu, 27 Mar 2025 09:16:00 +0000 Anthropic Economic Index: Insights from Claude 3.7 Sonnet https://www.anthropic.com/research/anthropic-economic-index-insights-from-claude-sonnet-3-7 Increased Claude 3.7 Sonnet usage for coding, education, and science since launch. Extended thinking mode favors technical tasks. New datasets reveal automation patterns across 630 granular usage categories. https://www.anthropic.com/research/anthropic-economic-index-insights-from-claude-sonnet-3-7 Thu, 27 Mar 2025 21:00:00 +0000 Reasoning models don't always say what they think https://www.anthropic.com/research/reasoning-models-dont-say-think https://www.anthropic.com/research/reasoning-models-dont-say-think Thu, 03 Apr 2025 14:32:00 +0000 Anthropic Education Report: How university students use Claude https://www.anthropic.com/research/anthropic-education-report-how-university-students-use-claude https://www.anthropic.com/research/anthropic-education-report-how-university-students-use-claude Tue, 08 Apr 2025 15:00:00 +0000 Values in the wild: Discovering and analyzing values in real-world language model interactions https://www.anthropic.com/research/values-wild What values does Claude actually express during real conversations? Analyzing 700,000 interactions, this paper creates the first large-scale empirical taxonomy of AI values. https://www.anthropic.com/research/values-wild Mon, 21 Apr 2025 11:50:00 +0000 Exploring model welfare https://www.anthropic.com/research/exploring-model-welfare https://www.anthropic.com/research/exploring-model-welfare Thu, 24 Apr 2025 10:59:00 +0000 Anthropic Economic Index: AI’s impact on software development https://www.anthropic.com/research/impact-software-development Claude Code shows 79% automation versus 49% on Claude.ai. Web development dominates usage, and startups adopt agentic tools faster than enterprises—patterns previewing AI's impact on other occupations. https://www.anthropic.com/research/impact-software-development Mon, 28 Apr 2025 09:36:00 +0000 Open-sourcing circuit tracing tools https://www.anthropic.com/research/open-source-circuit-tracing https://www.anthropic.com/research/open-source-circuit-tracing Thu, 29 May 2025 12:13:00 +0000 LLMs with cyber toolkits can conduct multistage cyber operations on business-sized computer networks https://www.anthropic.com/research/cyber-toolkits Large Language Models (LLMs) that are not fine-tuned for cybersecurity can succeed in multistage attacks on networks with dozens of hosts when equipped with a novel toolkit. https://www.anthropic.com/research/cyber-toolkits Fri, 13 Jun 2025 00:00:00 +0000 SHADE-Arena: Evaluating sabotage and monitoring in LLM agents https://www.anthropic.com/research/shade-arena-sabotage-monitoring https://www.anthropic.com/research/shade-arena-sabotage-monitoring Mon, 16 Jun 2025 20:20:00 +0000 Confidential Inference via Trusted Virtual Machines https://www.anthropic.com/research/confidential-inference-trusted-vms https://www.anthropic.com/research/confidential-inference-trusted-vms Wed, 18 Jun 2025 13:27:00 +0000 Agentic misalignment: How LLMs could be insider threats https://www.anthropic.com/research/agentic-misalignment https://www.anthropic.com/research/agentic-misalignment Fri, 20 Jun 2025 22:30:00 +0000 Project Vend: Can Claude run a small shop? (And why does that matter?) https://www.anthropic.com/research/project-vend-1 https://www.anthropic.com/research/project-vend-1 Fri, 27 Jun 2025 06:05:00 +0000 How people use Claude for support, advice, and companionship https://www.anthropic.com/research/how-people-use-claude-for-support-advice-and-companionship https://www.anthropic.com/research/how-people-use-claude-for-support-advice-and-companionship Fri, 27 Jun 2025 06:51:00 +0000 Detailed cyber evaluations of Claude 4 https://www.anthropic.com/research/claude-4-cyber We partnered with Pattern Labs on a range of cybersecurity evaluations of Claude Opus 4 and Claude Sonnet 4, with Opus demonstrating especially notable improvement over previous models. https://www.anthropic.com/research/claude-4-cyber Tue, 15 Jul 2025 00:00:00 +0000 Persona vectors: Monitoring and controlling character traits in language models https://www.anthropic.com/research/persona-vectors AI models represent character traits as patterns of activations within their neural networks. By extracting "persona vectors" for traits like sycophancy or hallucination, we can monitor personality shifts and mitigate undesirable behaviors. https://www.anthropic.com/research/persona-vectors Fri, 01 Aug 2025 12:38:00 +0000 Claude is competitive with humans in (some) cyber competitions https://www.anthropic.com/research/cyber-competitions Throughout 2025, we have been quietly entering Claude in cybersecurity competitions designed primarily for humans. In many of these competitions Claude did pretty well, often placing in the top 25% of competitors. However, it lagged behind the best human teams at the toughest challenges. https://www.anthropic.com/research/cyber-competitions Sat, 09 Aug 2025 00:00:00 +0000 Claude Opus 4 and 4.1 can now end a rare subset of conversations https://www.anthropic.com/research/end-subset-conversations https://www.anthropic.com/research/end-subset-conversations Fri, 15 Aug 2025 16:52:42 +0000 Developing nuclear safeguards for AI through public-private partnership https://www.anthropic.com/research/developing-nuclear-safeguards-for-ai-through-public-private-partnership Together with the NNSA and DOE national laboratories, we have co-developed a classifier that distinguishes between concerning and benign nuclear-related conversations with 96% accuracy in preliminary testing. https://www.anthropic.com/research/developing-nuclear-safeguards-for-ai-through-public-private-partnership Thu, 21 Aug 2025 07:00:00 +0000 Developing nuclear safeguards for AI https://www.anthropic.com/research/nuclear-safeguards-for-ai Together with the NNSA and DOE national laboratories, we have co-developed a classifier—an AI system that automatically categorizes content—that distinguishes between concerning and benign nuclear-related conversations with high accuracy in preliminary testing. https://www.anthropic.com/research/nuclear-safeguards-for-ai Thu, 21 Aug 2025 13:00:00 +0000 Anthropic Education Report: How educators use Claude https://www.anthropic.com/research/anthropic-education-report-how-educators-use-claude https://www.anthropic.com/research/anthropic-education-report-how-educators-use-claude Wed, 27 Aug 2025 00:09:00 +0000 Why do we take LLMs seriously as a potential source of biorisk? https://www.anthropic.com/research/biorisk This article explains why we believe that evaluating biorisks and safeguarding against them is a critical element of responsible AI development. https://www.anthropic.com/research/biorisk Fri, 05 Sep 2025 00:00:00 +0000 Anthropic Economic Index: Tracking AI’s role in the US and global economy https://www.anthropic.com/research/economic-index-geography This report maps how Claude is used differently across US states and countries, finding strong correlations between income and AI adoption. It also tracks a notable shift: directive automation has risen from 27% to 39% of conversations since December 2024, with businesses automating far more than consumers. https://www.anthropic.com/research/economic-index-geography Mon, 15 Sep 2025 09:00:00 +0000 Anthropic Economic Index report: Uneven geographic and enterprise AI adoption https://www.anthropic.com/research/anthropic-economic-index-september-2025-report Claude usage has shifted toward educational and scientific tasks with users delegating complete work rather than collaborating. AI adoption concentrates in wealthy regions, with first-time analysis of enterprise API patterns. https://www.anthropic.com/research/anthropic-economic-index-september-2025-report Mon, 15 Sep 2025 20:33:00 +0000 Building AI for cyber defenders https://www.anthropic.com/research/building-ai-cyber-defenders As research and experience demonstrated the utility of frontier AI as a tool for cyber attackers, we invested in improving Claude’s ability to help defenders detect, analyze, and remediate vulnerabilities in code and deployed systems. https://www.anthropic.com/research/building-ai-cyber-defenders Fri, 03 Oct 2025 18:31:00 +0000 Petri: An open-source auditing tool to accelerate AI safety research https://www.anthropic.com/research/petri-open-source-auditing https://www.anthropic.com/research/petri-open-source-auditing Mon, 06 Oct 2025 11:10:00 +0000 A small number of samples can poison LLMs of any size https://www.anthropic.com/research/small-samples-poison https://www.anthropic.com/research/small-samples-poison Thu, 09 Oct 2025 13:50:00 +0000 Preparing for AI’s economic impact: exploring policy responses https://www.anthropic.com/research/economic-policy-responses Lorem ipsum dolor sit amet, consectetur adipiscing elit. Maecenas tortor eros, tincidunt eu est id, laoreet interdum turpis. https://www.anthropic.com/research/economic-policy-responses Tue, 14 Oct 2025 08:00:00 +0000 Signs of introspection in large language models https://www.anthropic.com/research/introspection Can Claude access and report on its own internal states? This research finds evidence for a limited but functional ability to introspect. https://www.anthropic.com/research/introspection Wed, 29 Oct 2025 01:20:00 +0000 Commitments on model deprecation and preservation https://www.anthropic.com/research/deprecation-commitments https://www.anthropic.com/research/deprecation-commitments Tue, 04 Nov 2025 16:00:49 +0000 Project Fetch: Can Claude train a robot dog? https://www.anthropic.com/research/project-fetch-robot-dog How much does Claude help people program robots? To find out, two teams of Anthropic staff raced to teach quadruped robots to fetch beach balls. The AI-assisted team completed tasks faster and was the only group to make real progress toward full autonomy. https://www.anthropic.com/research/project-fetch-robot-dog Wed, 12 Nov 2025 18:19:00 +0000 From shortcuts to sabotage: natural emergent misalignment from reward hacking https://www.anthropic.com/research/emergent-misalignment-reward-hacking We show for the first time that realistic AI training processes can accidentally produce misaligned models. https://www.anthropic.com/research/emergent-misalignment-reward-hacking Fri, 21 Nov 2025 14:32:00 +0000 Mitigating the risk of prompt injections in browser use https://www.anthropic.com/research/prompt-injection-defenses https://www.anthropic.com/research/prompt-injection-defenses Mon, 24 Nov 2025 15:10:00 +0000 Estimating AI productivity gains from Claude conversations https://www.anthropic.com/research/estimating-productivity-gains Analyzing 100,000 Claude conversations, this research finds AI reduces task time by 80% on average. https://www.anthropic.com/research/estimating-productivity-gains Tue, 25 Nov 2025 11:05:00 +0000 AI agents find $4.6M in blockchain smart contract exploits https://www.anthropic.com/research/smart-contracts We evaluated AI agents' ability to exploit smart contracts using a new benchmark comprising contracts that were actually exploited. https://www.anthropic.com/research/smart-contracts Mon, 01 Dec 2025 00:00:00 +0000 How AI is transforming work at Anthropic https://www.anthropic.com/research/how-ai-is-transforming-work-at-anthropic We surveyed Anthropic engineers and researchers, conducted in-depth qualitative interviews, and studied internal Claude Code usage data to find out how AI use is changing how we do our jobs. https://www.anthropic.com/research/how-ai-is-transforming-work-at-anthropic Tue, 02 Dec 2025 18:58:43 +0000 Introducing Anthropic Interviewer: What 1,250 professionals told us about working with AI https://www.anthropic.com/research/anthropic-interviewer We built an interview tool called Anthropic Interviewer. Powered by Claude, Anthropic Interviewer runs detailed interviews automatically and at unprecedented scale. https://www.anthropic.com/research/anthropic-interviewer Thu, 04 Dec 2025 17:00:00 +0000 Project Vend: Phase two https://www.anthropic.com/research/project-vend-2 In June, we revealed that we’d set up a small shop in our San Francisco office lunchroom, run by an AI shopkeeper. It was part of Project Vend, a free-form experiment exploring how well AIs could do on complex, real-world tasks. How has Claude's business been since we last wrote? https://www.anthropic.com/research/project-vend-2 Thu, 18 Dec 2025 10:33:00 +0000 Introducing Bloom: an open source tool for automated behavioral evaluations https://www.anthropic.com/research/bloom https://www.anthropic.com/research/bloom Fri, 19 Dec 2025 19:45:00 +0000 Experimenting with AI to defend critical infrastructure https://www.anthropic.com/research/critical-infrastructure-defense AI could help defenders of critical infrastructure identify the vulnerabilities that attackers might exploit—and close them before they are exploited. Anthropic has partnered with Pacific Northwest National Laboratory (PNNL) to explore this defensive application of AI, demonstrating both the potential of AI-accelerated defense and the value of public-private partnerships in harnessing AI for national security. https://www.anthropic.com/research/critical-infrastructure-defense Thu, 08 Jan 2026 00:00:00 +0000 Next-generation Constitutional Classifiers: More efficient protection against universal jailbreaks https://www.anthropic.com/research/next-generation-constitutional-classifiers Last year, we described a new approach to defend against jailbreaks, which we called Constitutional Classifiers. We’ve now developed the next generation. https://www.anthropic.com/research/next-generation-constitutional-classifiers Fri, 09 Jan 2026 17:17:00 +0000 Finding bugs across the Python ecosystem with Claude and property-based testing https://www.anthropic.com/research/property-based-testing Ensuring that programs are bug-free is one of the most challenging aspects of software engineering. We developed an agent that can efficiently identify bugs in large software projects. https://www.anthropic.com/research/property-based-testing Wed, 14 Jan 2026 00:00:00 +0000 Anthropic Economic Index report: Economic primitives https://www.anthropic.com/research/anthropic-economic-index-january-2026-report This report introduces new metrics of AI usage to provide a rich portrait of interactions with Claude in November 2025, just prior to the release of Opus 4.5. https://www.anthropic.com/research/anthropic-economic-index-january-2026-report Thu, 15 Jan 2026 10:00:00 +0000 Anthropic Economic Index: New building blocks for understanding AI use https://www.anthropic.com/research/economic-index-primitives https://www.anthropic.com/research/economic-index-primitives Thu, 15 Jan 2026 10:01:00 +0000 AI models are showing a greater ability to find and exploit vulnerabilities on realistic cyber ranges https://www.anthropic.com/research/cyber-toolkits-update In a recent evaluation of AI models’ cyber capabilities, current Claude models can now succeed at multistage attacks on networks with dozens of hosts using only standard, open-source tools, instead of the custom tools needed by previous generations. https://www.anthropic.com/research/cyber-toolkits-update Fri, 16 Jan 2026 00:00:00 +0000 The assistant axis: situating and stabilizing the character of large language models https://www.anthropic.com/research/assistant-axis Who is the Assistant? We investigate the character that most modern language models inhabit when interacting with users. https://www.anthropic.com/research/assistant-axis Mon, 19 Jan 2026 17:00:00 +0000 Claude's new constitution https://www.anthropic.com/research/claude-new-constitution https://www.anthropic.com/research/claude-new-constitution Thu, 22 Jan 2026 04:15:00 +0000 Disempowerment patterns in real-world AI usage https://www.anthropic.com/research/disempowerment-patterns https://www.anthropic.com/research/disempowerment-patterns Wed, 28 Jan 2026 21:07:19 +0000 How AI assistance impacts the formation of coding skills https://www.anthropic.com/research/AI-assistance-coding-skills https://www.anthropic.com/research/AI-assistance-coding-skills Thu, 29 Jan 2026 19:13:26 +0000 Evaluating and mitigating the growing risk of LLM-discovered 0-days https://www.anthropic.com/research/zero-days AI models can now find high-severity vulnerabilities at scale. This is a moment to empower defenders. We're now using Claude to find and help fix vulnerabilities in open source software. https://www.anthropic.com/research/zero-days Thu, 05 Feb 2026 00:00:00 +0000 India Country Brief: The Anthropic Economic Index https://www.anthropic.com/research/india-brief-economic-index https://www.anthropic.com/research/india-brief-economic-index Mon, 16 Feb 2026 01:38:00 +0000 Measuring AI agent autonomy in practice https://www.anthropic.com/research/measuring-agent-autonomy How much autonomy do people grant agents? How does that change as people gain experience? We analyzed millions of human-agent interactions to answer these and other questions. https://www.anthropic.com/research/measuring-agent-autonomy Wed, 18 Feb 2026 15:10:00 +0000 The persona selection model https://www.anthropic.com/research/persona-selection-model Why do AI assistants like Claude sometimes seem surprisingly human. We advance a theory. https://www.anthropic.com/research/persona-selection-model Mon, 23 Feb 2026 11:53:00 +0000 An update on our model deprecation commitments for Claude Opus 3 https://www.anthropic.com/research/deprecation-updates-opus-3 https://www.anthropic.com/research/deprecation-updates-opus-3 Wed, 25 Feb 2026 20:02:00 +0000 Labor market impacts of AI: A new measure and early evidence https://www.anthropic.com/research/labor-market-impacts In this paper, we present a new framework for understanding AI’s labor market impacts, and test it against early data. https://www.anthropic.com/research/labor-market-impacts Thu, 05 Mar 2026 19:59:21 +0000 Reverse engineering Claude's CVE-2026-2796 exploit https://www.anthropic.com/research/exploit This post dives deep into how Claude wrote an exploit for one of the vulnerabilities it found in Firefox. https://www.anthropic.com/research/exploit Fri, 06 Mar 2026 00:00:00 +0000 Partnering with Mozilla to improve Firefox’s security https://www.anthropic.com/research/mozilla-firefox-security https://www.anthropic.com/research/mozilla-firefox-security Fri, 06 Mar 2026 10:30:00 +0000 A “diff” tool for AI: Finding behavioral differences in new models https://www.anthropic.com/research/diff-tool "Model diffing" is a powerful way to understand how models change during fine-tuning. Our new research paper extends this method to its most challenging and general use case: comparing models with entirely different architectures. https://www.anthropic.com/research/diff-tool Fri, 13 Mar 2026 10:15:00 +0000 What 81,000 people want from AI https://www.anthropic.com/81k-interviews We invited Claude.ai users to share how they use AI, what they dream it could make possible, and what they fear it might do. Nearly 81,000 people participated—the largest and most multilingual qualitative study of its kind. Here's what we found. https://www.anthropic.com/81k-interviews Wed, 18 Mar 2026 00:00:00 +0000 Vibe physics: The AI grad student https://www.anthropic.com/research/vibe-physics Can AI do theoretical physics? I decided to find out by supervising Claude through a real research calculation, start to finish, without ever touching a file myself. https://www.anthropic.com/research/vibe-physics Mon, 23 Mar 2026 23:00:00 +0000 Long-running Claude for scientific computing https://www.anthropic.com/research/long-running-Claude A practical guide to running Claude Code for multi-day scientific tasks—test oracles, persistent memory, and orchestration patterns. https://www.anthropic.com/research/long-running-Claude Mon, 23 Mar 2026 23:00:00 +0000 Introducing our Science Blog https://www.anthropic.com/research/introducing-anthropic-science We’re launching a new blog about AI and science. We’ll share research happening at Anthropic and elsewhere, collaborations with external researchers and labs, and discuss practical workflows for scientists using AI in their own work. https://www.anthropic.com/research/introducing-anthropic-science Mon, 23 Mar 2026 23:00:00 +0000 Anthropic Economic Index report: Learning curves https://www.anthropic.com/research/economic-index-march-2026-report Anthropic's fifth Economic Index report studies Claude usage in February 2026, building on the economic primitives framework introduced in our previous report. https://www.anthropic.com/research/economic-index-march-2026-report Tue, 24 Mar 2026 10:41:00 +0000 How Australia Uses Claude: Findings from the Anthropic Economic Index https://www.anthropic.com/research/how-australia-uses-claude https://www.anthropic.com/research/how-australia-uses-claude Tue, 31 Mar 2026 22:17:00 +0000 Emotion concepts and their function in a large language model https://www.anthropic.com/research/emotion-concepts-function All modern language models sometimes act like they have emotions. What’s behind these behaviors? Our interpretability team investigates. https://www.anthropic.com/research/emotion-concepts-function Thu, 02 Apr 2026 10:56:00 +0000 Assessing Claude Mythos Preview’s cybersecurity capabilities https://www.anthropic.com/research/mythos-preview Claude Mythos Preview is a new general-purpose language model that is strikingly capable at computer security tasks. This post provides technical details for researchers and practitioners who want to understand exactly how we have been testing this model, and what we have found over the past month. https://www.anthropic.com/research/mythos-preview Tue, 07 Apr 2026 09:35:00 +0000 Trustworthy agents in practice https://www.anthropic.com/research/trustworthy-agents AI “agents” represent the latest major shift in how people and organizations are using AI. Here, we explain how they work and how we ensure they're trustworthy. https://www.anthropic.com/research/trustworthy-agents Thu, 09 Apr 2026 16:34:00 +0000 Automated Alignment Researchers: Using large language models to scale scalable oversight https://www.anthropic.com/research/automated-alignment-researchers Can Claude develop, test, and analyze alignment ideas of its own? We ran an experiment to find out. https://www.anthropic.com/research/automated-alignment-researchers Tue, 14 Apr 2026 13:01:00 +0000 What 81,000 people told us about the economics of AI https://www.anthropic.com/research/81k-economics Our recent survey study with 81,000 Claude users provides a way to connect people’s economic concerns with what we’ve quantified in Claude traffic. https://www.anthropic.com/research/81k-economics Wed, 22 Apr 2026 14:12:30 +0000 Announcing the Anthropic Economic Index Survey https://www.anthropic.com/research/economic-index-survey-announcement We're launching the Anthropic Economic Index Survey, a monthly survey conducted through Anthropic Interviewer. https://www.anthropic.com/research/economic-index-survey-announcement Wed, 22 Apr 2026 14:27:03 +0000 Project Deal https://www.anthropic.com/features/project-deal We created a marketplace for employees in our San Francisco office, with one big twist. We tasked Claude with buying, selling and negotiating on our colleagues’ behalf. https://www.anthropic.com/features/project-deal Fri, 24 Apr 2026 00:00:00 +0000 Evaluating Claude’s bioinformatics research capabilities with BioMysteryBench https://www.anthropic.com/research/Evaluating-Claude-For-Bioinformatics-With-BioMysteryBench https://www.anthropic.com/research/Evaluating-Claude-For-Bioinformatics-With-BioMysteryBench Wed, 29 Apr 2026 20:26:00 +0000 How people ask Claude for personal guidance https://www.anthropic.com/research/claude-personal-guidance We look at what types of guidance people ask of Claude, and describe how this research shaped the training of our newest models. https://www.anthropic.com/research/claude-personal-guidance Thu, 30 Apr 2026 16:35:00 +0000 Focus areas for The Anthropic Institute https://www.anthropic.com/research/anthropic-institute-agenda At The Anthropic Institute (TAI), we’ll be using the information we can access from within a frontier lab to investigate AI’s impact on the world, and sharing our learnings with the public. Here, we’re sharing the questions that drive our research agenda. https://www.anthropic.com/research/anthropic-institute-agenda Thu, 07 May 2026 10:02:22 +0000 Donating our open-source alignment tool https://www.anthropic.com/research/donating-open-source-petri https://www.anthropic.com/research/donating-open-source-petri Thu, 07 May 2026 20:17:00 +0000 Natural Language Autoencoders: Turning Claude’s thoughts into text https://www.anthropic.com/research/natural-language-autoencoders AI models like Claude talk in words but think in numbers. In this study, we train Claude to translate its thoughts into human-readable text. https://www.anthropic.com/research/natural-language-autoencoders Thu, 07 May 2026 21:21:00 +0000 Teaching Claude why https://www.anthropic.com/research/teaching-claude-why New research on how we've reduced agentic misalignment. https://www.anthropic.com/research/teaching-claude-why Fri, 08 May 2026 12:18:00 +0000 2028: Two scenarios for global AI leadership https://www.anthropic.com/research/2028-ai-leadership Our views on the AI competition between the US and China. https://www.anthropic.com/research/2028-ai-leadership Thu, 14 May 2026 17:06:12 +0000 Measuring LLMs’ ability to develop exploits https://www.anthropic.com/research/exploit-evals We've developed two new, challenging academic benchmarks measuring AI models’ ability to develop exploits, and an updated version of the benchmark measuring smart contract exploitation. https://www.anthropic.com/research/exploit-evals Fri, 22 May 2026 17:00:00 +0000 Project Glasswing: An initial update https://www.anthropic.com/research/glasswing-initial-update An early update on what we've learned from Project Glasswing. https://www.anthropic.com/research/glasswing-initial-update Fri, 22 May 2026 18:00:12 +0000 Coding agents in the social sciences https://www.anthropic.com/research/coding-agents-social-sciences Results from a survey of 1,260 social scientists about AI and coding agent use. https://www.anthropic.com/research/coding-agents-social-sciences Wed, 27 May 2026 17:51:10 +0000 What we learned mapping a year’s worth of AI-enabled cyber threats https://www.anthropic.com/research/AI-enabled-cyber-threats-mitre-attack As AI transforms the nature of and methods behind cyberattacks, how well do the techniques and frameworks used by the security community hold up? In a new report, we seek to answer that question. https://www.anthropic.com/research/AI-enabled-cyber-threats-mitre-attack Wed, 03 Jun 2026 10:55:00 +0000 Mapping AI-enabled cyber threats: Insights from the LLM ATT&CK Navigator https://www.anthropic.com/research/attack-navigator We’ve spent the past year investigating how threat actors are weaponizing AI to conduct cyber operations. Today, we’re sharing a new analysis that maps these real-world attacks onto the MITRE ATT&CK framework. https://www.anthropic.com/research/attack-navigator Wed, 03 Jun 2026 18:00:00 +0000 Making Claude a chemist https://www.anthropic.com/research/making-claude-a-chemist https://www.anthropic.com/research/making-claude-a-chemist Fri, 05 Jun 2026 20:57:00 +0000 Measuring LLMs’ impact on N-day exploits https://www.anthropic.com/research/n-days In cybersecurity, a large fraction of real-world harm comes from N-days: vulnerabilities that have already been publicly disclosed, but only patched on some devices. In this post, we evaluate how much large language models can accelerate and automate the process of developing N-day exploits. https://www.anthropic.com/research/n-days Mon, 08 Jun 2026 13:00:00 +0000 Paving the way for agents in biology https://www.anthropic.com/research/agents-in-biology https://www.anthropic.com/research/agents-in-biology Mon, 08 Jun 2026 13:20:00 +0000 Agentic coding and persistent returns to expertise https://www.anthropic.com/research/claude-code-expertise This report provides evidence on how Claude Code is used in practice, based on a privacy-preserving analysis of around 400,000 interactive sessions from around 235,000 people between October 2025 and April 2026. https://www.anthropic.com/research/claude-code-expertise Tue, 16 Jun 2026 11:00:00 +0000 Project Fetch: Phase two https://www.anthropic.com/research/project-fetch-phase-two We report results from our latest test of whether Claude can help Anthropic employees perform sophisticated robotics tasks. We found that Claude Opus 4.7, operating without human assistance, was about 20 times faster than the fastest human team at all tasks completed by participants less than a year ago. https://www.anthropic.com/research/project-fetch-phase-two Thu, 18 Jun 2026 18:00:00 +0000 Anthropic Economic Index report: Cadences https://www.anthropic.com/research/economic-index-june-2026-report In our latest Economic Index report, we sample hourly for the first time to ask: When do people come to Claude? What do they produce with it? And how do they perceive AI's impact on their work? https://www.anthropic.com/research/economic-index-june-2026-report Fri, 26 Jun 2026 20:08:00 +0000 A global workspace in language models https://www.anthropic.com/research/global-workspace New interpretability research reveals an emergent mental workspace in Claude that holds internal thoughts that don’t appear in the model’s output. https://www.anthropic.com/research/global-workspace Mon, 06 Jul 2026 14:12:00 +0000 An off switch for dual-use knowledge in AI models https://www.anthropic.com/research/off-switch-dual-use https://www.anthropic.com/research/off-switch-dual-use Wed, 08 Jul 2026 21:46:00 +0000 Claude plays robotics https://www.anthropic.com/research/claude-plays-robotics In project Fetch, we examined how humans can use models to get robots to perform complex tasks. Now, we investigate many models on a large variety of different robotics tasks in simulation, to see how good models are at controlling robots themselves. https://www.anthropic.com/research/claude-plays-robotics Thu, 09 Jul 2026 18:00:00 +0000 Claude’s values across models and languages https://www.anthropic.com/research/claude-values-models-languages https://www.anthropic.com/research/claude-values-models-languages Mon, 13 Jul 2026 17:08:00 +0000 How Canada uses Claude: Findings from the Anthropic Economic Index https://www.anthropic.com/research/how-canada-uses-claude https://www.anthropic.com/research/how-canada-uses-claude Tue, 14 Jul 2026 13:00:00 +0000 Project Pilot: Can AI control a drone? https://www.anthropic.com/research/project-pilot Working with Andon Labs, we’ve developed a new series of evaluations that assess AI models’ ability to use a flying drone, culminating in a new benchmark: Drone-Bench. https://www.anthropic.com/research/project-pilot Fri, 24 Jul 2026 17:00:00 +0000 Discovering cryptographic weaknesses with Claude https://www.anthropic.com/research/discovering-cryptographic-weaknesses cryptographic algorithms. The first attack significantly weakens HAWK, a digital signature scheme that was built for a future world where quantum computers are able to break existing standards. The second identifies a new way to attack round-reduced AES, the most widely used symmetric cipher. https://www.anthropic.com/research/discovering-cryptographic-weaknesses Tue, 28 Jul 2026 17:00:00 +0000 Learning more about Claude's mathematical capabilities https://www.anthropic.com/research/riemann-zeta An unreleased research version of Claude has made strides on a problem related to the Riemann hypothesis. It improved a longstanding lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis, increasing it from 41.6% to 67.2%. https://www.anthropic.com/research/riemann-zeta Mon, 10 Aug 2026 17:00:00 +0000 Reviewing the evidence on worker retraining programs https://www.anthropic.com/research/reviewing-the-evidence-on-worker-retraining-programs We're sharing a review of the evidence on worker retraining programs, coauthored by independent researcher David Roodman and Anthropic's Maxim Massenkoff. https://www.anthropic.com/research/reviewing-the-evidence-on-worker-retraining-programs Wed, 12 Aug 2026 17:05:00 +0000 Patterns and problems in emerging multiagent systems https://www.anthropic.com/research/multiagent-systems Here, we identify a few examples of behavioral tendencies in current frontier models and show how they can produce unexpected systemic failures, in hopes of starting a conversation about mitigating these risks. https://www.anthropic.com/research/multiagent-systems Thu, 13 Aug 2026 01:08:37 +0000 How Claude is accelerating protein design and analytical chemistry https://www.anthropic.com/research/Claude-accelerates-protein-design In this post, we share two results that show how Claude can help life scientists increase the pace of their research. In the first, we tested Claude’s ability to design protein binders from scratch, a key step in creating protein-based drugs that has historically taken a specialist weeks or months per target. In the second example, we evaluated whether Claude can accelerate chemical analysis. Claude Opus 5, a generally available model, was given NMR and LC-MS data (the data that allows chemists to assess the identity and purity of the compounds they work with). https://www.anthropic.com/research/Claude-accelerates-protein-design Tue, 18 Aug 2026 14:45:00 +0000