ABAB Pedia

Goodfire

Goodfire is an AI interpretability research company founded in 2024 and based in San Francisco. Organized as a public benefit corporation, it was co-founded by Eric Ho, Tom McGrath and Daniel Balsam. Its research and products, including Ember and Silico, focus on understanding, debugging and monitoring neural networks.

Contents11 sections
Key facts

founding and research focus

Goodfire was founded by Eric Ho, Tom McGrath and Daniel (Dan) Balsam. Ho had co-founded AI recruiting company RippleMatch, McGrath had established an interpretability team at Google DeepMindGoogle DeepMindGoogle DeepMind is Google’s artificial intelligence research organization, formed in April 2023 by merging DeepMind, which Google acquired in 2014, with the Brain team from Google Research, and led by CEO Demis Hassabis. It developed AlphaGo, AlphaFold, and the Gemini family of models, and Hassabis and John Jumper received the 2024 Nobel Prize in Chemistry for AlphaFold’s protein structure prediction work.Open the full entry, and Balsam had been a founding engineer and AI leader at RippleMatch. They sought to turn analysis of neural-network internals into practical debugging tools.⁠[1][2]

$7 million seed round

Lightspeed announced that it had led Goodfire’s $7 million seed round. Its investment account described tools for examining, visualizing and adjusting concepts within models, alongside approaches based on training data and prompting. Menlo’s portfolio profile dates its own seed-stage investment to 2024.⁠[2][1]

open-source sparse autoencoders

Following Ember’s introduction, Goodfire released sparse autoencoders for Llama 3.1 8B and Llama 3.3 70B. These interpreter models decompose neural activations into more interpretable features and underpin feature discovery and behavioral steering in Ember’s API and SDK. The release included model weights, documentation and examples for researchers.⁠[3]

$50 million Series A

Goodfire announced a $50 million Series A led by Menlo Ventures, with participation from Lightspeed, AnthropicAnthropicAnthropic is an artificial intelligence research and product company founded in 2021 that develops the Claude models, assistant, and tools such as Claude Code, and operates as a public benefit corporation. Dario Amodei and Daniela Amodei are its co-founders, and its research covers model safety, interpretability, and steerability.Open the full entry, B Capital, Work-Bench, Wing and South Park Commons. It earmarked the funding for interpretability research and Ember, and described collaboration with Arc Institute on interpreting the Evo 2 DNA foundation model.⁠[4]

$150 million Series B

Goodfire announced a $150 million Series B at a $1.25 billion valuation, led by B Capital with participation from Juniper Ventures, DFJ Growth, Salesforce Ventures, Menlo, Lightspeed and others. It described two uses for its model design environment: understanding and changing model behavior, and extracting knowledge from scientific models for researchers to investigate. Named partners included Arc Institute, Mayo Clinic and Microsoft.⁠[5]

Silico research grants

Goodfire launched a research-grants program offering selected academic and nonprofit teams a total of $1 million in free Silico usage. It covered AI safety, fundamental interpretability and alignment, and interpretability for life sciences. The announced support consisted of access to the platform rather than an equivalent cash fund.⁠[6]

Baseten partnership

Goodfire and Baseten announced Project Beacon to combine activation-based monitoring with Baseten’s inference platform for risks such as prompt injection, sensitive-data exposure and cyber misuse. Small classifiers can provide an initial screen before further model evaluation or review, with configurable logging, refusal and rerouting. Detection figures in the announcement came from company-specific evaluations and do not establish complete risk elimination.⁠[7]

Topics: Silico and research workflows

Silico is Goodfire’s interpretability agent and model design environment. Given a research goal, it can plan experiments, coordinate concurrent runs, monitor progress and return inspectable results. Tools include SAEs, probes, causal analysis, model comparisons and data attribution, with applications in language models, life sciences, robotics and vision. The website offers a macOS client and team-deployment inquiries.⁠[8]

Topics: corporate form and technical goals

Based in San Francisco, Goodfire operates as a public benefit corporation with researchers and engineers from organizations including OpenAI and Google DeepMind. Its central field is mechanistic interpretability: examining internal representations and computations to explain, monitor and intervene in model behavior. Understanding and aligning models remains a research objective, rather than a claim that existing models are fully explained or guaranteed safe.⁠[9][5]

Official website and public accounts

Sources

  1. Deedy Das | Menlo Ventures (opens in a new window)
  2. Goodfire: Building Interpretable AI (opens in a new window)
  3. Announcing Open-Source SAEs for Llama 3.3 70B and Llama 3.1 8B (opens in a new window)
  4. Announcing Our $50M Series A to Advance AI Interpretability Research (opens in a new window)
  5. Understanding, Learning From, and Designing AI: Our Series B (opens in a new window)
  6. Announcing Goodfire Research Grants (opens in a new window)
  7. Goodfire and Baseten partner to bring frontier safety to open models (opens in a new window)
  8. Silico — your interpretability agent (opens in a new window)
  9. About Goodfire (opens in a new window)
  10. Goodfire (@GoodfireAI) (opens in a new window)
  11. Goodfire | LinkedIn (opens in a new window)