Skip to main content

Labs and companies

The labs, companies, universities and funders the history names: what each did for the field and who worked there. The ones in this site's hiring data link to their company page.

Labs and companies

In alphabetical order.

Amazon

1994LayersL6

One of the main employers for Applied Scientist roles. It publishes its hiring process and Leadership Principles as pages, which makes them primary source material for behavioral interview prep.

Related
Behavioral Rounds and the Values Round (set)
Sources
Official How we hire page: apply → assessment → phone screen → interview loop, with the Bar Raiser and Leadership Principles covered in the same place

Anthropic

2021LayersL6L5L4Hiring data

A frontier lab, and the maker of Claude. For the career layer, what matters is that it wrote its policy on candidates' use of AI as a public page. The hiring page itself is the best prep material.

In the history layer it stands for the branch of "alignment as a core research direction". It split off from OpenAI in 2021. Constitutional AI: Harmlessness from AI Feedback replaces most human preference labeling with a set of written principles, and has the model critique and rewrite its own answers. Beyond the Claude series, it released the protocol in MCP:给每个系统一个统一的插口, which made "models connecting to tools" an industry standard. In Opinions and predictions, Dario Amodei's long essay is the most often cited judgment for the frontier layer.

Related
Constitutional AI: Harmlessness from AI Feedback (released) · MCP: One Standard Socket for Every System (released) · Behavioral Rounds and the Values Round (set)
Linked from
Dario Amodei (works at) · Jack Clark (works at) · Jared Kaplan (works at) · Constitutional AI: Harmlessness from AI Feedback (released) · Anthropic 博客(news · engineering · research) (released)
Sources
Official hiring page: no degree required. It suggests putting independent research / blog / open source at the very top of your résumé. · Candidate AI use policy. The page is marked last updated 2025-07-10.

Bell Labs

1925LayersL5

In the 1990s, LeCun's convolutional networks and Vapnik's SVM competed in the same building. The two routes of the second wave (neural networks vs. statistical learning) met here.

Related
Yann LeCun (works at) · Vladimir Vapnik (works at) · Gradient-Based Learning Applied to Document Recognition (released) · Support-Vector Networks (released)
Linked from
Vladimir Vapnik (works at) · Yann LeCun (works at) · Gradient-Based Learning Applied to Document Recognition (works at) · Support-Vector Networks (works at)
Sources
Bell Labs 官网

Carnegie Mellon University

1956LayersL5

Newell and Simon's Logic Theorist (1956) was the first AI program. XCON, speech recognition (Sphinx) and self-driving (NavLab) all came out of this line of work.

Related
The first wave: symbols and expert systems (proposed) · The Second Wave: Statistical Learning and Deep Learning (discusses)
Sources
CMU School of Computer Science

Cursor

2022LayersL6L4Hiring data

An AI editor company (legal entity: Anysphere). It is one of the fastest-hiring AI-native startups. Its hiring page is a good sample of "work-sample style" role requirements.

Sources
Official hiring page. It lists roles only and says nothing about the process.

DARPA

1958LayersL5

The main funder of AI in the 1960s and 70s. Its 1974 cuts were the US side of the first AI winter. In the 1980s its Strategic Computing Initiative gave expert systems another push. When you read AI history, its name stands for the funding curve.

Related
The first wave: symbols and expert systems (funded) · The first wave: symbols and expert systems (discusses)
Sources
DARPA website

DeepSeek

2023LayersL5

A lab in Hangzhou funded by the quant fund High-Flyer. Its R1, released in January 2025 with open weights and a publicly stated low training cost, showed that reasoning models are not one company's secret.

Related
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (released) · Liang Wenfeng (works at)
Linked from
Liang Wenfeng (works at) · DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (released)
Sources
DeepSeek official website

Epoch AI

2022LayersL5L3

A nonprofit research group that tracks AI trends. It publishes open datasets on training compute, model size, chip cost, and when we run out of data. When I talk about scaling and compute economics, I check the numbers here first.

Related
Epoch AI 数据中心 (released) · Pretraining and scaling (measures) · The economics of compute: where training money goes, how inference is priced (discusses)
Linked from
Epoch AI 数据中心 (released)
Sources
Official site

Google

1998LayersL5

Google Brain (2011) turned deep learning into infrastructure. The Transformer, BERT and TPU all came from there. In 2023 it merged with DeepMind to form Google DeepMind, which entered the race with Gemini.

Related
Efficient Estimation of Word Representations in Vector Space (released) · Sequence to Sequence Learning with Neural Networks (released) · Attention Is All You Need (released) · BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (released) · Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (released) · Jeff Dean (works at)
Linked from
Ashish Vaswani (works at) · Geoffrey Hinton (works at) · Ian Goodfellow (works at) · Ilya Sutskever (works at) · Jeff Dean (works at) · Noam Shazeer (works at) · Tomáš Mikolov (works at) · BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (released) · Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (released) · Sequence to Sequence Learning with Neural Networks (released) · Efficient Estimation of Word Representations in Vector Space (released)
Sources
Google Research official website

Google DeepMind

2010LayersL6L5L2

Google's AI lab (DeepMind was founded in 2010 and merged with Google Brain in 2023). It has public official hiring pages for research roles and research engineering roles.

In terms of history, it is the standard-bearer of the second half of the second wave: Playing Atari with Deep Reinforcement Learning taught a network to play Atari from pixels. Mastering the game of Go with deep neural networks and tree search used self-play plus search to beat top human players. Highly accurate protein structure prediction with AlphaFold applied the same approach to protein structures and won a Nobel Prize. Training Compute-Optimal Large Language Models corrected the recipe for scaling: for the same compute, data has to grow together with parameters. After the 2023 merger with Google Brain, the Gemini series put it back on the front line of LLM competition.

Related
Playing Atari with Deep Reinforcement Learning (released) · Mastering the game of Go with deep neural networks and tree search (released) · Highly accurate protein structure prediction with AlphaFold (released) · Training Compute-Optimal Large Language Models (released)
Linked from
Arthur Mensch (works at) · David Silver (works at) · Demis Hassabis (works at) · Highly accurate protein structure prediction with AlphaFold (released) · Mastering the game of Go with deep neural networks and tree search (released) · Training Compute-Optimal Large Language Models (released) · Deep Reinforcement Learning from Human Preferences (released) · Playing Atari with Deep Reinforcement Learning (released) · Google DeepMind 博客 (released)
Sources
Official hiring page. It lays out a four-stage process and says it is tuned per role. · Google's How we hire page. The body text is rendered by script.

Hugging Face

2016LayersL5

The transformers library and the model hub standardized how pretrained models are distributed. It is the "GitHub" of the open-weights ecosystem.

Related
Clément Delangue (works at) · The third wave: scaling and LLM (discusses)
Linked from
Clément Delangue (works at) · Hugging Face Daily Papers (released)
Sources
Hugging Face official website

IBM Research

1945LayersL5

Samuel's checkers program (1959), Deep Blue (1997), Watson (2011), and statistical machine translation (1990s) are all here. In each of the three waves, it stood at the peak of the mainstream method of its day.

Related
Arthur Samuel (works at) · The first wave: symbols and expert systems (discusses)
Linked from
Arthur Samuel (works at)
Sources
IBM history page: Deep Blue

IDSIA

1988LayersL5

A small lab in Lugano, Switzerland, and the birthplace of the LSTM. Around 2011 it won several vision competitions with GPU convolutional networks, before AlexNet.

Related
Jürgen Schmidhuber (works at) · Long Short-Term Memory (released)
Linked from
Jürgen Schmidhuber (works at)
Sources
IDSIA 官网

Meta

2004LayersL6L5

One of the big-tech employers with the most AI roles. Besides FAIR, it has a superintelligence lab. Its careers page is a good sample of how a big company titles its AI roles.

Sources
Official careers site. It returned 429 / 400 to the scraping tool, so I could not verify the candidate guide page.

Meta AI (FAIR)

2013LayersL5

FAIR, which LeCun founded in 2013. PyTorch and LLaMA both came out of it. It is the biggest driver of the open-weights route.

Related
LLaMA: Open and Efficient Foundation Language Models (released) · Yann LeCun (works at) · The third wave: scaling and LLM (discusses)
Linked from
Kaiming He (works at) · Tomáš Mikolov (works at) · Yann LeCun (works at) · LLaMA: Open and Efficient Foundation Language Models (released) · Meta AI 博客 (released)
Sources
Meta AI official website

Microsoft

1975LayersL5

ResNet came out of Microsoft Research Asia. Since 2019 Microsoft has invested in OpenAI and been its exclusive compute provider, and Copilot made code generation a mass-market product for the first time.

Related
Deep Residual Learning for Image Recognition (released) · OpenAI (funded) · Kaiming He (works at)
Linked from
Kaiming He (works at) · Deep Residual Learning for Image Recognition (released)
Sources
Microsoft Research website

Mila – Quebec AI Institute

1993LayersL5

Bengio's lab. In 2014 it produced both the attention mechanism and GANs in the same year. One of the three places (Toronto, Montreal, Edmonton) where Canada kept deep learning alive in academia.

Related
Yoshua Bengio (works at) · Neural Machine Translation by Jointly Learning to Align and Translate (released) · Generative Adversarial Nets (released)
Linked from
Ian Goodfellow (works at) · Yoshua Bengio (works at) · Neural Machine Translation by Jointly Learning to Align and Translate (works at)
Sources
Mila website

Mistral AI

2023LayersL5Hiring data

Founded in Paris in 2023. Mistral 7B and Mixtral showed that a small model with good data can beat larger models. The main player in Europe's open-weights approach.

Related
Arthur Mensch (works at) · The third wave: scaling and LLM (discusses)
Linked from
Arthur Mensch (works at)
Sources
Mistral AI official website

MIT AI Lab / CSAIL

1959LayersL5

The lab Minsky and McCarthy founded in 1959. It was the stronghold of the symbolic school: ELIZA, SHRDLU and the Lisp machines all came out of it. It merged into CSAIL in 2003.

Related
The first wave: symbols and expert systems (proposed) · Perceptrons: An Introduction to Computational Geometry (funded) · Marvin Minsky (works at)
Linked from
John McCarthy (works at) · Joseph Weizenbaum (works at) · Kaiming He (works at) · Marvin Minsky (works at) · Terry Winograd (works at)
Sources
CSAIL website

NVIDIA

1993LayersL5

CUDA (2007) turned the GPU into a general-purpose compute device. After AlexNet, it became the physical foundation of deep learning. The compute economics of the third wave revolve around its supply cycle.

Related
Jensen Huang (works at) · The third wave: scaling and LLM (discusses)
Linked from
Jensen Huang (works at)
Sources
NVIDIA research page

OpenAI

2015LayersL6L5Hiring data

A frontier lab, the author of the GPT series. Research engineer roles are its main hiring line, and its official interview guide is public.

In the history layer, it is the lead of the third wave. Improving Language Understanding by Generative Pre-Training through Language Models are Few-Shot Learners showed that the road of "the same objective, scaled up ten times" could keep going. Scaling Laws for Neural Language Models wrote that down as a formula. Training language models to follow instructions with human feedback taught the continuation machine to follow instructions. ChatGPT then put it in everyone's hands. The later GPT-4 Technical Report no longer discloses the architecture, which marks research moving from papers to products. For the nodes, see The third wave: scaling and LLM and Pretraining and scaling.

Related
Improving Language Understanding by Generative Pre-Training (released) · Language Models are Unsupervised Multitask Learners (released) · Language Models are Few-Shot Learners (released) · GPT-4 Technical Report (released) · Training language models to follow instructions with human feedback (released) · Scaling Laws for Neural Language Models (released) · Learning Transferable Visual Models From Natural Language Supervision (released)
Linked from
Alec Radford (works at) · Andrej Karpathy (works at) · Daniel Kokotajlo (works at) · Dario Amodei (works at) · Ilya Sutskever (works at) · Jared Kaplan (works at) · Paul Christiano (works at) · Sam Altman (works at) · Learning Transferable Visual Models From Natural Language Supervision (released) · Deep Reinforcement Learning from Human Preferences (released) · Improving Language Understanding by Generative Pre-Training (released) · Language Models are Unsupervised Multitask Learners (released) · Language Models are Few-Shot Learners (released) · GPT-4 Technical Report (released) · Training language models to follow instructions with human feedback (released) · Scaling Laws for Neural Language Models (released) · Microsoft (funded) · OpenAI 博客 (released)
Sources
Official interview guide. The fetch tool got a 403, but the search index confirms the page exists (including multilingual versions).

Stanford AI Lab (SAIL)

1963LayersL5

Founded by McCarthy in 1963. Expert systems (DENDRAL, MYCIN) came out of it. Forty years later, ImageNet and DPO came from here too. It shows up in all the three waves.

Related
The first wave: symbols and expert systems (proposed) · Direct Preference Optimization: Your Language Model is Secretly a Reward Model (released) · ImageNet (proposed)
Linked from
Andrew Ng (works at) · Edward Feigenbaum (works at) · Fei-Fei Li (works at) · John McCarthy (works at) · Terry Winograd (works at) · Direct Preference Optimization: Your Language Model is Secretly a Reward Model (released)
Sources
SAIL website

University of Toronto

1827LayersL5

The home of Hinton's lab, where deep learning was kept alive through the winter. Deep belief networks (2006) and AlexNet (2012) were both built here.

Related
Geoffrey Hinton (works at) · ImageNet Classification with Deep Convolutional Neural Networks (released) · A fast learning algorithm for deep belief nets (released)
Linked from
Alex Krizhevsky (works at) · Geoffrey Hinton (works at) · A fast learning algorithm for deep belief nets (works at)
Sources
University of Toronto, Department of Computer Science