Labs and companies
The labs, companies, universities and funders the history names: what each did for the field and who worked there. The ones in this site's hiring data link to their company page.
Labs and companies
In alphabetical order.
Amazon
1994LayersL6
One of the main employers for Applied Scientist roles. It publishes its hiring process and Leadership Principles as pages, which makes them primary source material for behavioral interview prep.
Anthropic
2021LayersL6L5L4Hiring data
A frontier lab, and the maker of Claude. For the career layer, what matters is that it wrote its policy on candidates' use of AI as a public page. The hiring page itself is the best prep material.
In the history layer it stands for the branch of "alignment as a core research direction". It split off from OpenAI in 2021. Constitutional AI: Harmlessness from AI Feedback replaces most human preference labeling with a set of written principles, and has the model critique and rewrite its own answers. Beyond the Claude series, it released the protocol in MCP:给每个系统一个统一的插口, which made "models connecting to tools" an industry standard. In Opinions and predictions, Dario Amodei's long essay is the most often cited judgment for the frontier layer.
- Related
- Constitutional AI: Harmlessness from AI Feedback (released) · MCP: One Standard Socket for Every System (released) · Behavioral Rounds and the Values Round (set)
- Linked from
- Dario Amodei (works at) · Jack Clark (works at) · Jared Kaplan (works at) · Constitutional AI: Harmlessness from AI Feedback (released) · Anthropic 博客(news · engineering · research) (released)
- Sources
- Official hiring page: no degree required. It suggests putting independent research / blog / open source at the very top of your résumé. · Candidate AI use policy. The page is marked last updated 2025-07-10.
Bell Labs
1925LayersL5
In the 1990s, LeCun's convolutional networks and Vapnik's SVM competed in the same building. The two routes of the second wave (neural networks vs. statistical learning) met here.
- Related
- Yann LeCun (works at) · Vladimir Vapnik (works at) · Gradient-Based Learning Applied to Document Recognition (released) · Support-Vector Networks (released)
- Linked from
- Vladimir Vapnik (works at) · Yann LeCun (works at) · Gradient-Based Learning Applied to Document Recognition (works at) · Support-Vector Networks (works at)
- Sources
- Bell Labs 官网
Carnegie Mellon University
1956LayersL5
Newell and Simon's Logic Theorist (1956) was the first AI program. XCON, speech recognition (Sphinx) and self-driving (NavLab) all came out of this line of work.
- Related
- The first wave: symbols and expert systems (proposed) · The Second Wave: Statistical Learning and Deep Learning (discusses)
- Sources
- CMU School of Computer Science
Cursor
2022LayersL6L4Hiring data
An AI editor company (legal entity: Anysphere). It is one of the fastest-hiring AI-native startups. Its hiring page is a good sample of "work-sample style" role requirements.
DARPA
1958LayersL5
The main funder of AI in the 1960s and 70s. Its 1974 cuts were the US side of the first AI winter. In the 1980s its Strategic Computing Initiative gave expert systems another push. When you read AI history, its name stands for the funding curve.
- Related
- The first wave: symbols and expert systems (funded) · The first wave: symbols and expert systems (discusses)
- Sources
- DARPA website
DeepSeek
2023LayersL5
A lab in Hangzhou funded by the quant fund High-Flyer. Its R1, released in January 2025 with open weights and a publicly stated low training cost, showed that reasoning models are not one company's secret.
- Related
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (released) · Liang Wenfeng (works at)
- Linked from
- Liang Wenfeng (works at) · DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (released)
- Sources
- DeepSeek official website
Epoch AI
2022LayersL5L3
A nonprofit research group that tracks AI trends. It publishes open datasets on training compute, model size, chip cost, and when we run out of data. When I talk about scaling and compute economics, I check the numbers here first.
- Related
- Epoch AI 数据中心 (released) · Pretraining and scaling (measures) · The economics of compute: where training money goes, how inference is priced (discusses)
- Linked from
- Epoch AI 数据中心 (released)
- Sources
- Official site
1998LayersL5
Google Brain (2011) turned deep learning into infrastructure. The Transformer, BERT and TPU all came from there. In 2023 it merged with DeepMind to form Google DeepMind, which entered the race with Gemini.
- Related
- Efficient Estimation of Word Representations in Vector Space (released) · Sequence to Sequence Learning with Neural Networks (released) · Attention Is All You Need (released) · BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (released) · Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (released) · Jeff Dean (works at)
- Linked from
- Ashish Vaswani (works at) · Geoffrey Hinton (works at) · Ian Goodfellow (works at) · Ilya Sutskever (works at) · Jeff Dean (works at) · Noam Shazeer (works at) · Tomáš Mikolov (works at) · BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (released) · Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (released) · Sequence to Sequence Learning with Neural Networks (released) · Efficient Estimation of Word Representations in Vector Space (released)
- Sources
- Google Research official website
Google DeepMind
2010LayersL6L5L2
Google's AI lab (DeepMind was founded in 2010 and merged with Google Brain in 2023). It has public official hiring pages for research roles and research engineering roles.
In terms of history, it is the standard-bearer of the second half of the second wave: Playing Atari with Deep Reinforcement Learning taught a network to play Atari from pixels. Mastering the game of Go with deep neural networks and tree search used self-play plus search to beat top human players. Highly accurate protein structure prediction with AlphaFold applied the same approach to protein structures and won a Nobel Prize. Training Compute-Optimal Large Language Models corrected the recipe for scaling: for the same compute, data has to grow together with parameters. After the 2023 merger with Google Brain, the Gemini series put it back on the front line of LLM competition.
- Related
- Playing Atari with Deep Reinforcement Learning (released) · Mastering the game of Go with deep neural networks and tree search (released) · Highly accurate protein structure prediction with AlphaFold (released) · Training Compute-Optimal Large Language Models (released)
- Linked from
- Arthur Mensch (works at) · David Silver (works at) · Demis Hassabis (works at) · Highly accurate protein structure prediction with AlphaFold (released) · Mastering the game of Go with deep neural networks and tree search (released) · Training Compute-Optimal Large Language Models (released) · Deep Reinforcement Learning from Human Preferences (released) · Playing Atari with Deep Reinforcement Learning (released) · Google DeepMind 博客 (released)
- Sources
- Official hiring page. It lays out a four-stage process and says it is tuned per role. · Google's How we hire page. The body text is rendered by script.
Hugging Face
2016LayersL5
The transformers library and the model hub standardized how pretrained models are distributed. It is the "GitHub" of the open-weights ecosystem.
- Related
- Clément Delangue (works at) · The third wave: scaling and LLM (discusses)
- Linked from
- Clément Delangue (works at) · Hugging Face Daily Papers (released)
- Sources
- Hugging Face official website
IBM Research
1945LayersL5
Samuel's checkers program (1959), Deep Blue (1997), Watson (2011), and statistical machine translation (1990s) are all here. In each of the three waves, it stood at the peak of the mainstream method of its day.
- Related
- Arthur Samuel (works at) · The first wave: symbols and expert systems (discusses)
- Linked from
- Arthur Samuel (works at)
- Sources
- IBM history page: Deep Blue
IDSIA
1988LayersL5
A small lab in Lugano, Switzerland, and the birthplace of the LSTM. Around 2011 it won several vision competitions with GPU convolutional networks, before AlexNet.
- Related
- Jürgen Schmidhuber (works at) · Long Short-Term Memory (released)
- Linked from
- Jürgen Schmidhuber (works at)
- Sources
- IDSIA 官网
Meta
2004LayersL6L5
One of the big-tech employers with the most AI roles. Besides FAIR, it has a superintelligence lab. Its careers page is a good sample of how a big company titles its AI roles.
Meta AI (FAIR)
2013LayersL5
FAIR, which LeCun founded in 2013. PyTorch and LLaMA both came out of it. It is the biggest driver of the open-weights route.
- Related
- LLaMA: Open and Efficient Foundation Language Models (released) · Yann LeCun (works at) · The third wave: scaling and LLM (discusses)
- Linked from
- Kaiming He (works at) · Tomáš Mikolov (works at) · Yann LeCun (works at) · LLaMA: Open and Efficient Foundation Language Models (released) · Meta AI 博客 (released)
- Sources
- Meta AI official website
Microsoft
1975LayersL5
ResNet came out of Microsoft Research Asia. Since 2019 Microsoft has invested in OpenAI and been its exclusive compute provider, and Copilot made code generation a mass-market product for the first time.
- Related
- Deep Residual Learning for Image Recognition (released) · OpenAI (funded) · Kaiming He (works at)
- Linked from
- Kaiming He (works at) · Deep Residual Learning for Image Recognition (released)
- Sources
- Microsoft Research website
Mila – Quebec AI Institute
1993LayersL5
Bengio's lab. In 2014 it produced both the attention mechanism and GANs in the same year. One of the three places (Toronto, Montreal, Edmonton) where Canada kept deep learning alive in academia.
- Related
- Yoshua Bengio (works at) · Neural Machine Translation by Jointly Learning to Align and Translate (released) · Generative Adversarial Nets (released)
- Linked from
- Ian Goodfellow (works at) · Yoshua Bengio (works at) · Neural Machine Translation by Jointly Learning to Align and Translate (works at)
- Sources
- Mila website
Mistral AI
2023LayersL5Hiring data
Founded in Paris in 2023. Mistral 7B and Mixtral showed that a small model with good data can beat larger models. The main player in Europe's open-weights approach.
- Related
- Arthur Mensch (works at) · The third wave: scaling and LLM (discusses)
- Linked from
- Arthur Mensch (works at)
- Sources
- Mistral AI official website
MIT AI Lab / CSAIL
1959LayersL5
The lab Minsky and McCarthy founded in 1959. It was the stronghold of the symbolic school: ELIZA, SHRDLU and the Lisp machines all came out of it. It merged into CSAIL in 2003.
- Related
- The first wave: symbols and expert systems (proposed) · Perceptrons: An Introduction to Computational Geometry (funded) · Marvin Minsky (works at)
- Linked from
- John McCarthy (works at) · Joseph Weizenbaum (works at) · Kaiming He (works at) · Marvin Minsky (works at) · Terry Winograd (works at)
- Sources
- CSAIL website
NVIDIA
1993LayersL5
CUDA (2007) turned the GPU into a general-purpose compute device. After AlexNet, it became the physical foundation of deep learning. The compute economics of the third wave revolve around its supply cycle.
- Related
- Jensen Huang (works at) · The third wave: scaling and LLM (discusses)
- Linked from
- Jensen Huang (works at)
- Sources
- NVIDIA research page
OpenAI
2015LayersL6L5Hiring data
A frontier lab, the author of the GPT series. Research engineer roles are its main hiring line, and its official interview guide is public.
In the history layer, it is the lead of the third wave. Improving Language Understanding by Generative Pre-Training through Language Models are Few-Shot Learners showed that the road of "the same objective, scaled up ten times" could keep going. Scaling Laws for Neural Language Models wrote that down as a formula. Training language models to follow instructions with human feedback taught the continuation machine to follow instructions. ChatGPT then put it in everyone's hands. The later GPT-4 Technical Report no longer discloses the architecture, which marks research moving from papers to products. For the nodes, see The third wave: scaling and LLM and Pretraining and scaling.
- Related
- Improving Language Understanding by Generative Pre-Training (released) · Language Models are Unsupervised Multitask Learners (released) · Language Models are Few-Shot Learners (released) · GPT-4 Technical Report (released) · Training language models to follow instructions with human feedback (released) · Scaling Laws for Neural Language Models (released) · Learning Transferable Visual Models From Natural Language Supervision (released)
- Linked from
- Alec Radford (works at) · Andrej Karpathy (works at) · Daniel Kokotajlo (works at) · Dario Amodei (works at) · Ilya Sutskever (works at) · Jared Kaplan (works at) · Paul Christiano (works at) · Sam Altman (works at) · Learning Transferable Visual Models From Natural Language Supervision (released) · Deep Reinforcement Learning from Human Preferences (released) · Improving Language Understanding by Generative Pre-Training (released) · Language Models are Unsupervised Multitask Learners (released) · Language Models are Few-Shot Learners (released) · GPT-4 Technical Report (released) · Training language models to follow instructions with human feedback (released) · Scaling Laws for Neural Language Models (released) · Microsoft (funded) · OpenAI 博客 (released)
- Sources
- Official interview guide. The fetch tool got a 403, but the search index confirms the page exists (including multilingual versions).
Stanford AI Lab (SAIL)
1963LayersL5
Founded by McCarthy in 1963. Expert systems (DENDRAL, MYCIN) came out of it. Forty years later, ImageNet and DPO came from here too. It shows up in all the three waves.
- Related
- The first wave: symbols and expert systems (proposed) · Direct Preference Optimization: Your Language Model is Secretly a Reward Model (released) · ImageNet (proposed)
- Linked from
- Andrew Ng (works at) · Edward Feigenbaum (works at) · Fei-Fei Li (works at) · John McCarthy (works at) · Terry Winograd (works at) · Direct Preference Optimization: Your Language Model is Secretly a Reward Model (released)
- Sources
- SAIL website
University of Toronto
1827LayersL5
The home of Hinton's lab, where deep learning was kept alive through the winter. Deep belief networks (2006) and AlexNet (2012) were both built here.
- Related
- Geoffrey Hinton (works at) · ImageNet Classification with Deep Convolutional Neural Networks (released) · A fast learning algorithm for deep belief nets (released)
- Linked from
- Alex Krizhevsky (works at) · Geoffrey Hinton (works at) · A fast learning algorithm for deep belief nets (works at)
- Sources
- University of Toronto, Department of Computer Science