Geoffrey Everest Hinton (born December 6, 1947) is a British-Canadian computer scientist and cognitive psychologist whose work over five decades established much of the conceptual foundation of modern deep learning. With David Rumelhart and Ronald J. Williams he co-authored the 1986 Nature paper that popularized backpropagation for training multi-layer networks;1 with Terence Sejnowski and David Ackley he co-invented the Boltzmann machine, the stochastic generative network that anchored his lifelong work on unsupervised learning;2 and at the University of Toronto he led the research program, deep belief networks in 2006 and then the GPU-trained AlexNet built by his students Alex Krizhevsky and Ilya Sutskever in 2012, that restarted neural network research after its second long winter.34 For this body of work he shared the 2018 ACM A.M. Turing Award with Yoshua Bengio and Yann LeCun,5 and in 2024 he shared the Nobel Prize in Physics with John Hopfield "for foundational discoveries and inventions that enable machine learning with artificial neural networks."6

Hinton spent the core of his academic career at the University of Toronto, where he is University Professor Emeritus, and from 2013 to 2023 he divided his time between Toronto and Google, where he was a vice president and engineering fellow in the Google Brain organization.7 In 2017 he co-founded the Vector Institute in Toronto and became its chief scientific advisor.8 In May 2023 he resigned from Google, saying he wished to speak freely about the risks of artificial intelligence, and he has since been one of the most prominent public voices arguing that advanced AI poses serious, possibly existential, dangers.9 The press frequently describes him, with Bengio and LeCun, as one of the "godfathers of AI."7

Career

Hinton was born in Wimbledon, London, in 1947. He studied experimental psychology at King's College, Cambridge, receiving his bachelor's degree in 1970, spent a year apprenticing in carpentry, and from 1972 studied artificial intelligence at the University of Edinburgh, where he received his PhD in 1978 with a thesis on relaxation processes in vision supervised by Christopher Longuet-Higgins, a leading figure of the symbolic approach the young Hinton had already begun to dissent from.7 After postdoctoral work at the University of Sussex, at the Medical Research Council's Applied Psychology Unit in Cambridge, and at the University of California, San Diego, where he joined Rumelhart's Parallel Distributed Processing group, he took a faculty position at Carnegie Mellon University in 1982. There, with Sejnowski, Ackley, and Scott Fahlman, he developed the Boltzmann machine learning algorithm.72

In 1987 Hinton moved to the University of Toronto, attracted in part by funding from the Canadian Institute for Advanced Research, whose neural computation program he later directed. The 1986 Nature paper with Rumelhart and Williams had appeared just before the move, and through the late 1980s and 1990s his Toronto group worked on backpropagation variants, mixtures of experts, and generative models while neural networks fell out of fashion in the wider field.7 From 1998 to 2001 he was the founding director of the Gatsby Computational Neuroscience Unit at University College London, then returned to Toronto.7

YearsRoleOrganization
1982-1987Faculty member, computer scienceCarnegie Mellon University
1987-1998ProfessorUniversity of Toronto
1998-2001Founding director, Gatsby Computational Neuroscience UnitUniversity College London
2001-Professor (now University Professor Emeritus)University of Toronto
2013-2023Vice president and engineering fellowGoogle (Google Brain)
2017-Co-founder and chief scientific advisorVector Institute

The second phase of Hinton's career began in 2006, when he, Simon Osindero, and Yee-Whye Teh showed that stacking restricted Boltzmann machines into deep belief networks, trained greedily one layer at a time, made deep architectures trainable where plain backpropagation had stalled; a companion paper with Ruslan Salakhutdinov in Science demonstrated deep autoencoders for dimensionality reduction.1011 The phrase "deep learning" dates from this period. Between 2009 and 2012 his group, working with George Dahl and Abdel-rahman Mohamed, carried the same methods to speech recognition, replacing the Gaussian mixture models inside production recognizers at Microsoft, IBM, and Google.12

In 2012, Krizhevsky and Sutskever's AlexNet won the ImageNet Large Scale Visual Recognition Challenge by a margin of more than ten percentage points.4 Hinton founded a small company, DNNresearch, with his two students to commercialize the group's work, and Google acquired it in March 2013 for a reported $44 million; Hinton thereafter split his time between Toronto and Google, whose Google Brain group had been founded by Andrew Ng.713 At Google he produced the knowledge distillation paper of 2015, the capsule network line beginning in 2017, and work on contrastive learning, label smoothing, and the sparsely gated mixture-of-experts layer.141516 In 2017 he co-founded the Vector Institute, a Toronto nonprofit research institute, serving as its chief scientific advisor.8

On May 1, 2023, the New York Times reported that Hinton had resigned from Google, after a decade, in order to warn about the dangers of AI without constraint from an employer.9 In October 2024 the Royal Swedish Academy of Sciences awarded him half of the Nobel Prize in Physics, shared with Hopfield, citing the Boltzmann machine and the Hopfield network as the inventions that grounded machine learning in statistical physics.6 He received the award in Stockholm in December 2024.7

Research contributions

Boltzmann machines and stochastic generative models

Hinton's first landmark was the Boltzmann machine, developed at Carnegie Mellon with Sejnowski and Ackley and published in 1985. The model generalized Hopfield's 1982 network into a stochastic system with hidden units and gave it a learning rule: adjust weights so that the network's spontaneous statistics, in equilibrium, match the statistics it shows when clamped to observed data. The contrast between the clamped (wake) and free-running (sleep) phases anticipated the wake-sleep and contrastive methods Hinton pursued for the rest of his career.2 Sejnowski soon found ordinary backpropagation far faster for supervised tasks, using it to train NETtalk, but the Boltzmann machine remained Hinton's reference point for unsupervised learning, and its energy-based formulation, descending from the Hopfield network, was central to the 2024 Nobel citation.16

Backpropagation and distributed representations

The 1986 Nature paper "Learning representations by back-propagating errors," with Rumelhart and Williams, demonstrated that gradient descent through multiple layers could discover useful internal representations, answering the charge, pressed since Frank Rosenblatt's perceptron era, that layered networks could not be trained.1 The method had been derived earlier by Paul Werbos and had antecedents in control theory, but the 1986 paper, alongside the Parallel Distributed Processing volumes, made it the canonical training algorithm of connectionism.17 Hinton's companion work of the period introduced distributed representations, concepts encoded as patterns across many units rather than single symbols, an idea that runs directly through word2vec to the embeddings of modern language models.7

Mixtures of experts

In 1991, with Robert Jacobs, Michael Jordan, and Steven Nowlan, Hinton published "Adaptive Mixtures of Local Experts," in which specialized subnetworks train under a gating network that learns to route each case to the appropriate expert.17 The idea lay largely dormant until 2017, when Hinton co-authored the sparsely gated mixture-of-experts layer with Noam Shazeer and colleagues at Google, scaling a recurrent language model to 137 billion parameters by activating only a few experts per token.16 That conditional-computation design is now standard in frontier large language models.16

Deep belief networks and the deep learning revival

The practical revival of deep networks came from Hinton's return to the Boltzmann tradition. His contrastive divergence algorithm of 2002 made restricted Boltzmann machines trainable in reasonable time,18 and the 2006 deep belief network paper showed that greedy layer-wise pretraining followed by backpropagation fine-tuning could train deep architectures that had previously been intractable.10 The Science autoencoder paper with Salakhutdinov the same year gave the approach a striking demonstration, compressing documents and images into low-dimensional codes that outperformed principal component analysis.11 Bengio's group in Montreal pursued parallel results, and the 2015 Nature review by LeCun, Bengio, and Hinton codified "deep learning" as the name of the reunited field.19

Speech recognition

From 2009, Hinton's students Dahl and Mohamed applied deep belief networks to acoustic modeling, and by 2012 the approach had been validated at industrial scale: the IEEE Signal Processing Magazine paper that year, with co-authors across Microsoft, IBM, and Google, reported deep networks displacing three decades of Gaussian mixture modeling in production speech systems.12 Speech was the first large industry the revival conquered, and the credibility earned there fed directly into the vision work that followed.7

AlexNet, dropout, and the GPU turn

Hinton's supervisory role in AlexNet is well documented: Krizhevsky wrote the CUDA implementation and ran the experiments, Sutskever argued that a large network trained on enough data would win, and Hinton, the senior author, had championed exactly that bet for years.47 The same Toronto circle produced dropout, introduced in a 2012 preprint with Nitish Srivastava, Krizhevsky, Sutskever, and Salakhutdinov and consolidated in a 2014 journal article; randomly silencing units during training became one of the most used regularizers of the decade.20 Hinton also introduced, in his 2012 Coursera lectures and never in a formal paper, the RMSProp optimizer, whose per-coordinate learning rates directly inspired Adam.21 The episode validated the conviction, central to Hinton's advocacy and later codified in the scaling-law literature, that large networks, large data, and large compute would keep delivering capability.

Distillation, capsules, and later work

At Google, Hinton's 2015 paper with Oriol Vinyals and Jeff Dean introduced knowledge distillation, training a small network on the softened softmax outputs of a large model or ensemble; the technique, which Hinton called "dark knowledge," became a standard compression and transfer method and underlies later student-teacher systems such as DistilBERT.14 From 2017 he pursued capsule networks, architectures that attempt to represent part-whole hierarchies with groups of neurons whose activity vectors encode pose, an explicitly stated attempt to fix what he saw as convolutional networks' mishandling of viewpoint.15 His other late Google work includes the label smoothing analysis of 2019 and the SimCLR contrastive learning framework of 2020 with Ting Chen, Simon Kornblith, and Mohammad Norouzi.2223 In 2022 he proposed the Forward-Forward algorithm, a two-pass local learning rule offered as a plausible alternative to backpropagation for low-power hardware and as a more biologically credible account of learning in cortex.24

His doctoral students and postdoctoral researchers include Sutskever, Krizhevsky, Salakhutdinov, Brendan Frey, Radford Neal, Richard Zemel, George Dahl, Max Welling, Zoubin Ghahramani, and, as a postdoctoral visitor, LeCun.7

Selected publications

Ordered chronologically.

  1. Ackley, D. H., Hinton, G. E., and Sejnowski, T. J., "A Learning Algorithm for Boltzmann Machines," Cognitive Science 9(1), 1985. Introduced the Boltzmann machine and its wake-sleep learning rule.
  2. Rumelhart, D. E., Hinton, G. E., and Williams, R. J., "Learning Representations by Back-Propagating Errors," Nature 323, 1986. The paper that popularized backpropagation.
  3. Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E., "Adaptive Mixtures of Local Experts," Neural Computation 3(1), 1991. Founded mixture-of-experts modeling.
  4. Hinton, G. E., "Training Products of Experts by Minimizing Contrastive Divergence," Neural Computation 14(8), 2002. Made restricted Boltzmann machines practical to train.
  5. Hinton, G. E., Osindero, S., and Teh, Y.-W., "A Fast Learning Algorithm for Deep Belief Nets," Neural Computation 18(7), 2006. Greedy layer-wise pretraining; opened the deep learning revival.
  6. Hinton, G. E., and Salakhutdinov, R. R., "Reducing the Dimensionality of Data with Neural Networks," Science 313(5786), 2006. Deep autoencoders beating linear methods.
  7. Hinton, G., Deng, L., Yu, D., Dahl, G., Mohamed, A., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T., and Kingsbury, B., "Deep Neural Networks for Acoustic Modeling in Speech Recognition," IEEE Signal Processing Magazine 29(6), 2012. Carried deep learning into production speech.
  8. Krizhevsky, A., Sutskever, I., and Hinton, G. E., "ImageNet Classification with Deep Convolutional Neural Networks," NeurIPS, 2012. The AlexNet paper; NeurIPS Test of Time Award, 2022.
  9. Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R., "Dropout: A Simple Way to Prevent Neural Networks from Overfitting," Journal of Machine Learning Research 15, 2014. Consolidated dropout.
  10. Hinton, G., Vinyals, O., and Dean, J., "Distilling the Knowledge in a Neural Network," arXiv:1503.02531, 2015. Introduced knowledge distillation.
  11. Shazeer, N., Mirhoseini, A., Maziarz, K., Davis, A., Le, Q., Hinton, G., and Dean, J., "Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer," arXiv:1701.06538, ICLR 2017. Revived mixtures of experts at scale.
  12. Sabour, S., Frosst, N., and Hinton, G. E., "Dynamic Routing Between Capsules," arXiv:1710.09829, NeurIPS 2017. The capsule network proposal.

Views

Hinton's public positions on AI risk date chiefly from his 2023 departure from Google. In the May 2023 New York Times interview announcing his resignation, he said he wanted to "talk about the dangers of AI without considering how this impacts Google," listed near-term harms including misinformation and technological unemployment, warned that AI systems might soon exceed human capability, and remarked that part of him now regretted his life's work.9 The same week he signed the Center for AI Safety's statement that "mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war."25 He also clarified publicly that he had not left to criticize Google, which he said had acted responsibly.26 Accepting the Nobel Prize in December 2024, he estimated in remarks reported at the time a 10 to 20 percent chance that AI would take control from humanity within thirty years, and he has repeatedly called for government-mandated safety research and for limits on the release of the most capable systems.27

These positions are contested within the field he helped found. LeCun, his former postdoctoral researcher and Turing co-recipient, has publicly and repeatedly argued that warnings of existential risk from current technology are premature and that risks are best managed through open research and better architectures rather than restrictions on development.28 Hinton's rejoinder, in interviews through 2024 and 2025, is that digital intelligence has inherent advantages over biological intelligence in sharing learned knowledge, and that the alignment problem will not wait for better architectures.279

See also

References


  1. Rumelhart, D. E., Hinton, G. E., and Williams, R. J., "Learning Representations by Back-Propagating Errors," Nature 323, October 1986. 

  2. Ackley, D. H., Hinton, G. E., and Sejnowski, T. J., "A Learning Algorithm for Boltzmann Machines," Cognitive Science 9(1), January 1985. 

  3. Hinton, G. E., Osindero, S., and Teh, Y.-W., "A Fast Learning Algorithm for Deep Belief Nets," Neural Computation 18(7), July 2006. 

  4. Krizhevsky, A., Sutskever, I., and Hinton, G. E., "ImageNet Classification with Deep Convolutional Neural Networks," Advances in Neural Information Processing Systems 25, December 2012. 

  5. ACM, "Fathers of the Deep Learning Revolution Receive ACM A.M. Turing Award," Association for Computing Machinery, March 2019. 

  6. The Nobel Prize, "The Nobel Prize in Physics 2024," press release, NobelPrize.org, October 2024. 

  7. "Geoffrey Hinton," Wikipedia, accessed July 2026. 

  8. Vector Institute, founding announcements naming Geoffrey Hinton chief scientific advisor, March 2017. 

  9. Metz, C., "'The Godfather of A.I.' Leaves Google and Warns of Danger Ahead," The New York Times, May 2023. 

  10. Hinton, G. E., Osindero, S., and Teh, Y.-W., "A Fast Learning Algorithm for Deep Belief Nets," Neural Computation 18(7), July 2006. 

  11. Hinton, G. E., and Salakhutdinov, R. R., "Reducing the Dimensionality of Data with Neural Networks," Science 313(5786), July 2006. 

  12. Hinton, G., Deng, L., Yu, D., Dahl, G., Mohamed, A., Jaitly, N., Senior, A., Vanhoucke, V., Nguyen, P., Sainath, T., and Kingsbury, B., "Deep Neural Networks for Acoustic Modeling in Speech Recognition," IEEE Signal Processing Magazine 29(6), November 2012. 

  13. Hernandez, D., "The Man Behind the Google Brain: Andrew Ng and the Quest for the New AI," Wired, May 2013; acquisition figure as widely reported. 

  14. Hinton, G., Vinyals, O., and Dean, J., "Distilling the Knowledge in a Neural Network," arXiv:1503.02531, March 2015. 

  15. Sabour, S., Frosst, N., and Hinton, G. E., "Dynamic Routing Between Capsules," arXiv:1710.09829, NeurIPS 2017. 

  16. Shazeer, N., et al., "Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer," arXiv:1701.06538, ICLR 2017. 

  17. Jacobs, R. A., Jordan, M. I., Nowlan, S. J., and Hinton, G. E., "Adaptive Mixtures of Local Experts," Neural Computation 3(1), 1991. 

  18. Hinton, G. E., "Training Products of Experts by Minimizing Contrastive Divergence," Neural Computation 14(8), August 2002. 

  19. LeCun, Y., Bengio, Y., and Hinton, G., "Deep Learning," Nature 521, May 2015. 

  20. Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R., "Dropout: A Simple Way to Prevent Neural Networks from Overfitting," Journal of Machine Learning Research 15, 2014 (arXiv:1207.0580). 

  21. Hinton, G., Srivastava, N., and Swersky, K., "Neural Networks for Machine Learning, Lecture 6a: Overview of Mini-Batch Gradient Descent," Coursera lecture slides, 2012. 

  22. Müller, R., Kornblith, S., and Hinton, G., "When Does Label Smoothing Help?," arXiv:1906.02629, June 2019. 

  23. Chen, T., Kornblith, S., Norouzi, M., and Hinton, G., "A Simple Framework for Contrastive Learning of Visual Representations," arXiv:2002.05709, ICML 2020. 

  24. Hinton, G., "The Forward-Forward Algorithm: Some Preliminary Investigations," arXiv:2212.13345, December 2022. 

  25. Center for AI Safety, "Statement on AI Risk," May 2023. 

  26. Taylor, J., and Hern, A., "'Godfather of AI' Geoffrey Hinton quits Google and warns over dangers of misinformation," The Guardian, May 2023. 

  27. Hinton, G., Nobel Prize press conference and interview remarks, Stockholm, December 2024; coverage by BBC News and the Associated Press. 

  28. LeCun, Y., public statements and interviews, 2023-2025; see for example Heikkilä, M., "Computer scientist Yann LeCun: 'Intelligence really is about learning'," Financial Times, November 2025.