This week we turn to ideas of machine learning that borrow more directly from theories of human learning.
As with Week 2, a couple of the papers (Ouyang et al. 2022; Guo et al., 2025) are again quite technical. I’d encourage skimming each, and paying greater attention to the methods of training. Ouyang et al. (2022) for example applies Reinforcement Learning from Human Feedback to refine GPT-3 into GPT-3.5 – becoming, by the end of 2022, ChatGPT. I’d suggest reviewing Guo et al.’s discussion of DeepSeek-R1-Zero – a key moment in how *unsupervised* reinforcement learning is now becoming possible. Also worth glancing at is the DeepMind paper on AlphaZero – a seminal moment in unsupervised learning. Finally, on a technical front, the Delua paper describes the general differences between supervised and unsupervised learning.
I’ve included two psychology papers: one by Skinner, describing “operant behavior”, a classic paper of behaviouralism that reminds us – and helps to inspire – supervised learning. Kohler’s discussion of Pavlov’s dogs belongs to a similar tradition. Dewey’s approach to experiential learning seems similar, instead, to unsupervised learning.
Finally, Cope & Kalantzis’ history of cybernetics and the early days of AI makes some of these connections more explicit.
Again, there’s a lot - and no need to read everything. Particularly on the technical material, feel free to scan and select sections that seem to echo earlier theories of human learning.
A provocation for next week: machine learning borrows from theories of learning devised in the early part of the twentieth century. What has changed in how we view human learning since? And – if those changes are significant – why aren’t machines trained upon these updated theories and lessons?
Readings:
Cope, B., & Kalantzis, M. (2022). The Cybernetics of Learning. _Educational Philosophy and Theory_, 54(14), 2352-2388. https://doi.org/10.1080/00131857.2022.2033213
DeepMind, G. (2019). _Alphazero: Shedding new light on chess, shogi, and go_. https://deepmind.google/discover/blog/alphazero-shedding-new-light-on-chess-shogi-and-go/
Delua, J. (2021). Supervised versus unsupervised learning: What’s the difference? IBM. https://www.ibm.com/think/topics/supervised-vs-unsupervised-learning
Dewey, J. (1986, September). Experience and education. In The educational forum (Vol. 50, No. 3, pp. 241-252). Taylor & Francis Group. https://www.schoolofeducators.com/wp-content/uploads/2011/12/EXPERIENCE-EDUCATION-JOHN-DEWEY.pdf
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., … & He, Y. (2025). Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning. https://arxiv.org/abs/2501.12948
Kohler, I. (1962). Pavlov and his dog. _The Journal of Genetic Psychology_, _100_(2), 331-335. - https://doi.org/10.1080/00221325.1962.10533601
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., … & Lowe, R. (2022). Training language models to follow instructions with human feedback. _Advances in neural information processing systems_, _35_, 27730-27744. https://proceedings.neurips.cc/paper_files/paper/2022/file/b1efde53be364a73914f58805a001731-Paper-Conference.pdf
Skinner, B. F. (1963). Operant behavior. _American psychologist_, _18_(8), 503. http://pdfs.semanticscholar.org/36fd/0131b5ae8f78db85b321ef93da67f6e3c534.pdf