Reading notes
AI Safety Journal
A selective journal of recent publications in AI safety, AI welfare and the philosophy of AI, with concise summaries and commentary.
-
The Id, the Ego and the Superintelligence
Critchley argues that the hardest questions in advanced AI concern judgment, meaning and value, helping to explain why leading AI laboratories are hiring philosophers.
Summary
The article argues that AI’s capabilities strengthen rather than eliminate the need for philosophy. Systems can retrieve, summarise and recombine knowledge, yet the author distinguishes these abilities from judgment: the human capacity to decide what matters, weigh competing values and respond meaningfully to circumstances. Judgment is formed through lived experience, fallibility, memory and forgetting, and encounters with mortality, grief, boredom and social pressure – conditions an algorithm cannot acquire by absorbing internet data. This supposedly helps explain why Anthropic, Google DeepMind and OpenAI are bringing philosophers into work on machine character, ethics and principles for model behaviour. AI may perfect “artificial memory,” fulfilling an anxiety that reaches back to Plato (Phaedrus), but intelligence is not simply recall. Genuine thought requires selection, interpretation and concern. The central claim is that engineering alone cannot determine what intelligent systems ought to become; development must remain connected to philosophical reflection on meaning, value and judgment.
-
The Revenge of the Philosophy Majors
A survey of the new roles philosophers are beginning to play in frontier AI research – and of the limits of that emerging labour market.
Summary
The article traces demand for philosophers inside frontier AI labs. Through figures including Robert Long, Jeff Sebo, Amanda Askell and Henry Shevlin, Wallace shows how problems once confined to seminars now shape decisions about alignment, model behaviour, consciousness and AI welfare. Philosophers contribute by clarifying concepts, testing assumptions, distinguishing possibilities and reasoning under moral uncertainty – skills that complement, rather than replace, engineering. The piece also describes a reversal in philosophy’s employment story: although conventional academic jobs remain scarce, organisations working on advanced AI value people trained in ethics, epistemology, philosophy of mind and decision theory. Yet the trend is narrow and specialised. A philosophy degree alone is not a general passport into technology; the strongest demand is for researchers who can connect rigorous philosophical analysis with technical literacy and institutional questions. The article presents AI as both an object for philosophy and a possible new labour market for philosophers.
A Philosopher’s Thoughts
The article may mistakenly infer from the particular to the general in predicting improvements in philosophers’ employment prospects. Aaron Kagan offers a counterargument, claiming that very few philosophers have, in fact, been hired and that only one of these roles—Henry Shevlin’s at Google DeepMind—actually carries the title “Philosopher”. Kagan's take may be somewhat misleading, however, because even without a systematic search, I can identify several philosophers in relevant roles: Amanda Askell (Anthropic), Kyle Fish (Anthropic), Geoff Keeling (Google Research), Iason Gabriel (Google DeepMind), and Arianna Manzini (Google DeepMind).
Whilst only deep-pocketed AI laboratories may have both the resources and the incentive to ethically upskill, or more cynically “philosophy-wash”, most companies approach AI risk more narrowly, as a matter of regulatory compliance and avoiding litigation or criminal liability. Nevertheless, I think demand is growing for senior staff with philosophical training to occupy important roles in AI ethics and governance, and this is probably a good thing.

A Philosopher’s Thoughts
One objection might be that the contrast between machine reasoning and human judgment may be too categorical. Human judgment is shaped by pattern recognition, inherited concepts, social learning and imperfect memory, while AI systems can already be trained to rank values, revise conclusions and respond to context. Lacking mortality or grief does not by itself show that a system cannot exercise a functionally meaningful form of judgment. The stronger case for hiring philosophers is therefore not that AI can never judge, but that deciding what should count as good judgment is irreducibly normative.