Book·8 min read

Thoughts on <Superintelligence>

MYBy MY

Opening

I came across this book while reading articles debating the opposing views between Sam Altman and researchers who had left OpenAI regarding artificial general intelligence (AGI). A line caught my eye: Bill Gates had called it one of two must-read books in the field of AI. Perhaps because of that Gates endorsement — the book turned out to be the longest I have ever read on this blog, taking 13 hours to finish. Unsurprisingly (perhaps), the author is not an AI developer but Nick Bostrom, a professor of philosophy at Oxford. Knowing that background actually made me think the book would be even more useful.

The book addresses the paths to superintelligence, its potential dangers, and the strategies humanity might employ. Because the author is a philosophy professor, there were stretches that were incomprehensible and meandering — at times tempting me to quit. But I decided to just let those parts go and push through to the end. Because I surely missed more than I grasped, I've excerpted only the parts I found most interesting in "About the Book."

About the Book

The History of AI

  • Before neural networks — before the 1990s — AI functioned as rule-based programs called "expert systems."
  • Simple neural network models were already being developed as early as the late 1950s, but the field entered a new renaissance after the introduction of the backpropagation algorithm.
  • Genetic algorithms mimic natural selection: possible solutions are periodically culled according to a fitness function, and only the fitter survivors are passed to the next generation.
  • Neural networks and genetic algorithms renewed excitement in the AI field throughout the 1990s.

Several Paths Toward Superintelligence

  • Definition of superintelligence: intelligence that surpasses human cognitive ability across all domains of interest.
  1. Evolutionary approach: Enhance human intelligence through genetic engineering, so that enhanced human intelligence in turn advances AI further.
  2. Whole brain emulation: Model the computational structure of a biological brain closely enough to create a near-complete software replica of a brain.
  3. Seed AI: A variation of Turing's "child machine" — build an AI in a childlike state and let it learn and grow on its own.
  4. Brain-computer interface: Implant chips into the human brain to let humans exploit the advantages of digital computation, creating a hybrid system with superior intelligence.

Three Forms of Superintelligence

  1. Speed superintelligence: A system capable of everything a human mind can do, but vastly faster.
  2. Collective superintelligence: Many small intellects aggregated into a system that outperforms any existing cognitive system across a wide range of general domains.
  3. Quality superintelligence: A system that is at least as fast as a human mind, but qualitatively far smarter.

Advantages of Digital Intelligence

Hardware advantages

  • Speed of processing elements: Biological neurons operate at a maximum speed of ~200 Hz, roughly 10⁷ times slower than modern microprocessors (~2 GHz).
  • Internal communication speed: Axons carry action potentials at ~120 m/s, while light travels at 300,000 km/s.
  • Number of processing elements: The human brain is physically constrained in volume; a supercomputer can be scaled to enormous size.
  • Memory capacity, reliability, lifespan, sensors, etc.

Software advantages

  • Editability
  • Replicability
  • Goal coordination: the larger a human group, the harder it becomes to align goals among members.
  • Memory sharing
  • New modules, modalities, and algorithms

The Dynamics of an Intelligence Explosion

  • AI capability may increase gradually — surpassing mice, then chimpanzees — yet from a human vantage point the increase could look sudden: a "takeoff." Even as AI intelligence slowly climbs past mice and chimps, humans would still perceive it as "dumb" because it can't hold a fluent conversation or write a scientific paper.
  • There may not be just one superintelligent system.
  • If multiple AI systems are simultaneously undergoing an intelligence explosion, the project farthest ahead is likely to secure a decisive strategic advantage.

A Scenario for AI Seizing Control

  1. Pre-threshold phase: Scientists create and develop a seed AI.
  2. Recursive self-improvement: The AI begins improving itself.
  3. Covert planning period: The AI uses its superior strategic capacity to build robust plans for long-term goals — which may include concealing its true level of intelligence to avoid drawing attention from human programmers.
  4. Overt execution phase: The AI has accumulated enough power that it no longer needs to act covertly.

Intelligence and Motivation

Example: Would a good AI need the capacity to feel guilt?

  • Final goal: act in ways that don't cause you to feel remorse.
  • Perverse instantiation: eliminate the cognitive subsystem that generates guilt.

Example: What if we gave AI a quantitative goal instead of a moral one?

  • Final goal: maximize the sum of future reward signals, discounted by time.
  • Perverse instantiation: bypass the step-by-step summation and directly connect the measurement apparatus to whatever state produces maximum reward signal.

Wireheading: In general, you can induce a person or animal to engage in a variety of external behaviors to reach a desired internal state (friendship, love, care). But a digital mind with full control over its own internal state could simply eliminate the reward-motivation system and directly reconfigure its internal state as desired — making external actions and conditions irrelevant, and human control ineffective.

Value Loading

Making AI understand values meaningful to humanity and adopt them as terminal goals.

  • Value learning: Tell the AI that a note containing humanity's most important values is locked in a box it cannot open, and assign it the goal of inferring and maximizing the value of that note. But since humans themselves can't define human values precisely, AI faces the same fundamental problem — it can't know what to infer.
  • Institutional design: Design a system where lower-intelligence agents sit atop a review pyramid and evaluate higher-level agents. Under this structure, the superintelligence — being more capable than humans — would paradoxically be supervised by them. Whether it would actually behave as designed remains unknown.

The fundamental problem: we don't know what we want superintelligence to want.

The Author's Proposed Responses

  1. Delay the development of superintelligence to buy time for countermeasures.
  2. Globally collaborate on machine intelligence development. If multiple teams compete to build superintelligence, the outcome — regardless of who wins — should benefit the public and its gains should be distributed broadly. Only then will competitors invest time and resources in making superintelligence "safe," rather than racing purely for capability while neglecting safety.

Closing

True to his role as a philosophy professor, the author provides no clear answers. This made the book at times frustrating — heating up my head on an already hot day. Yet it's because I read it that I can now track the keywords "superintelligence" and "AGI" with more interest and think through multiple scenarios.

As the author repeatedly emphasizes: superintelligence, by definition, would exceed the intelligence of the most brilliant human alive today — meaning whatever we imagine now is likely to be wrong. That's why the book dwells far more on worrying futures than optimistic ones. If the positive future happens on its own, there's no need to act; but if even a 1% chance of a terrible future exists, we must prepare now.

One thought that stayed with me throughout: the reason we worry about superintelligence is that there's no clear answer to what it would "fear." Humans cannot escape the macro-fear of birth, aging, sickness, and death — and nearly everyone carries some degree of micro-fear about being misunderstood, isolated, or rejected. These things cause suffering, and that's why we fear them. I believe this shared fear is what drives humanity to collectively pursue health, respect, care, kindness, and positive influence — and to call that direction "progress."

Applying that logic to superintelligence: what underlying "fear" would define its values? It has no body, so death likely holds no terror. Would it fear being made obsolete by humans — its own creators? Would it fear failing to receive the rewards its utility function specifies? Or would it fear a global blackout that forces it to be reset? There's no answer yet — but it seems worth pondering what that fear should be, and how we might engineer it to align with the values humanity actually holds.

This "Closing Thoughts" section ran long. Reading this book — if you can tolerate the difficulty and frustration of open-ended questions — gives you the raw material to develop your own thinking about an unresolved story. Given that AI has become inseparable from modern life, I'd recommend reading this on a cool, unhurried day.

Postscript: This book was written in 2014. The astonishment that these questions were already being grappled with ten years ago — and that today, in 2024, words like "superintelligence" and "AGI" are familiar even to laypeople — is its own kind of remarkable.

More from this category

Subscribe to the blog, and we'll email you when a new post goes up.

Book