top of page

How I Became an AI Safety Person

Sep 6
12 min read

Updated: Sep 12

I'm currently dedicating my life to AI safety: trying to lower the chance that AI leads to human extinction and other catastrophic harm. I think AI risk is the most important problem in the world, so I'm proud to be working on it.


Why am I so worried about AI?


I think there are a variety of persuasive reasons to care about AI safety. AI companies are racing toward making AIs that are smarter than us but aligned to weird alien values. Experts are very concerned. But this post is not primarily about the rational reasons to care about AI safety. Instead, it's about my personal story.


This is how I, Jacob, started caring, and how it enveloped my life, and how I became an "AI safety person". Maybe you'll realize it's who you want to be, too.



2019: Not An AI-Worrying Prodigy

These days, when people worry about AI killing everyone, they often go to a website called LessWrong. I first stumbled upon LessWrong in 2019! So you might think that I was an AI-worrying prodigy. 


But in those days, I found the concept of AI utterly boring. 


(My dad was the AI enthusiast in the family, pushing me to watch Star Trek or Terminator, buying the new Amazon Alexa ... I was the household Luddite, fine living without new technology, at least until I got acclimated.)


When I saw a LessWrong post about AI, I would skim past it in order to get to my favorites: the posts about cognitive biases, and game theory, and odd topics, written by self-taught man with a cult following and the name of Eliezer Yudkowsky.

"Eliezer Yudkowsky over on LessWrong is really cool. I wouldn't recommend it to everyone, as it's a bit technical and jargony, but it's definitely insightful and deep if you can get past that part." —Me in 2019

Yudkowsky is better known as the world's foremost AI "doomer". But I didn't know anything about Yudkowsky's AI opinions at first. I just knew his writing was good. (In 2019, I wrote on my blog: "Eliezer Yudkowsky over on LessWrong is really cool. I wouldn't recommend it to everyone, as it's a bit technical and jargony, but it's definitely insightful and deep if you can get past that part. The link is right here.")


Instead of reading about AI risk, I would read about cognitive biases, like scope insensitivity.
Instead of reading about AI risk, I would read about cognitive biases, like scope insensitivity.

I even read 40 chapters of his Harry Potter fanfiction, Harry Potter and the Methods of Rationality. My favorite detail I learned is about the insane inflation of the Galleon. (Did you know that in Book 1 of Harry Potter, the Weasleys' Gringotts vault has only one Galleon, and Harry's wand costs seven Galleons, but by Book 4, Fred and George are buying a joke wand for five Galleons.)


This diagram sums up my beliefs about AI risk in my early days reading LessWrong:


2021-22: Twenty Thousand Years of Torment

During the pandemic, I was very disciplined about going on a walk every day. Over the course of these walks, I binged the entire catalog of CGP Grey and Brady Haran's podcast Hello Internet. One of the episodes was called 20,000 Years of Torment, which contained a review of the book Superintelligence: Paths, Dangers, Strategies by Nick Bostrom (which I added to my list of books to read). Apparently it was the first thing that made Grey really scared about technology. Maybe I should relisten to it, I'm kind of curious how it holds up.


I guess it inspired me to write this poem called "Twenty Thousand Years of Torment." It's on my Google Drive, anyway. Here's how it begins:


artifice*

intelligence

building up, those neural nets

harmless now, they surely seem

advantageous fantasies


(twenty thousand years of torment)


but Thunderheads**, they shall not be

optimizing life for we***

neurologic treachery

just-world cause fallacy

humanity incidentally condemned... to


TWENTY THOUSAND YEARS OF TORMENT

and you romanticize the pain?


*oh my gosh is this really how it begins

**I was reading the Scythe series, which features a mostly-benevolent AI called the Thunderhead

***it's "we" instead of "us" for ... the rhyme scheme ...


Aaaaaand okay that's enough poetry for now.


Well, I kept reading LessWrong, and it's hard to escape AI forever, even if sometimes the arguments are in intimidatingly verbose prose. Eventually, I came across an Eliezer Yudkowsky video (maybe this one? I'm not sure) where my fear about AI really sunk in. If an AI was trained to recognize strawberries, but in its training data strawberries were the only red things, then outside the training data, you might tell it to get a strawberry and it gets, like, human blood or something. This was an oversimplification. It was all so complicated. But I definitely was left with this fear that future AIs would destroy humanity not out of malice, but just because we were in the way.


But I was left with this feeling of worry. It definitely seemed far-fetched, but some of the smartest intellectual figures online were saying that AI alignment was a big deal.


(Tangent: "AI alignment", to be specific, is the problem of aligning AI values to human values, instead of some weird alien values. Of course, there's still the problem of figuring out which human values to align them to ... the values of SF tech billionaires? of the US military? of the CCP? ... but even any human values at all is a hard problem. "AI safety" is a bit more general than AI alignment, and includes other solutions like controlling misaligned AI or an international treaty to ban superintelligence.)


I distinctly remember that I nearly wrote a blog post talking about how AI alignment might be the most important issue no one was talking about.


But I didn't write it then because I didn't feel I understand the risk enough to explain it. 


I'd be biking home from the train station, thinking about AI, try to imagine what I'd write ... but I didn't feel like I could do better than linking Yudkowsky's prose. I didn't think I could respond to the objections.


I should have written it, I suppose. You could have been a bit more informed about the future, and for my part I would have looked so prescient. But I didn't.


2022-25: After ChatGPT

When I went to the library, I would often check the computer to see whether they happened to have any of the books on my list in stock. By coincidence, in late 2022, they had Superintelligence, so I checked it out. Just then, ChatGPT came out.


Here are some things I did during that period of time:

  1. Wrote on my LinkedIn that I started worrying about AI alignment six months before it was cool

  2. Kept reading LessWrong, including unnerving but crazy-seeming future scenarios like Daniel Kokotajlo's "What 2026 looks like"

  3. Took a machine learning class

  4. Applied to PAIR camp twice (they rejected me twice), and wrote on the application that "While I read news about large language models with interest, I haven’t really found them helpful in my daily life yet (too much hallucination!), so most of my experience with them comes from this ML class."

  5. Read the entirety of The Power Broker, and thought about the parallels between the New York government delegating power to Robert Moses and humanity delegating power to unaligned AI (I saw a tweet recently making this same analogy which made me so happy but now I can't find it)

  6. Wrote on my Stanford application that the most significant challenge society faced was "Harnessing advanced AI for mass societal benefit instead of letting it harness us. There’s great uncertainty on this issue—even experienced domain experts disagree—but I find this even more worrying. “AI alignment” will likely require both cross-disciplinary AI research and working together across the political divide to implement our best strategies."

  7. Got really into prediction markets, and went to Manifest where I met Eliezer Yudkowsky which felt crazy

  8. Was bored during the summer and randomly decided to volunteer at a theoretical AI alignment conference in Berkeley called ILIAD, where I was hopelessly confused and bonded with one of the other volunteers, Jo Jiao, over our shared impostor syndrome ... but I did have a great conversation where I learned about sparse autoencoders!

     On sign-making duty at the ILIAD conference, where I met Jo Jiao
    On sign-making duty at the ILIAD conference, where I met Jo Jiao
  9. Went to a bizarre auction party in Berkeley where I randomly spent the equivalent of $10 to win a year's worth of aerial silks lessons from someone named Sydney Von Arx, then got starstruck when I saw her discussing "Moore's law for AI" on Computerphile

  10. Developed a hobby of volunteering at Berkeley AI conferences, and found myself at The Curve 2024 (I took the Big Game shuttle ... my main memories are forecasters writing scenarios of crazy things like AGI by 2028 and Dyson spheres and such, and also having an interesting conversation with Evan Hubinger about alignment faking)

  11. ...and I also went to AI for Animals (now called Sentient Futures) which was very thought-provoking

  12. Started going to Stanford AI Alignment club meetings (learning about AI control from Buck Shlegeris!), but not that frequently

  13. My ILIAD friend Jo told me to come to a conference in Chicago for AI safety undergrads ... I felt underqualified ("I'm hardly an AI safety undergrad!"), but I would have gone if I didn't have a scheduling conflict

  14. Had a few spare days in the Midwest, so ended up visiting my ILIAD friend Jo at UChicago later. The first night I got there, she was like "Hi Jacob! I'm going to my friend's birthday party. You're not invited, but you can have dinner with my XLab friends." Then I had a great time! Playing poker with her XLab friends, curled up on a fifteenth-story windowsill trying to read "A Mathematical Framework for Transformer Circuits" ... those few summer days, more than almost anything else, made me feel like I could be an AI safety person

    Trying to make sense of "A Mathematical Framework for Transformer Circuits" on a windowsill in UChicago
    Trying to make sense of "A Mathematical Framework for Transformer Circuits" on a windowsill in UChicago
  15. Came back to Stanford for sophomore year, down of all my high school interests (puzzles, math, linguistics), and decided I wanted to lock in on AI safety


Late 2025: A Relatively Unintelligent Person Reviews A Book About Superhuman Intelligence

I preordered If Anyone Builds It, Everyone Dies, co-written by my old friend Eliezer Yudkowsky and Nate Soares, because I thought it would be "important" — a worthy book I'd have to turn on Serious Important Book Mode to read. 


I didn't expect it to be a page-turner. 


But it is! Most people don't say this, but IABIED reminded me of master nonfiction storytellers like Malcolm Gladwell and Randall Munroe. (I think it's the combination of Yudkowsky and Soares that made IABIED so good. It's often said that Eliezer wrote 300% of the book, and Nate wrote the other -200%, and I absolutely believe it.)


It's a book about ice cream and condoms. It's a book about nuclear meltdowns and cyber hacking. It's a book about the nature of intelligence itself. 


Reading If Anyone Builds It, Everyone Dies was a strange experience. Every moment where I felt the rush of understanding the arguments that had confused me — where I was like "wow, these authors are intelligent" — was a moment that increased my belief in apocalypse. 


A New York Times bestseller.
A New York Times bestseller.

This was another moment when I very nearly wrote a blog post about AI safety — it was gonna be a review of this book, to accompany The Guardian's review ("everyone with an interest in the future has a duty to read") and The New York Times' dumb review ("the most annoying students you met in college when they try mushrooms for the first time") and Asterisk's review ("More was Possible") and Kelsey Piper's review ("It is tempting when writing a book review to treat the book as a moment for reflection: on the author, the topic, the cultural moment. But when the book in question’s central claim is that we are all going to die horrible deaths if tech companies succeed at their current plans, the only question that matters is whether the book is correct") and Scott Alexander's review and Nina's review of Scott's review and GradientDissenter's review of Nina's review of Scott's review — but then I got busy with school, and again my attempt to write a blog post about AI safety fell by the wayside.


But I gave a talk at Stanford Dorm Lectures about this book! I titled it "A Relatively Unintelligent Person Reviews A Book About Superintelligence." I talked about the book and how I felt like Eliezer Yudkowsky was smarter* than me, but so were the experts who disagreed, like Fei-Fei Li or Yann LeCun. What was I supposed to believe in a world where I could be convinced of different things by different people? The talk was well received, I think, partially because I had people move around the room based on their p(doom), and partially because I brought Taki's as a reference to the evolutionary analogy that humans like all kinds of things that aren't just the food that gives them the most nutrients. But it really got people talking.


*Actually, SICKer, not smarter. I invented an acronym, SICK, that stood for Situational Intelligence / Contextual Knowledge, which captures what both the experts and the AIs would be ahead of me on.


I was applying to all manner of AI safety programs, but I hate applying to things, so I was scheming about what I could do instead. I thought I would become an AI safety YouTuber, call my channel "Jacob Doesn't Understand AI", and try getting interviews with the experts I'd met in Berkeley and San Francisco. I think I said I would release one video per day during Thanksgiving break.


I did manage to release an interview with Duncan Sabien, the author of "Deadly By Default" (one of my favorite essays explaining the case for AI risk), where I couldn't really think of good arguments in response.


But I put my YouTuber schemes on hold. (I didn't really feel like being a YouTuber anyway, I just felt like I should do something and it seemed more fun than applying to things.) I had actually gotten accepted to one of the AI safety programs I had been excited about. I was going to England. I was going to MARS!


December 2025: MARS

I flew across the ocean to do ambitious mechinterp research. ("Mechanistic interpretability, or "mechinterp", for short, is about understanding how AI systems work on a gears level.) The Holy Grail of AI alignment. I was convinced I was a personality hire because I didn't even do the task that I was supposed to do on time for the application. (Later, my mentor said that he picked me for my "neuroplasticity.") It was my first time in Europe; we would have one in-person week in Cambridge, and then we would work on our projects asynchronously.


The most iconic person in mechinterp is named Neel Nanda. (Okay, possibly Chris Olah, but those are really the two). On the very first day of MARS, he published a blog post, "A Pragmatic Vision for Interpretability", in which he explained that he was pivoting away from the more ambitious mechinterp work (trying to fully understand how AIs worked) as opposed to just doing things that wouldn't necessarily be on the road to full understanding but might be useful.


It was a portent. Mechinterp, so trendy, so the-thing-to-do, no longer felt like the thing to do. But I worked on it for 3 months, mostly feeling like I wasn't doing a very good job. I applied to more AI safety programs in the summer. MATS didn't take me, but MIRI did.


After the underwhelming release of GPT-5, it feels like maybe AI was hitting a wall? But in April, my world changed forever.


April 2026: The Month It Got Real

This is something I wrote down at the time:


I think April 2026 will go down in history (aka my ultimate personal spreadsheet) as The Month It Got Real...


April is the month AI became clearly a cyberweapon, the month that AI anxiety kept me from sleeping, the month where it became clear that risks I’d once only read about from nerds on a weird “rationalist” corner of the internet were being real.


April is the month I had [maybe the best day of my life] by the Yuba River realizing that life can be AWESOME and I really don’t want to die.


April is the month I became co-president of [Stanford AI Alignment club] — a platform where I could 2x or 10x my AI safety impact by getting some of the world’s smartest, most driven people working in impactful AI safety jobs instead of quant internships or GPT-wrapper startups.


April is the month where I got accepted to my very own impactful AI safety job at MIRI. Ok I lied, it was March. April was when I uhhhh joined the MIRI slack? My pattern is failing here.


April is the month where I've figured out my housing plans to move to Berkeley for the summer with AI safety friends and soon-to-be-friends. This is very exciting!


I’m having trouble enjoying normal life as much, though. I guess I feel like when I was young and I learned about all these large-scale problems, like factory farming or climate change, I wanted to make a difference but realistically I could only help at the margins.


But now somehow I have ACTUAL POWER??


I feel like the world is burning and somehow I became one of the firefighters … because I guess there aren’t enough volunteers? and somehow they couldn’t find anyone better??


Stressful


But I’m gonna have to figure out how to deal with it, because this is obviously the direction I want to take my life in for now. It’s just weird to be so extrinsically motivated, to be working on something not because I am passionate about it deep down but because I genuinely believe that it lowers the chance we all fucking die for real (or similar bad outcomes) by like 0.001% or whatever.


Summer 2026: Berkeley

Everyone in my shared house this summer. Works on AI safety. There's Jo (my first safety friend who I'm living with now!) who works at Redwood; Miles, who works at the AI Futures Project; Ariana, who was an Anthropic Fellow; Helena, who was working at Constellation (now METR). I consider them among my best friends in Berkeley.


The world kept getting crazier. The Hugging Face incident is the big famous one, but so many little things. AI risk is mainstream. And it feels like we might actually live in AI 2027 or something close (maybe more like AI 2028).




Many have pointed out the irony behind Steven Pinker's latest tweet. Just a few years ago, the immediate and obvious threats from AI were things like algorithmic bias and environmental impacts. These issues are real, but they'll pale in comparison to an AI-assisted pandemic, massive societal-scale hacking, and the very real possibility of human extinction in the next couple years.


It's crunch time in AI. You can get into AI safety in 3 months. Come join?


Note: I'm running a blog-a-thon today, and I am therefore deontologically obligated to publish this post, but I don't consider it done. I thought the HIPPOCAMPUS one was satisfactory but this is always a topic I've struggled to write about; you win some, you lose some. I plan to make further edits to it in the future. Please bother me to do this.

Comments


Logo art: The Magic: The Gathering card "Mindshrieker" illustrated by Dave Kendall. It's not that good or interesting a card, I've just always loved the art! 

 

©2019-23 by Chromatic Conflux. Created with Wix.com

bottom of page