20 Comments
User's avatar
Duane McMullen's avatar

I read Y&S to be saying:

1. A superintelligence will destroy humanity.

2. Current AIs are probably/probably not superintelligent.

3. We do not know when we will cross the threshold into superintelligence as we throw increasing resources into developing more capable AIs.

4. We are, at best, playing at the edge. It is highly plausible that we will cross over into superintelligence BEFORE we figure out what the edge is.

In that sense, your review does not contradict the argument.

Chaperpne's avatar

2 makes no sense,

1 is dumb, 3 is dumb, 4 is dumb,

and the fact that you spent time typing this comment implies that you are about the same amount of dumb as “Eliezer Yudkowsky” and “Nate Soares”

Patodesu's avatar

The first points Y&S are really saying are:

A Superintelligence with current AI Alignment techniques will destroy humanity

We are super mega far from having good enough techniques

So he is conterarguing the arguments. The second one in particular.

Simon Lermen's avatar

I responded to one of your claims, ie that we may be able to use AI to solve the alignment problem: https://simonlermen.substack.com/p/why-i-dont-believe-superalignment

Chaperpne's avatar

You're dumber than Eliezer Yudkowsky. He at least has the sense to exhibit his dumbness on Twitter. You're so dumb that you didn't even attempt a summary of your wall of text that has a 100% chance of being dumb by highlighting the things in your post that you consider least dumb

Jono's avatar

I think AI2027 well described a more gradual scale-up in AI capabilities. It's still blindingly fast, but it would be weird (of course) to get AI good at ~everything (especially software) except AI R&D.

Though I think IABIED's story is less about sudden increase and more about "what if the guardrails fail?" both with some detail about how failures would cascade and about how the end result would be a system that cannot be negotiated with (without any step being due to any irrationality of the system).

Rafael Ruiz's avatar

(Disclaimer: I'm only halfway through the book)

I think a key crux is whether takeoff will be fast or slow (continuous or discontinuous?). You say that we can use AI+Humans to align ASI-, then use ASI-+Humans to align ASI, and so on. But humans might not be able to be adding much to the conversation once we reach ASI. The analogy being that it's like putting a combination of Magnus Carlsen and Stockfish17 to make Stockfish18. Magnus Carlsen might have nothing to add to the conversation, and might even be dead weight at that point.

You might say that, well, at least humanity can put the humanistic values into the ASI. Like, Magnus can put its values and bend the way the future Stockfish18 will play. But that presupposes that technical alignment will be solved at least to a sufficient degree, and that no surprises arise from the behavior of Stockfish18 or its descendants once it's built. Once it's built, it might be running the show if it can do recursive self-improvement and process things must faster than humans.

Also, we have to assume that all ASIs will be aligned, or that the aligned ASI will destroy the others if they were to arise. If we have several ASIs by several different actors, some aligned, some misaligned, we would be living in a very fragile and unstable world.

For what it's worth, I think this is the most promising path to alignment. But it is a tricky path.

comex's avatar

But will we *need* to add to the conversation at that point?

Here’s a somewhat more optimistic chess analogy.

As you get better at chess, I think your move evaluation function mostly gets refined rather than replaced. One of the first things you learn in chess is not to give away your pieces for free. But if a beginner thinks a move is bad because it gives away pieces for free, chances are Magnus Carlsen will also think it’s bad, and so will Stockfish. Higher levels of play are mostly about finding the differences between moves that seem equally good at lower levels of play.

Sometimes that isn’t true. Sometimes Stockfish will make a move that seems outright bad to a lower-level player, seemingly giving away a piece for free – only to start some insane sequence that gives it an advantage 8 moves down the line. But this is rare. And when it does happen, the confusion is only temporary: the value of its move becomes clear once the sequence is finished.

The ASI equivalent would be something like: As ASI gets more advanced, it’ll start to come up with subtler and subtle moral preferences that we can’t understand or control. But those preferences will only come into play to distinguish between choices that seem morally equal (or near-equal) to us. In rare cases, the ASI will take actions that seem morally wrong to us at the time, but only because understanding their moral value requires predicting the future better than humans can. Those actions will be understandable in hindsight.

Again, this is an optimistic analogy. But I do think it’s a possible outcome, perhaps even a likely outcome.

Rafael Ruiz's avatar

Stockfish, if you let it think for a long time, can find checkmate in 16 and stuff like that. I don't think the heuristics human use are *that* similar to the way that Stockfish thinks. I don't think it usually thinks "Don't give pieces away", but "Nf3 has failed 37% of the time, and Be5 88% of the time, (and thousands of other combinations, anticipating several movements ahead), and the best one is Nf3, so I'll play Nf3"

I think your ASI equivalence already presupposes that we can solve the alignment problem. There's no reason an AI will be moral by default, and plenty of reason to think that such an alien non-biological intelligence might have very strange "preferences".

And, even if it's aligned, I mean, maybe it kills all humans because it sees what we're doing in terms of factory farming or other stuff that it considers morally atrocious.

DABM's avatar

Before AlphaZero, when chess engines were already extremely superhuman, they relied on a mixture of raw calculation of exact series of moves (which humans also do, just slow) and hard-coded heuristics similar to human ones. I don't *think* they actually did pattern recognition of the "this sort of move has succeeded x % of the time". Rather they used heuristics to prune which moves were important to check the result of when calculating long sequences (i.e. captures, checks, but also subtler things), as well as to evaluate the positions at the end of the calculated sequences for who is winning (i.e., things like generally bishops are a little better than knights, but this position is closed, so in fact knights might be slightly better here, therefore I will slightly tilt my eval towards saying the side with a knight and a bishop is doing better than the side with two bishops). For a long time, engines were (far) better than humans simply because they could do a huge amount of brute force calculations of exact sequences of moves, even as they were at best no stronger, and probably weaker than humans at evaluating things like subtle positional advantages, or what moves to look at first. Go can't be brute forced in this way to the same degree, being far, far more complex than chess, which is why it took ML techniques to produce a super-human bot.

Chaperpne's avatar

You, @Rafael Ruiz, are at least as dumb as Eliezer Yudkowsky, given the fact that you haven't considered even the basic “pull the plug” argument, and the fact (I know this, you don't, but I still blame you because you were dumb enough to write so many paragraphs without regurgitating Eliezer Yudkowsky's dumb counter-”pull the plug”-arguments) that the answer is indeed to pull the plug in a more sophisticated and cool fashion

Rafael Ruiz's avatar

There are plenty of arguments why “pull the plug” might not work, such as:

1. There probably won’t be a single plug to pull (there are different labs, after all)

2. We might not realize WHEN we need to shut it down, might be too late (AI might be decieving, scheming, or opaque)

3. A capable AGI might actively resist shutdown (e.g. copy itself out of the box)

4. Human vs AGI speed and capability are totally mismatched, particularly if we reach AGI "late" with incredibly powerful chips, like say, beyond 2040

5. You can’t be sure you’ve turned all of it off, particularly if they're making them in China. There's no central coordination

6. The catastrophic path may already be locked in in some near future (attractor state)

7. Humans and institutions might not want to pull the plug (due to economic incentives, or saying that it's going to be created in China soon anyways)

8. AI might persuade people to do it's bidding, particularly if it can do stuff like playing the stock market or crypto

9. “Pull the plug” doesn’t solve alignment, just postpones it. We will still live in a fragile world where technology keeps improving and creating AGI in more ordinary data centers gets easier and easier

9. Countermeasures like boxing or tool-AGI are fragile. And there's interest in AI being connected to the internet

10. This assumes that future AIs might be similar LLMs, but future AIs could follow a different architecture, some of which could be more dangerous

(Also, I didn't write that many paragraphs. Now I did! But before I had only made one single argument, so why so aggressive and condescending?)

Chaperpne's avatar

The aggression and condescension can't be avoided when a whole lot of people like yourself, starting with the authors of this book (‘Eliezer Yudkowsky’ and ‘Nate Soares’) start talking so loudly and confidently — enough to be given not just enough money to feed themselves and get by, but to further recruit, own property, etc. — about the production and ownership of ‘chips’ that can be used to build an ‘ASI’, inciting more gullible people into hunger strikes, unable to be completely ignored even by people like Bernie Sanders who obviously have far more important things to do with their time, all without taking even the minimum sufficient effort to understand what constitutes these chips and how they work, relying entirely instead on bs that you feed each other based on your own ignorant ‘research’.

In everything you've said here, you're arguing against yourself. I'll prove this by saying the same things in better ways than you. But first,

If anyone builds a computer powerful enough to either explicitly enact large-scale evil, or be too unpredictable, a lot of people will die. This has already happened with every type of technology that human beings have invented. A frat boy named Mark Zuckerberg got excited about an idea and built something that ended up catalyzing the suicides of so many children. You saw him turn and bow to the parents. You people (Yudkowsky, Soares et al) warn of nothing more serious than that.

What i’ve said in the previous paragraph is essentially what you say in points 2 and 8. It is too late. It has already happened. Now,

‘A capable AGI might actively resist shutdown (e.g. copy itself out of the box)’. Yes, that is the exact property endowed upon computer viruses that human beings build already. They are beaten by anti-virus software that runs on the same computers.

I've been to different parts of China. And Taiwan. They're making them there. I'm helping them a bit. They help me occasionally with my own stuff. But all of this collaboration happens in a very loose fashion. There is no central coordination.

A lot of things exist in China already. What's ‘going to be created’ is only more advanced things, like the full solution to chess, better protein-folding databases, cure to cancer (although a good number of jokes float around about that).

‘solve alignment, just postpones it’. https://ifanyonebuildsit.com/13/aligned-to-whom starts with the line ‘If humanity builds a superintelligence someday, we should make sure that it’s “aligned” with human values’. Those double quotes around that word are not mine. The existing of those double quotes says all that I need to say though.

None of you, starting with ‘Eliezer Yudkowsky’ and ‘Nate Soares', are going to understand electric currents, let alone electronics, if I don't start with an easier subject, because admittedly the way electric currents are taught in college is hard, whatwith Maxwell’s equations and everything.

So I'll tell you what a ‘plug’ is: it's like a faucet.

‘Countermeasures like boxing or tool-AGI are fragile’. Wtf are you trying to ‘counter’?

‘There probably won’t be a single plug to pull’. No, it was never like that. That's not how anything works. There never is a ‘master faucet’. Those things are called ‘dams’ and are not good analogies for plugs. Your problem is that you're old and your younger fellow-cultists have autism. I bet you're so old that you've seen this warning sent by superintelligence from the future flashing on your screen: “It is now safe to turn off your computer”. I'm joking. That warning was written by computer programmers who understand computer science. You lose your work if you irresponsibly shut off the dam. You never want to shut off the dam and prevent the water from entering the ocean and into the CLOUDs where it belongs.

Before I move on to electronics, i want to address your point 6. None of you, starting with ‘Eliezer Yudkowsky’ and ‘Nate Soares', have bothered to study the basics that are important to understanding AI. So I'm definitely not going to bother to find out what you've come to call ‘attractor’.

The place that I would recommend (this is somewhat subjective) to start learning about how computers work is this: https://en.wikipedia.org/wiki/Three-state_logic. I'm linking you here instead of to ‘triode’ or ‘transistor’, because if i do that, you'll fall asleep midway thinking, ‘there’s nothing in here that says how a bunch of transistors can constitute a T-1000 and a killing machine can emerge out of that, but it's so cool that it does, just give me the “sequences” according to which it'll kill me’. The word ‘binary’ is used by biologists in a less rigorous sense than computer scientists, but the number 3 is as important in electronics and computer science as the number 2. You start talking about neurons when you consider somewhat bigger numbers.

A triode has 3 tubes going into it, and is like a faucet controlled by hydraulics. Both of these are huge, of course. A transistor (or equivalently a NAND or NOR gate) is made of silicon. One of the 3 wires is almost always considered as some sort of output, but one of the inputs can be a plug.

Now, a programmer worth their salt can write a program with self-preserving routines that prevent plug-pulling at so many levels, which brings us finally to your point 4. I hope the analogies of ‘plug-pulling’ and surgical ‘faucet closing’ (the subject of more than one Resident Evil puzzle) have served their purpose of dispelling your preconceived idea that's more akin to ‘dam closing’. So I think we can change terminology to ‘wire-cutting’ which is not just closer to what people handling faulty circuits actually do, but also coincidentally sounds cooler. AI is trained on below-average human beings (eg. ‘Eliezer Yudkowsky’ and ‘Nate Soares') and above-average human beings alike, but built and maintained by competent computer programmers. I watched most of the same movies that you did: The Matrix, Terminator, iRobot, Appleseed, Nier: Automata, but I didn't watch them everyday, again and again, an order of magnitude more times than Hinckley watched ‘Taxi Driver' instead of going to high school and college and learning how computers are made to exhibit emergent behavior until my Mom kicked me out of her basement with a sufficient allowance (I'm guessing 6 figures USD?) to start a cult. But given that that's what the authors of this book did, I'm making this and the next paragraph a bit flashy in a way that should appeal to people like them and you.

If anyone builds a program that teaches itself how to safeguard itself from wire-cutting at every interface on multiple levels (I'm not in the mood to explain what I mean by levels), its wires can still be crossed. Because wire-crossing is exactly what even slightly sophisticated code does anyway. ‘Eliezer Yudkowsky’ and ‘Nate Soares' won't be able to do it. You won't be able to do it. But as long as a rogue AI is operating fully autonomously (with no malicious human being actively working to counter countermeasures and actively updating its code to that end), a Hokage-class human being always knows how to cross the wires of a fully autonomous AI to stop it from pursuing its preprogrammed malicious objectives, or any alien preferences it may develop along the way: https://youtu.be/ANF_ltuydmE?si

Howard Hansen's avatar

If someone builds it and we all die, will it matter?

Chaperpne's avatar

No, Devon. No one cares about you

And no, not everyone will die

Chaperpne's avatar

You're dumb, like all doomers

Devon Fritz's avatar

I am not a doomer - read about conditionals.

Chaperpne's avatar

Look up the word "and"