AI Safety Contribution Thesis

Last updated:

If you're reading this, it's probably to help me decide how to contribute to avoiding ASI misalignment; to that end, you can likely treat everything but the "So, what should I do" section as optional.

Foreword

This essay is an attempt to solidify my own opinion around how I can contribute to avoiding ASI misalignment, where "avoiding ASI misalignment" means avoiding human extinction or permanent mass suffering that occurs as a direct or indirect result of the effects of ASI. It's written in a confident style in order to make it easier for a reader to disagree, not because I'm confident of everything I've written.

There are important sub-definitions to "avoiding ASI misalignment" that deserve documents in their own right, like who can use the aligned ASI, and for what. I think these are mostly implementation details downstream from what I'm trying to decide right now, because I don't think these clarifications really change much about the overall problem (e.g. aligning an ASI to one human seems roughly as hard as all humans). More words in the appendix.

I also have a high-level posture that predicting the future is hard. Many folks in the AI safety community have collapsed to certain predictions of the future - like whether ASI alignment is easy, tractable, or intractable - and I more or less believe that all of these predictions do not meet the burden of proof required, especially given they are predictions over scientific progress that we have little precedent for. More words in the appendix.

Viable Approaches

There are two categories of ways to avoid the emergence of misaligned ASI: 1) achieving alignment of all ASI allowed to exist, or 2) permanently stopping ASI development.

Achieving ASI Alignment

Achieving ASI alignment requires such a thing to be achievable (AKA alignment is easy or tractable). Assuming it is, it requires the sort of technological breakthroughs it seems we aren't on track to achieve yet. It's possible that's not true, but since we have so little evidence one way or another, we can't be sure.

So if alignment is easy or tractable, by definition we'll get there by discovering and implementing the means of achieving that alignment, AKA doing alignment research.

This assumes that an aligned ASI, by definition, can and will monitor for and eliminate emergent misaligned ASI.

It also assumes that creating an aligned ASI is not "too much" harder than creating a misaligned one. In theory, humanity could dedicate all of its resources solely toward creating an aligned ASI, but in practice, we cannot count on that. Hand-wavily, 99% allocation of resources towards aligned ASI still loses if misaligned ASI is 100x easier to build.

Permanently Stopping ASI Development

If alignment is intractable, the only way humanity can see a good outcome is by permanently stopping the development of ASI. There are other potentially valid permutations of this conclusion - only improve AI capabilities when we're sure we can align said more-capable AI; only stop the emergence/development of misaligned ASI while somehow leaving aligned ASI development alone - that are similar because they all require similar levels of global coordination, at least for some time. I'll only discuss a permanent stop in this section; see here for (temporary) slow/pause.

The convenience of stopping ASI development is that it is both a total solution (works whether ASI alignment is tractable or intractable) and a legible goal: everyone knows what will happen if ASI development is stopped (we won't get ASI, misaligned or otherwise). You don't have to hand-wave anything; the actions and their outcomes are clear.

However, how to achieve this stop is deeply unclear, certainly hard, and maybe intractable. Many of society's incentive systems do not support it: capitalism, geopolitics, and indeed a finite-resource universe incentivize the acquisition of power. Many actors would prefer to gamble on ASI alignment for the chance to be the ones "owning" the ASI at the end of the race (as in, defection is encouraged).

In spite of these incentives, many also recognize the prisoner's dilemma they are in - that is, collaboration to halt ASI development is likely in everyone's best interests. Unfortunately, political will is not enough: it is not yet obvious how to implement an anti-ASI-development monitoring and enforcement apparatus. In the near term, monitoring seems hard, but possible: datacenters are big resource sinks, chips come from few places; with enough monitoring investment, it should be very hard to hide both (and what one is doing with them) from state intelligence.

Enforcement seems harder: it seems all major geopolitical entities would need to be ever-ready to gang up against any other entity on relatively short notice - and prevent the emergence of a too-powerful, unstoppable entity - for as long as humanity would like to stay alive. There do not appear to be any convenient mutually-assured-distruction-like incentives this time: I disagree that MAIM will work. MAD works because a nuclear power is incentivized not to do The Dangerous Action of firing their nukes, so nobody fires their nukes, but all powers are still incentivized to make nukes. In AI's case, everyone is incentivized to do The Dangerous Action of making ASI; it's not enough that actors are incentivized to prevent other actors from doing The Dangerous Action.

So, what should humanity do?

Evidently, both visible routes sit somewhere between "hard" and "impossible", at least on the timelines we appear to need them. Therefore, it seems we seek a stroke of luck. The implication of this is that we don't yet know enough to even know which area of investigation is the "right one" to put all our resources behind. This implication seems to lend itself to a try-many-things portfolio approach that neglects neither route, though there are 1) many considerations once you get specific on which path, and 2) several seemingly-useful helper actions that indirectly improve our chances of achieving success on one of the main paths.

A Partial List of Considerations

Helper Action: Play For Time

The main issue I see in a "permanent stop" approach is the difficulty in ensuring its permanence. To that end, aiming for a temporary pause or slowdown instead of a permanent one ameliorates this. Such a goal still needs to achieve near-future enforcement, of course, but this seems tractable. Implementing this approach would obviously buy safety-related research more time and increase its chance of success. In light of that, it seems independently worth striving for.

Helper Action: Lock Down Inference Access

As capabilities increase, preventing intentional misuse of model capabilities increases in importance. We've already seen two types of effort in this direction with Fable's recent release: delaying and pausing public access, and guardrails that attempt to block some capabilities of a model, while allowing others. I see misuse prevention as a subfield of alignment research; if it continues to be insufficient, blocking public access to capable-enough models seems necessary. And I see stably implementing denial of public model access as another option in the umbrella of international-coordination-heavy slow/pause/stop solutions, with similar work to be done.

So, what should I do?

Timeline Considerations

My choice is probably most impacted by how much time I believe is left to contribute. I claim this isn't something we can decisively predict yet (see appendix for slightly more discussion) and I have no reason to throw my hat in the forecasting ring other than because my career's impact depends on it. So I'd call myself medium-term on ASI time horizon: I put some weight on a <5-year-to-ASI fast takeoff-y scenario, some weight on a <20-year medium takeoff-y scenario, and the rest on slower than that. Call it ⅓ to each bucket to put numbers to it.

But even if my p(doom in 5 years) were 95%, that would not preclude some path that optimized for impact >5 years away if the magnitude of impact justified the 20x lower probability. Calling that out because there's a lot of urgency in the air right now to get your AI safety impact in before the singularity hits in a couple years, and frankly I worry that that line of reasoning would lead me and other new entrants to the field to underinvest or act rashly, jeopardizing their impact in >5 year scenarios.

Other considerations

Mid-July Conclusion

Updates (from mid-June)

Non-updates:

Mid-June Conclusion (old)

I am beginning to suspect that the right course of action for me is to aim for contributing to ASI (mis)alignment research directly. Main reasons:

Example paths to achieving that transition:

Next steps (last updated: mid-July)

While I felt I had come to a good enough stopping point to draft this doc and solicit takes, I still have a lot more to do, and readers shouldn't take my conclusions as set in stone yet. Todos:

Appendix

(Say It With Me) Predicting the Future is Hard

Generally, it is difficult to predict the future. It seems we don't yet have sufficient evidence to meaningfully collapse on any of the following. I'll add to/defend this list as relevant.

Define "alignment"

What does it look like for an ASI to be "aligned" to humanity? Half answers/relevant beliefs I loosely hold:

That's all for now