AI Safety Contribution Thesis
Last updated:
If you're reading this, it's probably to help me decide how to contribute to avoiding ASI misalignment; to that end, you can likely treat everything but the "So, what should I do" section as optional.
Foreword
This essay is an attempt to solidify my own opinion around how I can contribute to avoiding ASI misalignment, where "avoiding ASI misalignment" means avoiding human extinction or permanent mass suffering that occurs as a direct or indirect result of the effects of ASI. It's written in a confident style in order to make it easier for a reader to disagree, not because I'm confident of everything I've written.
There are important sub-definitions to "avoiding ASI misalignment" that deserve documents in their own right, like who can use the aligned ASI, and for what. I think these are mostly implementation details downstream from what I'm trying to decide right now, because I don't think these clarifications really change much about the overall problem (e.g. aligning an ASI to one human seems roughly as hard as all humans). More words in the appendix.
I also have a high-level posture that predicting the future is hard. Many folks in the AI safety community have collapsed to certain predictions of the future - like whether ASI alignment is easy, tractable, or intractable - and I more or less believe that all of these predictions do not meet the burden of proof required, especially given they are predictions over scientific progress that we have little precedent for. More words in the appendix.
Viable Approaches
There are two categories of ways to avoid the emergence of misaligned ASI: 1) achieving alignment of all ASI allowed to exist, or 2) permanently stopping ASI development.
Achieving ASI Alignment
Achieving ASI alignment requires such a thing to be achievable (AKA alignment is easy or tractable). Assuming it is, it requires the sort of technological breakthroughs it seems we aren't on track to achieve yet. It's possible that's not true, but since we have so little evidence one way or another, we can't be sure.
So if alignment is easy or tractable, by definition we'll get there by discovering and implementing the means of achieving that alignment, AKA doing alignment research.
This assumes that an aligned ASI, by definition, can and will monitor for and eliminate emergent misaligned ASI.
It also assumes that creating an aligned ASI is not "too much" harder than creating a misaligned one. In theory, humanity could dedicate all of its resources solely toward creating an aligned ASI, but in practice, we cannot count on that. Hand-wavily, 99% allocation of resources towards aligned ASI still loses if misaligned ASI is 100x easier to build.
Permanently Stopping ASI Development
If alignment is intractable, the only way humanity can see a good outcome is by permanently stopping the development of ASI. There are other potentially valid permutations of this conclusion - only improve AI capabilities when we're sure we can align said more-capable AI; only stop the emergence/development of misaligned ASI while somehow leaving aligned ASI development alone - that are similar because they all require similar levels of global coordination, at least for some time. I'll only discuss a permanent stop in this section; see here for (temporary) slow/pause.
The convenience of stopping ASI development is that it is both a total solution (works whether ASI alignment is tractable or intractable) and a legible goal: everyone knows what will happen if ASI development is stopped (we won't get ASI, misaligned or otherwise). You don't have to hand-wave anything; the actions and their outcomes are clear.
However, how to achieve this stop is deeply unclear, certainly hard, and maybe intractable. Many of society's incentive systems do not support it: capitalism, geopolitics, and indeed a finite-resource universe incentivize the acquisition of power. Many actors would prefer to gamble on ASI alignment for the chance to be the ones "owning" the ASI at the end of the race (as in, defection is encouraged).
In spite of these incentives, many also recognize the prisoner's dilemma they are in - that is, collaboration to halt ASI development is likely in everyone's best interests. Unfortunately, political will is not enough: it is not yet obvious how to implement an anti-ASI-development monitoring and enforcement apparatus. In the near term, monitoring seems hard, but possible: datacenters are big resource sinks, chips come from few places; with enough monitoring investment, it should be very hard to hide both (and what one is doing with them) from state intelligence.
Enforcement seems harder: it seems all major geopolitical entities would need to be ever-ready to gang up against any other entity on relatively short notice - and prevent the emergence of a too-powerful, unstoppable entity - for as long as humanity would like to stay alive. There do not appear to be any convenient mutually-assured-distruction-like incentives this time: I disagree that MAIM will work. MAD works because a nuclear power is incentivized not to do The Dangerous Action of firing their nukes, so nobody fires their nukes, but all powers are still incentivized to make nukes. In AI's case, everyone is incentivized to do The Dangerous Action of making ASI; it's not enough that actors are incentivized to prevent other actors from doing The Dangerous Action.
So, what should humanity do?
Evidently, both visible routes sit somewhere between "hard" and "impossible", at least on the timelines we appear to need them. Therefore, it seems we seek a stroke of luck. The implication of this is that we don't yet know enough to even know which area of investigation is the "right one" to put all our resources behind. This implication seems to lend itself to a try-many-things portfolio approach that neglects neither route, though there are 1) many considerations once you get specific on which path, and 2) several seemingly-useful helper actions that indirectly improve our chances of achieving success on one of the main paths.
A Partial List of Considerations
- What if well-meaning ASI alignment research uncovers (and publicizes) capabilities improvements, such that it speeds up capabilities timelines? Or, more precisely, what if it uncovers "more" capabilities improvements than alignment improvements? Ex: reducing AI's hallucination rate is likely to improve both alignment and capabilities - how can one decide if that's "worth it"?
- We don't have too many shots at implementing the global coalition necessary to enforce a pause/slow/stop on ASI development, and lots could go wrong on any given shot. If some head of state pisses off another head of state, and they call off the coalition, that may cause a delay in implementation that proves fatal.
- Research is a tool, not a result. Research into misalignment, for example, may be more useful in increasing political will for a pause/slow/stop policy than contributing to ASI alignment progress directly.
Helper Action: Play For Time
The main issue I see in a "permanent stop" approach is the difficulty in ensuring its permanence. To that end, aiming for a temporary pause or slowdown instead of a permanent one ameliorates this. Such a goal still needs to achieve near-future enforcement, of course, but this seems tractable. Implementing this approach would obviously buy safety-related research more time and increase its chance of success. In light of that, it seems independently worth striving for.
Helper Action: Lock Down Inference Access
As capabilities increase, preventing intentional misuse of model capabilities increases in importance. We've already seen two types of effort in this direction with Fable's recent release: delaying and pausing public access, and guardrails that attempt to block some capabilities of a model, while allowing others. I see misuse prevention as a subfield of alignment research; if it continues to be insufficient, blocking public access to capable-enough models seems necessary. And I see stably implementing denial of public model access as another option in the umbrella of international-coordination-heavy slow/pause/stop solutions, with similar work to be done.
So, what should I do?
Timeline Considerations
My choice is probably most impacted by how much time I believe is left to contribute. I claim this isn't something we can decisively predict yet (see appendix for slightly more discussion) and I have no reason to throw my hat in the forecasting ring other than because my career's impact depends on it. So I'd call myself medium-term on ASI time horizon: I put some weight on a <5-year-to-ASI fast takeoff-y scenario, some weight on a <20-year medium takeoff-y scenario, and the rest on slower than that. Call it ⅓ to each bucket to put numbers to it.
But even if my p(doom in 5 years) were 95%, that would not preclude some path that optimized for impact >5 years away if the magnitude of impact justified the 20x lower probability. Calling that out because there's a lot of urgency in the air right now to get your AI safety impact in before the singularity hits in a couple years, and frankly I worry that that line of reasoning would lead me and other new entrants to the field to underinvest or act rashly, jeopardizing their impact in >5 year scenarios.
Other considerations
- I have no experience in either high-level approach (governance/policy or alignment
research), so it's unlikely that I would correctly invest in the best sub-route if left
to my own devices, at first. I will learn faster and do more under the auspices of
others.
- That said, there is a wide spread in learning efficiency across jobs. A product engineer at an AI safety org will learn little relevant to AI safety research. An infra engineer may learn slightly more, depending, but not be able to make the transition to research quickly enough to be worth it. So in both cases, if AI safety researcher is the goal, those jobs are poor starts.
- Doing work someone else tells me to do cedes some control to the expert on what I learn and on where I have impact. This is probably not my strongest counterfactual position in the long run.
- My talents/experiences/desires align more closely to technical topics. Contributing to research (for some definition of research) thus seems like a more productive way to leverage myself, no matter which high-level approach it is in service of.
- I am not financially independent, though current costs are low and can be sustained for a while.
- SF is where my support system is.
- There is a moral hazard in doing AI safety research for a company that makes money
from AI capabilities research (i.e. all frontier labs), as I would become incentivized
to see the company succeed.
- Though, just because there's more hazard in working for a frontier lab, doesn't mean there isn't some hazard pretty much everywhere. Where would most AI safety organizations be if the frontier labs didn't exist? What if the useful impact to AI safety is great enough to risk some moral hazard? And so on.
Mid-July Conclusion
Updates (from mid-June)
- I personally view working on slow/pause proposals to be a lot more salient than I
used to, primarily because I've been incrementally more persuaded on their feasibility.
Reasons for updating:
- Feedback that a slow/pause would directly increase the chance of ASI alignment research succeeding. Previously, I simply hadn't been thinking much about slow/pause vs. stop.
- The US government's willingness to pause Fable and the interest it increasingly signals in regulating inference is more than I was expecting.
- The release of ai-2040.com gave me a detailed scaffold upon which I could construct my belief on feasibility in greater detail (though there's much there I don't agree with).
Non-updates:
- The observation that I could work on increasing alignment research funding didn't ultimately update my beliefs much, because I think I'm poorly positioned for multiple reasons, but especially because at least some research experience seems better. I feel similarly about field-building. I am generally amenable to more assist-style actions, though, and am still keeping an eye out for them.
Mid-June Conclusion (old)
I am beginning to suspect that the right course of action for me is to aim for contributing to ASI (mis)alignment research directly. Main reasons:
- my background and interests align me towards technical work more so than politics
- I am accumulating a growing volume of living counterexamples to the intuition that one needs a PhD-length time investment before one can meaningfully contribute to AI safety research
- Developing the capability to do research is a foundational skill for AI Safety as an overall field; it preserves optionality. I can always go implement something that needs implementing after some time doing research, but the opposite is not true.
Example paths to achieving that transition:
- Transitioning to research engineer seems feasible for me, perhaps with some 0-12
months of investment, depending on the org. I have a few data points of orgs where the
research engineers are/essentially become researchers in a way that I feel has just as
much impact as the scientists.
- For example, METR's eval and frontier lab auditing work is mostly engineering work (no math formulas required for 99% of the labor). I'm counting this type of work in my sweeping generalization of "AI safety research".
- Fieldbuilding programs exist to grow AI safety researchers that could conceivably accept me after some demonstration of conviction in the transition, such as the beginnings of a body of completed research.
- I have a small handful of examples in the wild of relatively isolated software engineers doing research on-own and getting research science jobs. I think this is a bad path, as it's too long and doesn't have enough feedback, but it still deserves mention.
Next steps (last updated: mid-July)
While I felt I had come to a good enough stopping point to draft this doc and solicit takes, I still have a lot more to do, and readers shouldn't take my conclusions as set in stone yet. Todos:
- Investigate meaningful work to be done in furtherance of slow/stop policy.
- Filtering for people that actually care, I've only heard about a dozen people's
opinions/theories of change on AI safety, and need to push those numbers up.
- I also need reviews of this doc from such individuals (update: I think it would be better to hear people's theories than have them hear/read mine, now).
- Build a deeper understanding of the state of AI safety research: determine which subfields/topics seems like the best choice to invest in, or determine the set of good-enough subfields/topics if I have to compromise on perceived best in order to ramp faster.
Appendix
(Say It With Me) Predicting the Future is Hard
Generally, it is difficult to predict the future. It seems we don't yet have sufficient evidence to meaningfully collapse on any of the following. I'll add to/defend this list as relevant.
- How "hard" ASI alignment will be (e.g. easy, tractable but hard, intractable)
-
Obviously, there's a coherent, logical argument to be made that ASI alignment is functionally impossible, even with what we know now (see If Anyone Builds It Everyone Dies). And to be clear, I mostly agree with this - I'm interested in AI safety for a reason.
Unfortunately, the argument gets most of its strength not from a rigorous proof of ASI alignment impossibility, but from persuasiveness of the "safety" of the stop-ASI-development plan. Generally, the burden of proof for something being impossible should be quite high, and always treated with suspicion; humans have a long track record of producing persuasive, and ultimately false, impossibility arguments (see epistemic learned helplessness). So we can't eliminate the other possibilities just yet. And if stop-ASI-development were easy, I'd be in favor of a total redirect to that approach, but as I said above, I don't think that's the situation.
-
- a timeline of when ASI will happen (though certain capabilities increases follow predictable curves)
- what level of capabilities pushes us into extinction-level ASI
- What area of AI safety investment, if any, is highest ROI
Define "alignment"
What does it look like for an ASI to be "aligned" to humanity? Half answers/relevant beliefs I loosely hold:
- Perfect alignment of ASI itself seems to be a superhuman capability, and the
difficulty of achieving this capability scales with the rest of the ASI's capability.
- CEV-shaped meta-alignment or similar seems necessary to avoid monkey's paw-like side effects we wouldn't have wanted. No 10 perfectly-followed rules will do, here.
- The more capable a model, the more easily it can cause harm, obviously. Going one step further: the more capable a model, the harder it will be to avoid accidental extinction as a side effect of exercising its capabilities, likely requiring ever-increasing capabilities in things like outcome prediction, etc. For example, humans today are capable enough to manufacture pandemic-grade pathogens, and so are also capable enough to accidentally release one; this was not a problem 500 years ago.
- For an ASI to be aligned to any individual, it will still need to be CEV-shaped,
just with a population size of one.
- An ASI aligned to any individual is almost certainly not satisfactory ASI alignment for most, if not many. But it probably achieves some "at least not all of us are dead"-tier alignment.
- There are an infinite range of other scenarios that have an accompanying range of
acceptability, with regard to the goal of "no human extinction, no extreme suffering"
(examples: earth is a zoo, humans are hooked up to dopamine pumps, humans are forced
into a simulation, etc.). I'm not sure how to reason about these as a set.
- With no evidence, I loosely believe that all of these eventualities aren't too likely; that our future collapses onto either "we have nailed alignment as perfectly as is possible" or "we are dead". And I'm not sure how the existence of this set could meaningfully update any decisions I need to make now-ish (my choices should still aim toward a "nail alignment" outcome).
That's all for now