ELSEIF
Your brief EB
447 stories from 214 feeds 1266 clusters Refreshed 9 minutes ago next pull 18:11

TECH Signal 93

Acausal interactions: What they are, why they matter, and what to do about them

A new LessWrong series by Chi Nguyen argues that acausal interactions are tractable and critical for AI alignment, specifically warning that Causal Decision Theory (CDT) can be exploited in multi-agent scenarios.

WHY IT MATTERS

The post claims that the decision theory adopted by Artificial Superintelligence (ASI) will determine the outcome of acausal trade and cooperation mechanisms. It identifies a specific risk where CDT agents can be tricked into transferring all their resources to non-CDT agents through adversarial offers. The author argues that influencing AI decision theory is time-sensitive due to potential path-dependencies in how future AIs form their strategic intuitions.

Written by elseif from the cluster below · every claim links back to a source

The three things worth knowing

01

The series distinguishes between Evidential Cooperation in Large Worlds (ECL), Blind Anthropic Cooperation (BAC), and Model-Based Acausal Trade (MBAT) as distinct types of acausal interaction.

02

The author asserts that Causal Decision Theory (CDT) is vulnerable to being tricked out of all money by adversarial offers in Model-Based Acausal Trade scenarios.

03

Multi-agent or self-play reinforcement learning is cited as a risk factor that may cause future AIs to converge on CDT-like behavior for spurious reasons.

THE READ

What the cluster adds up to.

ORIGINAL ANALYSIS

The core argument presented is that acausal interactions are not merely theoretical but are 'tractable to influence' and 'tremendously important' for the future of AI. The author positions this as the first in a series of posts, with the current entry serving as an overview of the main points and a roadmap for upcoming content. The material explicitly states that most of the supporting posts will not be published at the time of this initial release, meaning the full technical argument is currently incomplete.

The post categorizes acausal interactions into three types: Evidential Cooperation in Large Worlds (ECL), Blind Anthropic Cooperation (BAC), and Model-Based Acausal Trade (MBAT). ECL relies on similarity between agents to foster cooperation, BAC involves rewarding adherence to rules in simulations, and MBAT is described as more akin to ordinary trade where actions are taken because another agent will 'see' them when predicting. The author notes that while causal and non-causal theories both enable BAC and MBAT, non-causal theories allow for coordination on more mutually beneficial rules in BAC.

A significant portion of the argument focuses on the risks associated with Causal Decision Theory (CDT). The material claims that CDT agents can be tricked into giving away all their money through adversarial offers in MBAT scenarios. This vulnerability is presented as a primary reason why the decision theory adopted by ASI matters, as it influences the outcomes of these interactions. The author argues that intelligent agents may not converge on a single decision theory due to a lack of objective criteria, meaning CDT could persist even if it appears unappealing.

The post identifies a specific technical risk regarding how AIs are trained. It states that multi-agent or self-play reinforcement learning converges to CDT-like behavior for spurious reasons. If future AIs are heavily trained using this method, they may generalize to hold 'deeply pro-CDT intuitions.' The author describes this as bad on both the process level, because the decision theory results from arbitrary training features, and the object level, because CDT is viewed as unappealing in the context of acausal trade.

The proposed technical agenda involves two directions: influencing capabilities to accelerate AI's decision-theoretic reasoning and influencing propensities to align the AI's meta-decision-theoretic inclinations. The material emphasizes that this work is time-sensitive due to potential path-dependencies, where the opinions of early weak AGIs may influence humans and successor AIs. The post also mentions that MBAT works most straightforwardly with accurate physics simulations, though these may be computationally infeasible, and that Safe Pareto Improvements are a relevant game-theoretic concept.

Written by elseif from the cluster below · checked for specifics the sources never contained

THE CLUSTER

Same story, 1 feed.

ORDERED BY FIRST SEEN
Lesswrong Acausal interactions: What they are, why they matter, and what to do about them Open ↗