Pith. sign in

REVIEW 7 cited by

Multi-agent Reinforcement Learning in Sequential Social Dilemmas

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1702.03037 v1 pith:YJGTBYAM submitted 2017-02-10 cs.MA cs.AIcs.GTcs.LG

Multi-agent Reinforcement Learning in Sequential Social Dilemmas

classification cs.MA cs.AIcs.GTcs.LG
keywords dilemmassocialgamepoliciessequentialagentsgamesintroduce
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Matrix games like Prisoner's Dilemma have guided research on social dilemmas for decades. However, they necessarily treat the choice to cooperate or defect as an atomic action. In real-world social dilemmas these choices are temporally extended. Cooperativeness is a property that applies to policies, not elementary actions. We introduce sequential social dilemmas that share the mixed incentive structure of matrix game social dilemmas but also require agents to learn policies that implement their strategic intentions. We analyze the dynamics of policies learned by multiple self-interested independent learning agents, each using its own deep Q-network, on two Markov games we introduce here: 1. a fruit Gathering game and 2. a Wolfpack hunting game. We characterize how learned behavior in each domain changes as a function of environmental factors including resource abundance. Our experiments show how conflict can emerge from competition over shared resources and shed light on how the sequential nature of real world social dilemmas affects cooperation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. The Reciprocity Gradient

    cs.LG 2026-05 unverdicted novelty 7.0

    The reciprocity gradient allows agents to learn near-optimal context-sensitive policies by analytically propagating reward gradients through reputation chains in multi-agent settings.

  2. Unsupervised Causal Abstractions Discovery

    cs.LG 2026-06 unverdicted novelty 6.0

    Low-rank graphs induce latents that form causal abstractions, with identifiability results and a practical objective enabling unsupervised learning of high-level SCMs from low-level measurements.

  3. Investigating the Impact of Subgraph Social Structure Preference on the Strategic Behavior of Networked Mixed-Motive Learning Agents

    cs.MA 2026-04 unverdicted novelty 6.0

    Preferences over local subgraph structures cause distinct changes in reward collection and strategic actions for agents playing sequential social dilemmas in Harvest and Cleanup environments.

  4. Do Agent Societies Develop Intellectual Elites? The Hidden Power Laws of Collective Cognition in LLM Multi-Agent Systems

    cs.MA 2026-04 unverdicted novelty 6.0

    LLM agent societies develop power-law coordination cascades and intellectual elites through an integration bottleneck that grows with system size.

  5. Memory-Induced Supra-Competitive Outcomes Between Deep Reinforcement Learning Agents in Optimal Trade Execution

    q-fin.CP 2026-05 unverdicted novelty 5.0

    In a two-agent Almgren-Chriss liquidation game, deep RL agents given intra-episode history of prices and own actions achieve supra-competitive outcomes more frequently and persistently than agents without such memory.

  6. Learning Incentive Structures for Cooperative Resilience in Multi-Agent Systems under Social Dilemmas

    cs.MA 2026-01 unverdicted novelty 5.0

    A method infers resilience-promoting reward functions via trajectory scoring and integrates them into MARL, with hybrid incentives shown to reduce collapse in disrupted resource environments.

  7. AI Researchers Must Help Lead Arms Control to Mitigate Military AI Risks

    cs.CY 2026-06 unverdicted novelty 2.0

    AI researchers must lead technical research in arms control to mitigate risks from military AI systems, drawing lessons from nuclear deterrence.