Uniform ε\varepsilon-equilibrium in finite multiplayer stochastic games

OPENMajorConjectureProposed c. 1981 · Full conjecture

Canonical statement

Consider any stochastic game with a finite player set II, finite state set SS, finite nonempty action set Ai(s)A_i(s) for each player ii at state ss, bounded stage payoff gi(s,a)g_i(s,a) for each action profile aiIAi(s)a\in\prod_{i\in I}A_i(s), and transition law q(s,a)q(\,\cdot\mid s,a) on SS. For every ε>0\varepsilon>0 and initial state s0s_0, there exist a behavioral-strategy profile σ\sigma and NNN\in\mathbb N such that, for every horizon nNn\ge N, every player ii, and every unilateral behavioral deviation τi\tau_i,
Es0,σ ⁣[1nt=1ngi(st,at)]Es0,(τi,σi) ⁣[1nt=1ngi(st,at)]ε, \mathbf E_{s_0,\sigma}\!\left[\frac1n\sum_{t=1}^n g_i(s_t,a_t)\right]\ge \mathbf E_{s_0,(\tau_i,\sigma_{-i})}\!\left[\frac1n\sum_{t=1}^n g_i(s_t,a_t)\right]-\varepsilon,
where (st,at)(s_t,a_t) is the state--action process generated by the indicated strategy profile and transition law.
View source LaTeX
Consider any stochastic game with a finite player set \(I\), finite state set \(S\), finite nonempty action set \(A_i(s)\) for each player \(i\) at state \(s\), bounded stage payoff \(g_i(s,a)\) for each action profile \(a\in\prod_{i\in I}A_i(s)\), and transition law \(q(\,\cdot\mid s,a)\) on \(S\). For every \(\varepsilon>0\) and initial state \(s_0\), there exist a behavioral-strategy profile \(\sigma\) and \(N\in\mathbb N\) such that, for every horizon \(n\ge N\), every player \(i\), and every unilateral behavioral deviation \(\tau_i\),
\[
\mathbf E_{s_0,\sigma}\!\left[\frac1n\sum_{t=1}^n g_i(s_t,a_t)\right]\ge \mathbf E_{s_0,(\tau_i,\sigma_{-i})}\!\left[\frac1n\sum_{t=1}^n g_i(s_t,a_t)\right]-\varepsilon,
\]
where \((s_t,a_t)\) is the state--action process generated by the indicated strategy profile and transition law.

A stochastic game, introduced by Shapley [Shapley1953Stochastic], is played in stages: a state from a finite set determines the finite action sets of finitely many players, whose joint action produces bounded stage payoffs and a lottery over the next state. The conjecture asserts that for every ε>0\varepsilon>0 and initial state there is a uniform ε\varepsilon-equilibrium: one behavioral-strategy profile that no player can improve upon by more than ε\varepsilon, in expected average payoff, simultaneously in every sufficiently long horizon. The question took shape around 1981, emerging from the uniform-value theory of stochastic games rather than from a single dated statement.

Mertens and Neyman proved that every finite two-player zero-sum stochastic game has a uniform value, settling that case [MertensNeyman1981Games]. Deep positive results also cover two-player nonzero-sum games; see the survey by Vieille [Vieille2002Recent]. More recent work includes equilibrium existence for two-player games with shift-invariant payoffs [FleschSolan2023Equilibrium].

For three or more players the two-player and zero-sum techniques do not directly apply to an arbitrary finite player set, and neither a proof nor a counterexample is known; existence of uniform ε\varepsilon-equilibria for an arbitrary finite number of players remains open.

The boxed statement is the canonical open formulation — not a stronger variant or a related research program. The status reflects the catalog's last review; do your own literature search before investing serious effort.