article
by Sam Altman
Reinforcement Learning Progress
Today, OpenAI released a new result. We used PPO (Proximal Policy Optimization), a general reinforcement learning algorithm invented by OpenAI, to train a team of 5 agents to play Dota and beat semi-pros.
We are not affiliated with Sam Altman. This unofficial audio edition is available to listen to and share for free. Read the original post on Sam Altman's blog.
Duration
~1 minutes (1K characters)
Release date
Jun 25, 2018