A deep reinforcement learning algorithm that uses two critic networks and delayed policy updates for stable continuous control.