7 , which plots the average reward computed within a time window of 10 minutes, we observe that the average reward for DQN has an increasing trend in the Straight and Right-Turn routes, and it is consistently higher than the average reward obtained by Q and RANDOM.
← all excerpts
Reinforcement learning for online testing of autonomous driving systems: a replication and extension study.
2
—
—
The sentences
Again, despite the lack of statistical significance, DQN shows a positive trend in finding more violations quicker than RANDOM, especially in the Right-Turn route.